Sadu
DocsTrust
Sign in

Arabic routing benchmark

The exam every routing policy sits before it goes live

Sadu routes each request to the model most likely to answer it well. To prove a new routing policy is actually better — not just different — every candidate sits the same fixed exam: an Arabic-first set spanning Modern Standard Arabic, Gulf dialect, Arabic–English code-switching, plus code, math, extraction, and Saudi-context questions. Most items carry a reference answer, so the LLM judge grades against ground truth instead of vibes.

Methodology

  • The exam is FIXED and versioned — scores are only comparable because the questions never drift.
  • Each item is routed by the candidate policy exactly as live traffic would be, then judged (with the reference answer when one exists) on a 0–1 scale.
  • Spend is penalized: a policy that buys quality with brute-force cost scores lower.
  • Promotion additionally requires real-traffic evidence: a canary slice of production requests runs the candidate and its observed reward is compared against live.

Exam composition

arabic · 7 itemsgeneral · 5 itemsarabic-dialect · 2 itemsmath · 2 itemscode · 2 itemsextraction · 1 itemscreative · 1 items

Latest published run

Results publish here after each evaluated policy run. The exam and methodology are public now; scores follow with the first funded evaluation.