Arabic routing benchmark
The exam every routing policy sits before it goes live
Sadu routes each request to the model most likely to answer it well. To prove a new routing policy is actually better — not just different — every candidate sits the same fixed exam: an Arabic-first set spanning Modern Standard Arabic, Gulf dialect, Arabic–English code-switching, plus code, math, extraction, and Saudi-context questions. Most items carry a reference answer, so the LLM judge grades against ground truth instead of vibes.
Methodology
- •The exam is FIXED and versioned — scores are only comparable because the questions never drift.
- •Each item is routed by the candidate policy exactly as live traffic would be, then judged (with the reference answer when one exists) on a 0–1 scale.
- •Spend is penalized: a policy that buys quality with brute-force cost scores lower.
- •Promotion additionally requires real-traffic evidence: a canary slice of production requests runs the candidate and its observed reward is compared against live.
Exam composition
arabic · 7 itemsgeneral · 5 itemsarabic-dialect · 2 itemsmath · 2 itemscode · 2 itemsextraction · 1 itemscreative · 1 items
Latest published run
Results publish here after each evaluated policy run. The exam and methodology are public now; scores follow with the first funded evaluation.