
Starchild's Smart Router benchmarks #1 on a value index rank of top models
We benchmarked 9 systems, including 6 flagships, 2 competing routers, and Smart Routing, on the same agent loop with live tools. We then ranked them on a Value Index that combines difficulty-weighted accuracy with cost per correct answer.
01Accuracy versus cost
Models were scored on accuracy (vertical axis) and cost (horizontal axis). The ideal position is top-left: high accuracy at low cost. Flagships sit in the top-right, with nearly the same accuracy at 3–9× the cost.
02Value Index ranking
The ranking uses a Value Index: 65% difficulty-weighted accuracy and 35% cost-efficiency. Each metric is scaled so the best system in the field scores 1 and the worst scores 0, then the two scores are combined.
| # | Model | Value | Weighted acc. | $/correct |
|---|---|---|---|---|
| 1 | ★ Smart Router | 0.822 | 91.5% | $0.0032 |
| 2 | DeepSeek V4 Pro | 0.720 | 93.7% | $0.0100 |
| 3 | GPT-5.5 | 0.677 | 93.9% | $0.0210 |
| 4 | Kimi K2.6 | 0.671 | 92.4% | $0.0095 |
| 5 | Qwen3.7 Max | 0.663 | 91.5% | $0.0071 |
| 6 | OpenRouter Auto | 0.550 | 82.8% | $0.0024 |
| 7 | Claude Opus 4.8 | 0.382 | 86.9% | $0.0290 |
| 8 | Grok 4.3 | 0.379 | 83.3% | $0.0051 |
| 9 | OpenRouter Fusion | 0.000 | 77.9% | $0.0500 |
Reference only, unranked: GLM-5.2 (router component, 91.9%, $0.0035) · MiniMax M3 (component, 81.1%, $0.0020) · Gemini 3.5 Flash (92.4%, cost excluded because a logging bug returned $0).
03The hardest suite: where we don't win
On AIME 2025, three flagships reached 100% versus our 96.7%, while costing 3–6× more per correct answer. Our claim is the best accuracy per dollar across the full test set, rather than a win on every suite.
04Routing by difficulty
Across the difficulty gradient, Smart Router stays near the top of the field on medium and hard tiers while keeping its cost curve low and flat. It pays frontier prices only when the estimated difficulty warrants escalation and routes everything else to cheaper models. Flagships pay frontier prices on every request, including the easy tier, where 97% of models clear 80% and accuracy does not distinguish between them.
05Limitations and interpretation
Not for every workload
Where a wrong answer costs far more than compute, paying 9× for a flagship can be rational. Re-run the index with your own α. The sensitivity check lets you test the result directly.
AIME is saturating
Several models cluster at ~100% on AIME 2025, so that suite mostly separates on cost, not capability. AIME 2026 (uncontaminated) is in the dashboard.
Snapshot, not a durable ranking
Models and pricing move fast. This describes the field on July 1, 2026. We'll re-run as it changes.