Starchild's Smart Router benchmarks #1 on a value index rank of top models

Starchild's Smart Router benchmarks #1 on a value index rank of top models

We benchmarked 9 systems, including 6 flagships, 2 competing routers, and Smart Routing, on the same agent loop with live tools. We then ranked them on a Value Index that combines difficulty-weighted accuracy with cost per correct answer.

Scoring methodology: 126 problems, using AIME 2025 and LiveCodeBench for objective scoring, data accurate as of July 1, 2026
#1Value Index rank of 9 systems (0.822)
91.5%weighted accuracy, within about 2 points of the top flagships
$0.0032per correct answer vs $0.021 (GPT-5.5), $0.029 (Opus 4.8)
63smedian task time, the second fastest in the field

01Accuracy versus cost

Models were scored on accuracy (vertical axis) and cost (horizontal axis). The ideal position is top-left: high accuracy at low cost. Flagships sit in the top-right, with nearly the same accuracy at 3–9× the cost.

Fig 1 · Capability vs cost per correct answer
Weighted accuracy (y) vs $/correct, log scale (x) · everyday + AIME, 126 problems
Smart Routing Flagships Competing routers Components (reference, unranked)

02Value Index ranking

The ranking uses a Value Index: 65% difficulty-weighted accuracy and 35% cost-efficiency. Each metric is scaled so the best system in the field scores 1 and the worst scores 0, then the two scores are combined.

Fig 2 · Value Index, all ranked systems
α = 0.65 · components and Gemini 3.5 Flash excluded from ranking (see notes)
Table 1 · Full leaderboard
Value Index ranking · α = 0.65 · everyday + AIME · 126 common problems
#ModelTypeValueWeighted acc.Raw acc.$/correctSpeed
1★ Smart Routerrouter
0.822
91.5%85.5%$0.003263.3s
2DeepSeek V4 Proflagship
0.720
93.7%87.4%$0.0100163.9s
3GPT-5.5flagship
0.677
93.9%87.4%$0.0210116.3s
4Kimi K2.6flagship
0.671
92.4%85.5%$0.0095116.7s
5Qwen3.7 Maxflagship
0.663
91.5%87.1%$0.0071244.6s
6OpenRouter Autorouter
0.550
82.8%82.7%$0.002497.0s
7Claude Opus 4.8flagship
0.382
86.9%84.6%$0.029041.0s
8Grok 4.3flagship
0.379
83.3%81.5%$0.005169.0s
9OpenRouter Fusionrouter
0.000
77.9%78.6%$0.0500175.4s

Reference only, unranked: GLM-5.2 (router component, 91.9%, $0.0035) · MiniMax M3 (component, 81.1%, $0.0020) · Gemini 3.5 Flash (92.4%, cost excluded because a logging bug returned $0).

Low rank ≠ low capability. Claude Opus 4.8 is solidly mid-pack on accuracy (90% AIME). It ranks 7th for one reason: at $0.029/correct, it is the most expensive system in the field, costing about 9× more than Smart Router despite slightly lower accuracy. With accuracy this close, cost drives the ranking.

03The hardest suite: where we don't win

On AIME 2025, three flagships reached 100% versus our 96.7%, while costing 3–6× more per correct answer. Our claim is the best accuracy per dollar across the full test set, rather than a win on every suite.

Fig 3 · AIME 2025: accuracy by system
30 competition problems, objective scoring · bars: accuracy (left axis) · dashed line: $/correct (right axis)

04Routing by difficulty

Across the difficulty gradient, Smart Router stays near the top of the field on medium and hard tiers while keeping its cost curve low and flat. It pays frontier prices only when the estimated difficulty warrants escalation and routes everything else to cheaper models. Flagships pay frontier prices on every request, including the easy tier, where 97% of models clear 80% and accuracy does not distinguish between them.

Fair-comparison rules. Router components (GLM-5.2, MiniMax M3) are excluded from the ranking because a router beating its own components is true by construction. Gemini 3.5 Flash is excluded from cost/value math due to a $0 cost-logging bug; its accuracy is reported for reference. Every system ran the identical production agent loop with live tools and objective scoring.

05Limitations and interpretation

Not for every workload

Where a wrong answer costs far more than compute, paying 9× for a flagship can be rational. Re-run the index with your own α. The sensitivity check lets you test the result directly.

AIME is saturating

Several models cluster at ~100% on AIME 2025, so that suite mostly separates on cost, not capability. AIME 2026 (uncontaminated) is in the dashboard.

Snapshot, not a durable ranking

Models and pricing move fast. This describes the field on July 1, 2026. We'll re-run as it changes.

Starchild Internal Research · Value Index = α·norm(capability) + (1−α)·norm(cost-efficiency), α = 0.65. Capability = difficulty-weighted accuracy; cost-efficiency = $/correct answer. Suite-level leaderboards (AIME 2026, LiveCodeBench, Everyday) available in the underlying dashboard.