
Claude Opus 5 Is Live on Starchild: What It Beats and What It Doesn't
TL;DR: Anthropic's new flagship, Claude Opus 5, is live on Starchild. It scores a point above Claude Fable 5 on independent testing while charging half as much per token. It is also slower and wordier than the models it competes with, so it earns its place on hard work rather than high-volume work.
What actually changed
Opus 5 replaces Opus 4.8 at exactly the same price, $5 per million input tokens and $25 per million output tokens.
There was a modest jump on benchmarks. Opus 5 scored 61 on the Intelligence Index, eclipsing Opus 4.8 by 5 points.
For reference, Claude Fable 5 costs double and only scores 60. Running the full benchmark suite cost $3,836 on Opus 5 against $5,631 on Fable 5.
Where it does well
Anthropic's pitch is that Opus 5 checks its own work instead of stopping at the first plausible answer. It gave some examples as part of the launch:
Given a drawing of a machine part and no way to view it, Opus 5 wrote its own image-processing code to read the geometry out of the pixels, then rebuilt the part in 3D. No competing model in the same test solved it in five attempts.
Given a real bug in a widely used open-source tool, it found the underlying cause and closed an edge case that the community's own patch had missed.
Early users reported the same pattern in numbers:
| Reported by | Result |
|---|---|
| Zapier | 100% pass rate on its automation benchmark, without spending more tokens than earlier Claude models |
| Box | 8% better than Opus 4.8 on content analysis, 17% on due diligence work |
| Lovable | 22% better than Opus 4.7 on the hardest coding tasks, with far less variance between runs |
| A financial modeling team | 9 points more accurate, using a third fewer steps and 60% less time |
The time and step reductions matter as much as the accuracy gains, because the cost of running additional steps on a flagship model can become astronomical.
Where it falls short
Opus 5 is slow. It produces 55 tokens per second, whereas Fable 5 manages 74 and a fast model like Gemini 3.6 Flash does 251. You definitely feel that when vibe-coding something. You will want to get a few tabs going and walk away for a bit.
It is also wordy. Running the same benchmark suite, Opus 5 generated 100 million tokens where Fable 5 needed 87 million and the field average is 63 million.
Looking at OpenAI, GPT-5.6 Sol scores 59 on the Intelligence Index, two points below Opus 5, and costs $2,824 to run the same suite against Opus 5's $3,836. If your work is flexible, Sol is the better value.
How it compares
| Model | Price per 1M in / out | Intelligence | Speed |
|---|---|---|---|
Claude Opus 5 |
$5 / $25 | 61 | 55/s |
Claude Fable 5 |
$10 / $50 | 60 | 74/s |
GPT-5.6 Sol |
$5 / $30 | 59 | 65/s |
Kimi K3 |
$3 / $15 | 57 | 32/s |
Claude Opus 4.8 |
$5 / $25 | 56 | 57/s |
GLM-5.2 |
$1.40 / $4.40 | 51 | 170/s |
Gemini 3.6 Flash |
$1.50 / $7.50 | 50 | 251/s |
Intelligence scores and speeds from Artificial Analysis, measured at each model's maximum reasoning setting.
When to pick it
Opus 5 is worth its price on work where a wrong answer costs more than a slow one: debugging that needs a root cause, long coding sessions, analysis across large document sets, and any workflow that has to run end to end without someone checking each stage.
For everything that runs often and fast, keep a cheaper model on the job. On Starchild your agent can switch between them, so the practical setup is a fast model handling the routine loops and Opus 5 handling the steps that decide whether the work was any good.
Theoretically, you can also turn its effort down. Anthropic says that even at its lowest setting, Opus 5 passes more Zapier automation tasks than any other model. In Starchild you get low, medium, and high, starting on medium.
Claude Opus 5 is available now. Select it in your agent's model settings, or use Conductor Mode to let Starchild choose per task.
Try Claude Opus 5 on Starchild
Sources: Anthropic's Claude Opus 5 announcement, Anthropic's model documentation, and Artificial Analysis.
Claude Opus 5
GPT-5.6 Sol
Kimi K3
GLM-5.2
Gemini 3.6 Flash