Analysis · 29 September 2026 · 5 min read
Claude Sonnet 5.5 vs Opus 5.5
Near identical on professional work, half the price per token. Where Opus 5.5 still leads, and why Sonnet 5.5 is not always cheaper per finished task.
Sonnet 5.5 and Opus 5.5 score almost the same on professional work (1,844 against 1,846 on GDPval) and Sonnet costs half as much per token. Opus is still ahead on the hardest coding and reasoning tests. Whether Sonnet is actually cheaper per finished task is less clear than the price list suggests.
Side by side
| Sonnet 5.5 | Opus 5.5 | |
|---|---|---|
| Released | 28 Sept 2026 | 22 Sept 2026 |
| Input, per million tokens | $2 | $4 |
| Output, per million tokens | $10 | $20 |
| GDPval-AA v2.1 | 1,844 | 1,846 |
| Terminal-Bench 4.0 | 70.6% | 66.4% |
| CursorBench 4.0 | 55.5% | 57.8% |
| FrontierCode 1.1 (max) | 46.2% | 54.4% |
| Humanity's Last Exam | 64.5% | 67.7% |
| Artificial Analysis Intelligence Index | 56 | 58 |
Benchmark and price figures from Anthropic's Sonnet 5.5 and Opus 5.5 pages. Index scores from Artificial Analysis.
Which is better?
For most work they are close enough that price decides. Sonnet leads on Terminal-Bench, long command line work, and ties on GDPval, which is built from real deliverables such as reports and spreadsheets. Opus leads by about eight points on FrontierCode and three on Humanity's Last Exam, the tests closest to hard, open ended problems. Anthropic says Opus remains stronger at work "requiring sustained judgment".
Is Sonnet 5.5 actually cheaper?
Per token, yes: exactly half. Per finished task, it depends on how much the model thinks.
Anthropic says Sonnet 5.5 costs up to 30 percent less than Sonnet 5 for most work. Artificial Analysis found the opposite at maximum effort: about 193,000 output tokens per task, the most it has measured, around 60 percent more than Opus 5.5 at the same setting, for a cost of 7.60 dollars per task, roughly 50 percent above Sonnet 5.
Two things reconcile these. Anthropic describes typical use, while Artificial Analysis ran the model at its maximum effort setting, where it spends far more tokens. And Artificial Analysis tested a pre release version with a bug affecting structured outputs, so its figures may change on a rerun.
The practical point stands either way: the price that matters is the price of a finished piece of work, not of a token. We made the same argument about GPT-6 Astra and Opus 5.5.
What this means for jobs
Two releases in one week have put professional level output at two price points instead of one. The cheaper that output gets, the more tasks cross the line where handing them to a model beats paying a person to do them. That line moves fastest for work that is high volume, clearly specified and easy to check.
Our index covers 961 occupations and does not change with model launches. It tells you which roles are built around that kind of work. The assessment tells you how much of your own week is, in about two minutes and without an account.
Where does your job sit?
Every number above is an occupation average. Your own exposure depends on your seniority, your sector and how much of your day is routine. Answer a few questions and get your personal score free.
Get my personal scoreFree · No credit card · No signup
Figures come from the AI Job Risk Index and were current when this was published. Scores change as new research is ingested, so the index is always the live version. See how scoring works. Informational guidance based on published research, not professional career or financial advice.