Skip to content
AI Job Risk

Analysis · 29 September 2026 · 5 min read

Claude Sonnet 5.5 benchmarks and your job

Sonnet 5.5 scores 1,844 on GDPval, two points behind Opus 5.5, at half the price. The numbers, what independent testing found, and what it means for jobs.

Claude Sonnet 5.5, released by Anthropic on 28 September 2026, scores within two points of Opus 5.5 on the benchmark built from real professional work, at half the price. For anyone tracking what AI means for jobs, that is the part of the release that matters: flagship level output is now available at the mid tier price.

The benchmark numbers

All figures are Anthropic's reported results.

BenchmarkSonnet 5.5Opus 5.5Sonnet 5
GDPval-AA v2.1 (Elo)1,8441,8461,449
Terminal-Bench 4.070.6%66.4%10.3%
CursorBench 4.055.5%57.8%34.1%
FrontierCode 1.1 (max)46.2%54.4%42.4%
Humanity's Last Exam64.5%67.7%54.9%

Sonnet 5.5 beats Opus 5.5 on Terminal-Bench, a test of sustained multi step work in a terminal, and trails it on the harder coding and reasoning tests. Anthropic's own description is that Opus stays "clearly stronger at complex, open-ended work requiring sustained judgment".

Why GDPval is the row to watch

GDPval is made of 1,320 real work tasks across 44 occupations, written by people with an average of fourteen years in the job: reports, spreadsheets, presentations, legal reviews. Artificial Analysis grades models on it by blind comparison of the finished deliverables.

Sonnet 5.5 moved from 1,449 to 1,844 in one version, almost exactly matching Opus 5.5. One caution on the scale: the current version is anchored to another model, not to human experts, so the score ranks models against each other rather than against a person. On the earlier human anchored version, top models already rated above the expert baseline, which our Fable 5.1 post covers.

What changed for jobs

Nothing about what AI can do changed much this week. Opus 5.5 did most of this six days earlier. What changed is the price of doing it. Sonnet 5.5 is listed at 2 dollars per million input tokens and 10 per million output, half of Opus 5.5, and the same price as Sonnet 5.

Two of Anthropic's early tester reports are about ordinary office work rather than code: a support team whose tickets were "processed 20% faster", and a financial services test where Sonnet 5.5 used about 121,000 tokens per answer against Sonnet 5's 497,000. Both are vendor selected examples. Both point at the same jobs our index already scores highest.

For reference, our scores for the two roles those examples describe: Customer Service Representatives (68) and Financial and Investment Analysts (56).

What the release does not show

Independent testing is less tidy than the launch post. Artificial Analysis ranked Sonnet 5.5 second on its intelligence index, two points behind Opus 5.5, but found that at maximum effort it used more output tokens per task than any model it has measured, and cost more per task than Sonnet 5. It tested a pre release version with a known bug and plans to rerun. Sonnet 5.5 vs Opus 5.5 goes through what that means for cost.

And no benchmark measures employment. Anthropic's own labour research found no systematic rise in unemployment among the most exposed workers so far, with an early slowdown in hiring for workers aged 22 to 25.

Where your job sits

Our index scores 961 occupations from published research, and no score moves because a model launched. What a release like this changes is how cheaply the tasks at the top of that index can be done. The assessment shows how much of your own week looks like those tasks. It takes about two minutes, free, no account.

Where does your job sit?

Every number above is an occupation average. Your own exposure depends on your seniority, your sector and how much of your day is routine. Answer a few questions and get your personal score free.

Get my personal score

Free · No credit card · No signup

Figures come from the AI Job Risk Index and were current when this was published. Scores change as new research is ingested, so the index is always the live version. See how scoring works. Informational guidance based on published research, not professional career or financial advice.