Skip to content
AI Job Risk

Analysis · 22 September 2026 · 6 min read

Claude Opus 5.5 benchmarks and your job

Opus 5.5 now leads GDPval, the benchmark built from real professional work. What its numbers say about jobs, and what they cannot.

Anthropic released Claude Opus 5.5 on 22 September 2026, three weeks after Claude Fable 5.1 and GPT-6 Astra. The headline is unusual for a model launch: the smaller, cheaper model now beats the larger flagship on many of the tests that matter for work.

This post reads the release the way this site reads every release. Not which model wins, but whether anything moved the measurement of what a machine can do inside a job.

The numbers Anthropic published

All figures below are Anthropic's own reported results. Vendor tables flatter the vendor, so treat them as the upper end of what independent testing will find.

BenchmarkOpus 5.5Fable 5.1Opus 5GPT-6 Astra
Terminal-Bench 4.066.4%55.8%52.3%57.9%
FrontierCode v1.154.4%50.3%48.0%53.3%
GDPval-AA v2.1 (Elo)1,8461,7351,7081,542
OSWorld 2.0, partial81.8%80.7%74.0%not reported

The largest move is on Terminal-Bench 4.0, a test of sustained multi-step work in a real terminal: from 52.3 percent for Opus 5 to 66.4 percent in one release, ten and a half points clear of Fable 5.1.

The row that is about your job

GDPval is the one benchmark in that table built from real occupational work: 1,320 tasks across 44 occupations, written by professionals averaging fourteen years in the role, graded on the deliverable rather than on an answer. Artificial Analysis runs a version of it and scores models by blind pairwise comparison.

Opus 5.5 now leads that leaderboard at 1,846, ahead of Fable 5.1 at 1,735 and far ahead of GPT-6 Astra at 1,542.

One honest note on the scale, because it matters for how you read it. The current version, v2.1, is anchored to another model rather than to human experts, so 1,846 tells you Opus 5.5 beats other models, not by how much it beats a person. On the previous version, which was anchored to human experts, the top models already sat well above the human baseline, a result our Fable 5.1 post covers. Opus 5.5 is ahead of those models. The direction is not in doubt; the exact distance from a human professional is not something this number measures.

What changed, in one sentence

The gains keep landing in the same place: long, multi-step, tool using work, the kind that fills a working week rather than answering a question. That is the pattern since Fable 5.1, and Opus 5.5 extends it while costing less to run and finishing long tasks in a fraction of the time.

For jobs, that combination matters more than any single score. Capability tells you a task can be handed over. Price and speed decide whether it is.

What it does not tell you

A benchmark is a controlled test, not a workplace. Nothing in this release is evidence about employment. The evidence on employment says something narrower: Anthropic's own economists found no systematic rise in unemployment among highly exposed workers so far, alongside an early slowdown in hiring at the entry level, and only eleven of 756 occupations show AI doing more than half their tasks in real use.

Our index does not move on launches. It aggregates published research across 961 occupations. What a launch like this changes is the direction of travel for the roles built on bounded, well specified output. Computer Programmers (59) and Customer Service Representatives (68) already sit near the top of our ranking for reasons this release only reinforces.

Where you sit

Every figure above describes models, and every figure in our index describes occupations. Neither describes you. Two people with the same title and different weeks have different exposure, and the difference is usually larger than the difference between two jobs.

Our free assessment measures that difference. It asks about your seniority, your sector and how much of your week is routine, then scores your actual task mix against the research baseline for your occupation. It takes about two minutes, with no signup and no card.

Where does your job sit?

Every number above is an occupation average. Your own exposure depends on your seniority, your sector and how much of your day is routine. Answer a few questions and get your personal score free.

Get my personal score

Free · No credit card · No signup

Figures come from the AI Job Risk Index and were current when this was published. Scores change as new research is ingested, so the index is always the live version. See how scoring works. Informational guidance based on published research, not professional career or financial advice.