Analysis · 4 September 2026 · 7 min read
Claude Fable 5.1 benchmarks and your job
Fable 5.1 posts 1,853 Elo on GDPval against a human expert baseline of 1,000. What that measures, where the model still fails, and what both mean for jobs.
Anthropic released Claude Fable 5.1 on 1 September 2026, two days before OpenAI shipped GPT-6 Astra. Most of the coverage that followed ranked the two against each other. That is the wrong question for anyone reading this site. The useful question is narrower: did anything in this release move the measurement of whether a machine can do a job, and if so, which jobs?
One number says yes, and it is not a coding score.
The benchmark that is actually about work
Almost every AI benchmark measures something adjacent to work. Exam questions, puzzle grids, code snippets. GDPval is the exception, which is why it is the only benchmark on this page worth your attention.
OpenAI built GDPval from 1,320 tasks drawn from 44 occupations across the nine largest sectors of the US economy. The tasks were written by professionals averaging fourteen years of experience in the occupation, and they are real deliverables rather than questions: analyse this financial report and build the stakeholder deck, review this contract for risk, produce the spreadsheet. A model's output and a human expert's output are then compared blind, and the score is how often the model's work is judged as good or better.
Artificial Analysis runs a version of it, GDPval-AA v2, over 220 of those tasks with shell and browser access, and converts the blind comparisons into an Elo rating. The human expert baseline is anchored at 1,000.
Claude Fable 5.1 scores 1,853. Claude Opus 5 scores 1,824. GPT-6 Astra, released two days later, scores 1,629.
Read that anchor again, because it is the entire story. A rating above 1,000 means the model's deliverables were preferred to the human expert's more often than not, on tasks that experts in those occupations wrote to represent their own work. Artificial Analysis notes the models complete some of these tasks roughly a hundred times faster than the human experts and at a fraction of the cost.
Where the jump actually landed
The rest of the Fable 5.1 numbers matter mainly for what they say about the shape of the improvement. Anthropic's reported figures cluster the gains in one place: long, multi-step, tool-using work.
- Terminal-Bench 4.0: 55.8 percent, against 42.0 for Fable 5 and 52.3 for Opus 5. A near fourteen point jump in one generation on sustained terminal work.
- Terminal-Bench-Science 0.1: 52.6 percent, up 27.9 points on Fable 5. The largest single move in the release.
- OSWorld 2.0: 77.9 percent on the partial setting, 41.7 on strict. Real desktop tasks, driven through a real computer.
- AutomationBench: 31.4 percent, up 14.3 points on Fable 5.
- Humanity's Last Exam: 60.9 percent without tools, up only 3.1 points.
That last line is the tell. Three points on a hard knowledge exam, twenty eight on sustained scientific terminal work. The models are not getting much smarter in the way a test measures smart. They are getting substantially better at finishing long pieces of work without a person holding the thread, which is the capability that maps onto a job rather than onto a question.
Anthropic shipped the product to match. Claude Cowork launched in January 2026 as an agent aimed at people who do not write code, reading and editing files in folders you point it at, and expanded to web and mobile in July. Anthropic's own usage data showed most Cowork users were not coding. The agentic capability left the engineering department some time ago.
What the same numbers say against the story
Take the strict OSWorld figure seriously: 41.7 percent. Under strict grading, the best model available fails the majority of ordinary computer tasks. AutomationBench sits at 31.4. These are not the numbers of a system that can be handed a job.
There is a real tension between that and the GDPval result, and it resolves in a way that matters for your planning. Models are at or above expert level on bounded deliverables with clear inputs, the kind of thing that arrives as a well specified request. They remain unreliable at open ended operation over a real environment, the kind of thing that fills an actual working week. Most jobs are a mix, and the mix is what decides your exposure, not the job title.
That is the same split our index is built on, and it is why we publish a task level view rather than a verdict. Two people with the same title and different weeks sit in different places. What the research agrees on covers how little of this literature supports the confident framing it usually gets, and two different things are called AI exposure covers why two studies of the same occupation routinely disagree.
Capability is not the clock
A benchmark result is the earliest possible signal, not the arrival. Between a model scoring well and a role changing sit procurement, integration, liability, retraining and the cost of the tokens, and those run on different schedules in different industries. An MIT study of computer vision automation found that at then-current costs only about a quarter of the wages attached to technically automatable vision tasks would have been economical to automate at all.
The 2026 layoff wave is where the two lines finally met for some functions. AI or automation was cited in US job cuts covering 205,000 workers through August, concentrated in customer service, data operations, entry level software and finance back offices. Our analysis of that wave compares those categories with our scores occupation by occupation, and they line up.
Fable 5.1 does not change any score in our index, because our scores come from published research on occupations rather than from model releases. What a release like this changes is the direction of travel behind those numbers. If your week is mostly bounded deliverables produced to a specification, the GDPval line is about you, and it moved this month.
Across the 961 occupations we publish, the roles built around exactly that kind of output already carry the highest scores. Customer Service Representatives (68) and Data Entry Keyers (63) sit near the top for reasons that predate this release and are only reinforced by it.
Your own number is not the occupation average, though. It depends on how much of your week is the bounded part.
Where does your job sit?
Every number above is an occupation average. Your own exposure depends on your seniority, your sector and how much of your day is routine. Answer a few questions and get your personal score free.
Get my personal scoreFree · No credit card · No signup
Figures come from the AI Job Risk Index and were current when this was published. Scores change as new research is ingested, so the index is always the live version. See how scoring works. Informational guidance based on published research, not professional career or financial advice.