Skip to content
AI Job Risk

Analysis · 7 September 2026 · 6 min read

GPT-6 Astra vs Claude Fable 5.1

Two flagship models, same headline price, opposite strengths. Compared on the benchmarks built from real occupational work rather than on exam scores.

OpenAI shipped GPT-6 Astra on 3 September 2026, two days after Anthropic shipped Claude Fable 5.1. Both cost the same at the headline rate, 10 dollars per million input tokens and 50 per million output. Both were announced as the best model available. Only one of them can be.

The comparison below is not about which lab wins. It is about a narrower question this site exists to answer: if these systems are going to absorb parts of jobs, which one is actually better at work, and how would you tell from a benchmark table.

The scoreboard, and why it splits

MeasureGPT-6 AstraClaude Fable 5.1
Artificial Analysis Intelligence Index v4.1.16166
GDPval-AA v2, Elo, human expert = 1,0001,6291,853
AutomationBench41.431.4
Terminal-Bench 4.057.755.8
Humanity's Last Exam, with tools57.2%65.0%
FrontierMath Tier 497.6%87.8%
ExploitBench100%70%
SRE-Bench, single attempt88.0%12.5%

Figures as reported by Artificial Analysis and compiled from the vendors' own tables by Vellum. One benchmark is deliberately absent from the table: OSWorld. Anthropic reports Fable 5.1 at 77.9 percent on a partial setting and 41.7 on strict, while OpenAI reports Astra at 72.6 percent as a single headline figure alongside a claim about task time. Those are not the same measurement and lining them up in a column would be the exact trick this post is arguing against.

Read the rows and the pattern is not "one model is better". Astra wins the frontier reasoning rows outright, saturating FrontierMath and effectively finishing ARC-AGI-3, and it wins security operations by margins that are hard to overstate: 88.0 against 12.5 on single-attempt site reliability work. Fable 5.1 wins the aggregate index, wins tool-using exam work, and wins GDPval by 224 Elo.

The row that matters for jobs

If you only look at one line, look at GDPval.

OpenAI built it from 1,320 tasks across 44 occupations in the nine largest sectors of the US economy, written by professionals averaging fourteen years in the role. They are deliverables rather than questions: the stakeholder deck, the contract review, the model in a spreadsheet. Artificial Analysis runs 220 of them with shell and browser access and scores blind pairwise comparisons as an Elo rating, anchoring the human expert baseline at 1,000.

Fable 5.1 sits at 1,853 and Astra at 1,629. Both are far above the human anchor, which is the fact worth carrying away from this entire comparison. On the benchmark built specifically to look like professional output, the argument stopped being whether models reach expert level and became which model is further past it.

Artificial Analysis also recorded Astra dropping around 80 Elo points on this benchmark relative to expectations set by its other results. A model can saturate frontier mathematics and still lose ground on producing the deliverable a working professional would produce. Those are different skills and GDPval is the one that looks like a job.

Where Astra genuinely leads

Two rows deserve more weight than the index gap suggests.

AutomationBench: 41.4 against 31.4. Among widely reported benchmarks this is the closest thing to a test of whether a model can carry a workflow rather than answer questions about one, and Astra leads it clearly.

Token efficiency, which is really a cost story. Artificial Analysis found Astra uses roughly a tenth fewer tokens than its predecessor at max effort on the Index, and about a third of the tokens in Codex coding tests. The consequence is that despite identical headline pricing, Astra comes in at less than half the cost of Fable 5.1 for equivalent performance on their Coding Agent Index. Cheaper per finished piece of work at the same list price is exactly how automation gets adopted, and it does not show up anywhere in an intelligence ranking. What that costs against a wage is the question underneath it.

What neither table tells you

Both labs published numbers that flatter their own model, and the independent aggregate disagrees with OpenAI's "most intelligent" framing by putting Astra third behind Fable 5.1 and Opus 5. That is worth knowing, and it still does not answer the question most readers arrived with.

A benchmark measures capability under test conditions. Whether your role changes depends on adoption, and adoption runs on procurement, integration, liability and cost against a wage, on schedules that differ by industry and lag capability by years. Our post on timing covers why. The evidence that anything is happening to employment comes from payroll records and job cut filings, not launches: AI was cited in US job cuts covering 205,000 workers through August 2026, concentrated in a handful of functions.

Neither release moves a single score in our index, because we aggregate published research across 961 occupations and a launch is not research. What both releases do is confirm the direction: the capability gains are landing on bounded, tool using, multi-step work delivered fast. Read each model's results in full in our companion posts on GPT-6 Astra and Claude Fable 5.1, or start with where your own occupation sits.

Where does your job sit?

Every number above is an occupation average. Your own exposure depends on your seniority, your sector and how much of your day is routine. Answer a few questions and get your personal score free.

Get my personal score

Free · No credit card · No signup

Figures come from the AI Job Risk Index and were current when this was published. Scores change as new research is ingested, so the index is always the live version. See how scoring works. Informational guidance based on published research, not professional career or financial advice.