Skip to content
AI Job Risk

Analysis · 4 September 2026 · 7 min read

GPT-6 Astra benchmarks and your job

OpenAI says Astra opens the AGI era. The benchmarks that bear on work tell a narrower story, and independent aggregates disagree with the vendor tables.

OpenAI released GPT-6 Astra on 3 September 2026 and announced it as the world's most intelligent and aligned model, with a launch line welcoming everyone to the AGI era. Two days earlier Anthropic had shipped Claude Fable 5.1.

Strip the framing and there are two separate questions. Is Astra the strongest model available? And does anything in it change the case that AI is coming for particular kinds of work? The answers are not the same, and the second one does not depend on the first.

The claim, and who disagrees with it

OpenAI's own tables are impressive and mostly real. The reported figures include 97.6 percent on FrontierMath Tier 4 against 87.8 for Fable 5.1, effectively saturating ARC-AGI-3 at 99.9 percent where Opus 5 manages 30.2, and 96.0 percent on GPQA Diamond. On long context it holds 96.3 percent on an eight needle retrieval test between 512K and 1M tokens, where the previous OpenAI model dropped to 73.8.

Independent measurement is less flattering. In the Artificial Analysis Intelligence Index, Astra scores 61.2 against Fable 5.1's 65.7 and Opus 5's 63.1, which puts the world's most intelligent model third on the aggregate that nobody at either lab controls. On Humanity's Last Exam with tools it posts 57.2 percent against Fable 5.1's 65.0. On GDPval-AA v2, the benchmark built from real occupational deliverables, it scores 1,629 against Fable 5.1's 1,853.

So the superlative is a claim about a particular set of tables. That is worth knowing before you read anything else about this launch, including ours.

The number that actually bears on work

Now the part that does matter, and it is not on the leaderboard everyone quoted.

AutomationBench: 41.4. Fable 5.1 scores 31.4 on the same test and the previous OpenAI model scores 18.1. Astra leads it outright, and among widely reported benchmarks this is the one closest to the question of whether a model can carry a workflow rather than answer about one.

OSWorld 2.0: 72.6 percent, at roughly 40 minutes per task, against 65.7 percent at roughly 75 minutes for GPT-5.6 Sol. The accuracy gain is modest. The time is nearly halved. On browser tasks OpenAI reports completion around 1.9 times faster.

That second pair of numbers is the economically interesting one, and it is routinely ignored because speed is boring next to a maths score. A model that does a task at the same quality in half the time halves the cost of doing it. Adoption decisions are made on cost against a wage, not on capability in the abstract, so a speed halving moves the line for a set of tasks that were previously too expensive to hand over. Nothing about the intelligence ranking changes that.

Astra also became the first model OpenAI has designated as crossing the Critical cybersecurity threshold under its Preparedness Framework, reporting 100 percent on ExploitBench and 88.0 percent on a single attempt site reliability benchmark where Fable 5.1 scores 12.5. Whatever you make of the framing, security operations is a profession, and that is a professional capability appearing suddenly.

What it does not show

Two guardrails, both from the same tables.

Terminal-Bench 4.0 puts Astra at 57.7 and Fable 5.1 at 55.8, a two point gap on sustained multi-step work after a generational release. On DeepSWE, Astra's 74.1 percent is beaten by Meta's Muse Spark 1.3 at 75.4. At the frontier of long horizon work the labs are trading small margins, not running away.

And a benchmark is still a benchmark. Nothing in this release is evidence about employment. The evidence about employment comes from payroll records and job cut filings, and it says something narrower and more specific than any launch does: AI was cited in US job cuts covering 205,000 workers through August 2026, concentrated in customer service, data operations, entry level software and finance back offices, while Stanford payroll data shows the decline landing on workers aged 22 to 25 inside exposed occupations rather than across them.

How to read a model launch against your own job

Three habits, worth more than any single score.

Ignore the aggregate, find the task. Whether a model ranks first or third changes very little. Whether it can now finish a specific piece of your week unattended changes everything. Read for benchmarks built on work products, mainly GDPval, OSWorld and AutomationBench, and skip the exam scores.

Watch cost and speed, not only accuracy. Astra is priced at 10 dollars per million input tokens and 50 per million output. The task time halving is the release's real economic content.

Distinguish capability from deployment. Capability arrives at a launch. Deployment arrives when the standard tools of your field ship a feature that does your task rather than assists it. Our post on timing covers why the second lags the first by years, and unevenly.

Our index does not move on model releases. It aggregates published research across 961 occupations, and a launch is not research. What a launch does is tell you which direction the underlying capability is travelling, and this month it travelled toward bounded, tool using, multi-step work performed fast. The occupations built on that kind of output already sit high in our ranking: Data Entry Keyers (63) and Bookkeeping, Accounting, and Auditing Clerks (58) among them. Fable 5.1's results, read for the same thing, point the same way from the other lab.

The occupation average is a floor, not a verdict on you. What decides your exposure is how much of your week is the part these benchmarks now measure.

Where does your job sit?

Every number above is an occupation average. Your own exposure depends on your seniority, your sector and how much of your day is routine. Answer a few questions and get your personal score free.

Get my personal score

Free · No credit card · No signup

Figures come from the AI Job Risk Index and were current when this was published. Scores change as new research is ingested, so the index is always the live version. See how scoring works. Informational guidance based on published research, not professional career or financial advice.