Skip to content
AI Job Risk

Analysis · 29 September 2026 · 5 min read

Claude Sonnet 5.5 vs GPT-6 Astra

Sonnet 5.5 beats GPT-6 Astra on terminal work and professional work benchmarks at a fifth of the token price. Where Astra is still stronger.

On the published numbers, Claude Sonnet 5.5 beats GPT-6 Astra on both the terminal work benchmark and the professional work benchmark, at a fifth of Astra's price per token. Astra, OpenAI's flagship, launched on 3 September 2026. Sonnet 5.5, Anthropic's mid tier model, launched on 28 September.

The comparison

Claude Sonnet 5.5GPT-6 Astra
Input, per million tokens$2$10
Output, per million tokens$10$50
GDPval-AA v2.1 (Elo)1,8441,542
Terminal-Bench 4.070.6%57.9%

Sources: Anthropic for Sonnet 5.5's prices and scores and for Astra's Terminal-Bench score as listed on the Opus 5.5 page; Artificial Analysis for Astra's GDPval score; Astra's price as reported at its launch, covered in our Astra post.

The table is short on purpose. Most other benchmarks were reported by one lab for its own model and not the other, or on different versions of a test. Putting those side by side would compare numbers that were never measured the same way.

Where Astra is stronger

Astra's own launch figures lead on things this table leaves out: frontier mathematics, where it scored 97.6 percent on FrontierMath Tier 4, and security operations, where OpenAI rated it the first model past its Critical cybersecurity threshold. It also has a one million token context window. Our Astra post covers those results.

Neither model is simply better. Astra is stronger at the frontier of reasoning. Sonnet 5.5 is stronger at the everyday professional output that makes up most office jobs, and much cheaper to run.

Why the price gap matters more than the scores

For deciding whether a task moves from a person to a model, the relevant cost is per finished task, not per token, and it depends on how many tokens each model spends. Still, a five times difference in list price is large enough to change which tasks are worth automating.

The work that moves first is high volume, well specified and cheap to check. The 205,000 US job cuts citing AI through August 2026 were concentrated in exactly that work: customer service, data operations, entry level software and finance back offices.

Our scores for two of those roles: Customer Service Representatives (68) and Data Entry Keyers (63).

Which one to use

For long command line work and document style deliverables, the published numbers favour Sonnet 5.5, at a fraction of the cost. For the hardest maths and security work, Astra. For most people reading this, the more useful question is which parts of their own job look like the tasks both models now do well.

Our index scores 961 occupations from published research. The assessment scores yours, from your seniority, your sector and how routine your week is. About two minutes, no account needed.

For the same comparison inside Anthropic's range, see Sonnet 5.5 vs Opus 5.5.

Where does your job sit?

Every number above is an occupation average. Your own exposure depends on your seniority, your sector and how much of your day is routine. Answer a few questions and get your personal score free.

Get my personal score

Free · No credit card · No signup

Figures come from the AI Job Risk Index and were current when this was published. Scores change as new research is ingested, so the index is always the live version. See how scoring works. Informational guidance based on published research, not professional career or financial advice.