Skip to content
AI Job Risk

Analysis · 8 October 2026 · 5 min read

Claude Haiku 5.5 benchmarks and your job

Haiku 5.5 scores 1,620 on GDPval, above GPT-6 Astra, at a hundredth of the output price. The numbers, the catch, and what it means for jobs.

Claude Haiku 5.5, released by Anthropic on 7 October 2026, is the cheapest model Anthropic sells: 10 cents per million input tokens and 50 cents per million output. On GDPval, the benchmark built from real professional deliverables, it scores higher at maximum effort than GPT-6 Astra, which charges a hundred times more per output token. For anyone tracking what AI means for jobs, that is the part of the release that matters. What counted as flagship work in early September now sits at the bottom of the price list.

The benchmark numbers

All figures are Anthropic's reported results.

BenchmarkHaiku 5.5Haiku 4.5GPT-6 LunaSonnet 5.5
GDPval-AA v2.1 (Elo)1,6207351,4371,840
OSWorld 2.1, offline (computer use)72.4%15.7%48.9%83.9%
Terminal-Bench 4.039.2%0.0%16.4%70.6%
Humanity's Last Exam, with tools57.4%18.7%not given64.5%

The jump from Haiku 4.5 is large on every row, and in Anthropic's table Haiku 5.5 beats GPT-6 Luna, OpenAI's model at the same price, on every test where both are listed. Anthropic is also clear about the limits. Its launch page says Sonnet 5.5 and Opus 5.5 remain the better choice for complex agentic coding, and positions Haiku 5.5 for narrowly scoped work: summarising, classifying, and running as a subagent under a larger model.

The comparison that matters

GDPval is made of real work tasks across 44 occupations, written by people with years in the job: reports, spreadsheets, presentations, legal reviews. Artificial Analysis grades models on it by blind comparison of the finished deliverables. Its current leaderboard, at each model's highest setting:

ModelGDPval-AA v2.1 (Elo)Output price per million tokens
Claude Opus 5.51,866$20
Claude Sonnet 5.51,839$10
Gemini 4 Argon (high)1,626$10
Claude Haiku 5.51,620$0.50
GPT-6.1 Sol1,575$10
GPT-6 Astra1,542$50
GPT-6 Luna1,432$0.50

Argon's price is Google's introductory rate, and Argon is not yet generally available. The scale is anchored to another model, not to human experts, so it ranks models against each other rather than against a person. Within that ranking, the cheapest Claude now sits level with Google's new frontier model and above OpenAI's.

The catch: effort and tokens

That 1,620 is at maximum effort, and the setting matters. Haiku 5.5 is the first Haiku with adjustable effort, and on the same leaderboard it scores 1,420 at high and 1,277 at medium, which is below Luna at its top setting.

Maximum effort also means a lot of thinking. Artificial Analysis measured about 162,000 output tokens per task on its intelligence index at max effort, more than Opus 5.5 and roughly three times GPT-6 Luna. So the price per task falls less than the price per token suggests. Artificial Analysis has not yet published a cost per task, because its site does not reflect Anthropic's new tiered pricing, where requests above 100,000 tokens cost five times more.

Who tested it, and on what

Anthropic's early tester reports are about office work, not code. HubSpot reported 92.8 percent on its CRM test suite, the best score it has seen there. Rogo called it accurate enough to trust for pulling segment revenue out of annual 10-K filings. AlphaSense measured 0.84 against 0.76 for Haiku 4.5 across 400 research queries, and Box reported 11 points higher than Haiku 4.5 at about half the latency. All of these are vendor selected examples.

For reference, our scores for the two roles that kind of work describes: Financial and Investment Analysts (56) and Customer Service Representatives (68).

One name on that list was in the news for another reason. The day before launch, HubSpot announced it was cutting about 660 jobs, 7 percent of its staff. Its CEO, Yamini Rangan, told CBS Boston the decision was not driven by efficiencies from AI, and described it as a shift toward delivering customer outcomes with AI. That is a company reorganising around an AI strategy, which is a different thing from replacing staff with a model. Our tracker of AI layoffs in 2026 keeps the two apart for the same reason.

What the release does not show

Independent testing is less tidy than the launch page. Artificial Analysis gives Haiku 5.5 a score of 43 on its intelligence index, up 26 points from the last Haiku a year ago and 13 behind Sonnet 5.5. On its own run of AutomationBench, which tests whole business workflows, Haiku 5.5 scored 35 percent, against between 53 and 60 percent for GPT-6 Luna and other small models. Artificial Analysis says a pre release refusal bug likely understates that score and plans to rerun it.

And no benchmark measures employment. Anthropic's own labour research found no systematic rise in unemployment among the most exposed workers so far, with an early slowdown in hiring for workers aged 22 to 25.

Where your job sits

The pattern across the last five weeks is the story. GPT-6 Astra launched on 3 September at 50 dollars per million output tokens. Sonnet 5.5 brought flagship scores to the mid tier on 28 September, GPT-6.1 Sol did the same for OpenAI the next day, and now a model at a hundredth of Astra's output price beats it on real work deliverables. Capability arrives at the top and moves down the price list within weeks, and adoption decisions are made on cost against a wage.

Our index scores 961 occupations from published research, and no score moves because a model launched. What a release like this changes is how cheaply the tasks at the top of that index can be done. The assessment shows how much of your own week looks like those tasks. It takes about two minutes, free, no account.

Where does your job sit?

Every number above is an occupation average. Your own exposure depends on your seniority, your sector and how much of your day is routine. Answer a few questions and get your personal score free.

Get my personal score

Free · No credit card · No signup

Figures come from the AI Job Risk Index and were current when this was published. Scores change as new research is ingested, so the index is always the live version. See how scoring works. Informational guidance based on published research, not professional career or financial advice.