Analysis · 8 October 2026 · 5 min read
GPT-6.1 Sol benchmarks and your job
GPT-6.1 Sol nearly matches GPT-6 Astra at a fifth of the price and beats it on GDPval. What OpenAI reported, what testing found, and what it means for jobs.
OpenAI released GPT-6.1 Sol on 29 September 2026, one week after GPT-6 Sol, and priced it at a fifth of GPT-6 Astra: 2 dollars per million input tokens and 10 per million output. OpenAI's pitch is near Astra intelligence for a fifth of the price. Independent testing mostly agrees, and on the benchmark built from real professional work it does slightly better than near. For jobs, the release is less about what AI can do than about what it now costs to do it.
What OpenAI reported
OpenAI's launch post states most of its results as comparisons rather than raw scores:
- DeepSWE v1.1, software engineering in real codebases: matches Astra at roughly a fifth of the cost.
- GDP.pdf, answering professional questions from complex PDFs in finance, healthcare, law and seven other fields: scores above Claude Opus 5.5 at less than half the cost per task, and approaches Astra at about a fifth.
- AutomationBench 1.0.6, whole workflows using 47 tools across sales, marketing, operations, support, finance and HR: 2.2 points above Opus 5.5 at medium effort, at roughly a third of the cost.
- OSWorld 2.0, long computer use tasks: within 2.1 points of Astra at maximum effort, at about a seventh of the cost per task.
- Terminal-Bench Science 0.1, data analysis and simulation: 5.47 dollars per task against 23.21 for Opus 5.5 and 23.80 for Astra, though Astra still has the top score at 68.1 percent.
Each comparison is at the effort setting OpenAI chose to highlight, so read them as the best case. OpenAI also reports fewer factual errors on hard prompts, down from 11.4 to 7.7 percent at low effort, measured on conversations where users had flagged an earlier model's mistake.
The independent check
Two measurements from Artificial Analysis bear on work.
On its intelligence index, GPT-6.1 Sol scores 52 at maximum effort, one point below Astra, at less than a quarter of Astra's cost per task. It costs 72 cents per task to run, and it is slow, at about 55 output tokens per second.
On GDPval, real work deliverables across 44 occupations graded blind, the current leaderboard reads:
| Model | GDPval-AA v2.1 (Elo) | Price per million tokens, in / out |
|---|---|---|
| Claude Sonnet 5.5 | 1,839 | $2 / $10 |
| GPT-6.1 Sol | 1,575 | $2 / $10 |
| GPT-6 Astra | 1,542 | $10 / $50 |
| GPT-6 Sol | 1,510 | $2 / $10 |
GPT-6.1 Sol beats Astra on real deliverables at a fifth of the price, which is the result OpenAI's pitch implies. The same table shows the limit of the pitch: Claude Sonnet 5.5, at exactly the same list price, scores 264 points higher. The difference shows up in the bill instead. Artificial Analysis puts Sonnet 5.5 at 5.46 dollars per index task against Sol's 72 cents, because Sonnet used about six times as many output tokens across the index, 420 million against 67 million. At the same price per token, one model is better at the work and the other is far cheaper per task, and which matters depends on the job.
Why the price is the story
Astra launched on 3 September at 50 dollars per million output tokens. Four weeks later, a model at a fifth of that price matches it on the aggregate index and beats it on GDPval. Claude Haiku 5.5 then went further on 7 October, at a hundredth of Astra's output price. Capability arrives at the top of the price list and moves down within weeks.
That matters because adoption is decided on cost against a wage, not on capability in the abstract. A task that was not worth handing to Astra at 23.80 dollars a run looks different at 5.47. Our post on what Opus 5.5 costs per piece of work goes through that arithmetic.
AutomationBench is the result closest to ordinary office jobs, since its workflows are built on the tools of sales, marketing, support, finance and HR.
Two of the occupations doing that work, as our index scores them: Market Research Analysts and Marketing Specialists (64) and Human Resources Specialists (53).
What the release does not show
Sol is available in ChatGPT Work and Codex and through the API, not yet in the standard chat app, so most people will not meet it directly. And no benchmark measures employment. The evidence on jobs still comes from payroll records and job cut filings, and it is narrower than any launch: Anthropic's labour research found no systematic rise in unemployment among the most exposed workers so far, with an early slowdown in hiring for workers aged 22 to 25. Our post on timing covers why deployment lags capability by years.
Where your job sits
Our index scores 961 occupations from published research, and no score moves because a model launched. What a cheaper model changes is which of the tasks at the top of that index are now worth automating. The assessment shows how much of your own week looks like those tasks. It takes about two minutes, free, no account.
Where does your job sit?
Every number above is an occupation average. Your own exposure depends on your seniority, your sector and how much of your day is routine. Answer a few questions and get your personal score free.
Get my personal scoreFree · No credit card · No signup
Figures come from the AI Job Risk Index and were current when this was published. Scores change as new research is ingested, so the index is always the live version. See how scoring works. Informational guidance based on published research, not professional career or financial advice.