Analysis · 7 September 2026 · 7 min read
Can GPT-6 Astra do your job?
Astra finishes a benchmark task for between $0.63 and $2.57. That number, not the benchmark score, decides which work gets handed over and which does not.
Every model launch produces the same question and the same unsatisfying answers. OpenAI released GPT-6 Astra on 3 September 2026 with a line about the AGI era. The coverage argued about benchmark tables. Almost nobody quoted the number that actually decides whether a machine does a piece of your work, which is what the machine charges to finish it.
That number is now public, and it is small.
What Astra costs per finished task
Astra is priced at 10 dollars per million input tokens and 50 per million output, up two and a half times from the 4 and 20 dollars its predecessor charged. Per token, it got considerably more expensive.
Per finished piece of work, Artificial Analysis measures the cost of completing one task on its Intelligence Index:
| Effort setting | Cost per completed task |
|---|---|
| Low | $0.63 |
| Medium | $1.16 |
| High | $1.41 |
| Extra high | $1.85 |
| Maximum | $2.57 |
Between sixty three cents and two dollars fifty seven, depending on how hard you let it think, for a task drawn from a benchmark suite of hard reasoning problems.
Hold that against the thing it is implicitly being compared to. Not another model. An hour of a person's time.
Why the per token price is the wrong number
The token price went up and the cost of finished work went down, which sounds contradictory until you notice they measure different things.
Artificial Analysis found Astra uses roughly a tenth fewer tokens than its predecessor at maximum effort, and about a third of the tokens in coding tests. Fewer tokens to reach the same finished output means the list price can rise while the bill per completed job falls. On their Coding Agent Index, Astra comes in at less than half the cost of Claude Fable 5.1 for equivalent performance, despite the two carrying identical headline pricing.
This is the mechanism worth understanding, because it is the one that will keep operating quietly after the launch coverage stops. Every efficiency gain lowers the price of a finished task without any change to the sticker price, and the price of a finished task is what a manager compares to a wage.
The comparison that actually decides jobs
Capability decides whether a task can be handed over. Cost decides whether it is.
An MIT study of computer vision automation put a number on that gap: at the costs prevailing when it was run, only about a quarter of the wages attached to technically automatable vision tasks would have been economical to automate at all. Three quarters of what was already technically possible stayed with people, purely on price.
That is the calculation now being redone across every function, with a much lower number on the machine side. And it predicts something specific about which work goes first, which is not the work that is hardest or most valuable. It is work that is:
- high volume, so the per task saving multiplies
- well specified, so the request can be written down without a meeting
- cheap to verify, so a wrong answer is caught before it costs anything
- already text or screen bound, so no physical step blocks it
Compare that list against what actually happened this year. AI was cited in US job cuts covering 205,000 workers through August 2026, concentrated in customer service, data operations, entry level software and finance back offices. Our analysis of that wave sets those categories against our scores, and they are the same four properties in different clothes. The occupations sitting highest in our index, Customer Service Representatives (68) and Data Entry Keyers (63) among them, are high volume, well specified, text bound roles.
What the price does not tell you
Three limits, and they matter as much as the headline.
A benchmark task is not your task. The costs above are for problems on a reasoning benchmark, not for the specific, context heavy, half defined work that fills most jobs. Anything requiring someone to explain the situation first carries a specification cost that never appears in a token bill, and for small jobs that cost exceeds the saving.
The model is not always the right tool at this price. Reporting on Astra's rollout notes it is a poor default for short rewrites, routine classification and simple extraction, where smaller and cheaper models already clear the quality bar. Frontier pricing buys frontier reasoning, which most work does not need.
Availability is not adoption. Astra reached ChatGPT paid tiers, the API, Azure and Bedrock over the days following launch, after a staged release to a limited set of organisations. Reaching a workspace is not the same as being wired into a workflow, and the gap between the two is measured in quarters.
So the honest answer to the question in the title: Astra can probably do several tasks in your week, at a cost per task low enough that the arithmetic works, and it cannot do your job, because a job is a bundle of tasks and only some of them are the shape this technology absorbs. Which is why the useful number is not a benchmark score but the fraction of your own week that looks like the four bullets above. Across the 961 occupations in our index that fraction is what separates the top of the ranking from the bottom.
For the benchmark results themselves see our post on Astra's numbers, and for how it compares with Anthropic's release the same week, Astra against Claude Fable 5.1.
Where does your job sit?
Every number above is an occupation average. Your own exposure depends on your seniority, your sector and how much of your day is routine. Answer a few questions and get your personal score free.
Get my personal scoreFree · No credit card · No signup
Figures come from the AI Job Risk Index and were current when this was published. Scores change as new research is ingested, so the index is always the live version. See how scoring works. Informational guidance based on published research, not professional career or financial advice.