Data study · 9 August 2026 · 7 min read
Two different things are called AI exposure
Some studies measure what language models can do with a job. Others measure how much of it machines could take over. Across 961 scored occupations the two barely correlate, yet both are published under one label.
Every number in this field ships under the same label. A study says your occupation is 72 percent exposed to AI, a chart ranks accountants above nurses, a headline rounds it all to "AI is coming for these jobs". The label is always AI exposure.
Behind the label sit two different questions.
The first: how much of this job is reading, writing and reasoning that a language model can do? The second: how much of this job could machines take over at all, physical work included? A machinist scores low on the first and high on the second. Both answers are right. They are answers to different questions.
We maintain an index built from published research on AI and work, and reconciling those sources forced us to confront this directly. This post shows, with measurements you can check, that the research base splits into two families along exactly this line, that the families barely agree with each other, and what that means for reading any exposure number, including ours.
As with every data post here, the figures below are generated from the live corpus rather than written by hand, so they move as research is ingested.
The two questions
Sources in the first family measure the overlap between a job's tasks and what language models do. GPTs are GPTs, the Science paper this whole literature leans on, asks which tasks an LLM could do in half the time. Felten, Raj and Seamans map language modelling capability onto occupational abilities. The ILO's global index rates task level exposure to generative AI. Microsoft's applicability study reads a large sample of real assistant conversations to see which work activities people actually bring to a model.
Sources in the second family measure how much of a job machines could absorb, with no restriction to language. The Global Automation Atlas rates task substitution across the whole task set, physical work included. WORKBank has experts rate how far current AI systems, not just models, could automate each task.
Fifteen of the sources behind the index ask the first question. Five ask the second. The first family is where most of the corpus sits, which will matter later.
Sources that ask the same question agree
If the two families were roughly interchangeable ways of measuring one thing, sources from different families would agree about as well as sources from the same family. They do not come close.
We measure agreement as rank correlation across the published occupations each pair of sources shares: 1.00 means two sources order occupations identically, 0.00 means knowing one tells you nothing about the other. Measured on the live corpus, sources asking the same question agree at 0.48 to 0.84, while sources asking different questions agree at 0.01 to 0.15.
| Source pair | Ask the same question? | Rank agreement | Shared occupations |
|---|---|---|---|
| GPTs are GPTs (OpenAI and Penn) × Language modeling AIOE (Felten) | Yes | 0.84 | 912 |
| GPTs are GPTs (OpenAI and Penn) × Working with AI (Microsoft) | Yes | 0.75 | 893 |
| Global Automation Atlas × WORKBank | Yes | 0.48 | 104 |
| Global Automation Atlas × Language modeling AIOE (Felten) | No | 0.01 | 912 |
| Global Automation Atlas × GPTs are GPTs (OpenAI and Penn) | No | 0.15 | 923 |
| WORKBank × GPTs are GPTs (OpenAI and Penn) | No | 0.11 | 104 |
The contrast is the finding. GPTs are GPTs and Felten's index were built by different teams with different methods, and they rank occupations almost identically, because they ask the same question. The Global Automation Atlas is a careful piece of work, and its agreement with either of them is indistinguishable from noise, because it asks a different one.
This is not one family being right and the other being wrong. Try both questions on a plumber. A language model can handle quoting, scheduling and compliance paperwork, a modest slice of the job, so the language family scores plumbers low. Could machines take over more of the physical work in principle? A different and larger number. Neither answer is a mistake. Publishing either one bare, labelled AI exposure, is where the trouble starts.
One occupation, both answers
Here is what the split looks like on a single role: Political Scientists (54).
| Source | Question it asks | Estimate |
|---|---|---|
| Working with AI (Microsoft) | What can a language model do? | 74 |
| Language modeling AIOE (Felten) | What can a language model do? | 73 |
| Anthropic Economic Index | What can a language model do? | 68 |
| Refined global index (ILO) | What can a language model do? | 55 |
| GPTs are GPTs (OpenAI and Penn) | What can a language model do? | 54 |
| Global Automation Atlas | What could machines take over? | 8 |
| Published score | 54 |
Five sources that ask the language question, built independently by two AI labs, an academic team, a UN agency and Microsoft, land in a band you could summarise honestly. The one source asking the machine question sits far below them, and it is not wrong: political science is talk, text and judgment, and very little of it is the kind of work you would hand to a machine wholesale.
A reader who saw only the last row would call political scientists moderately exposed. A reader who saw only the automation estimate would call them nearly safe. Same occupation, same scale, same label on both numbers.
Why we do not average the two
The obvious fix is to treat the families as two measurements and give each half the weight. We built exactly that in August 2026, measured it against the corpus, and reverted it the same week.
The problem is coverage. Right now the median published occupation is covered by four language model sources and one automation source. Weighting the family means equally therefore hands one automation source half of nearly every score it appears in. When we measured the effect, 89 percent of published scores moved and 104 occupations changed risk band, most of that on the strength of a single preprint outvoting four or five independent sources that agreed with each other. Political scientists dropped 18 points and a whole band.
Five corroborating estimates should not lose to one because of a labelling scheme. So the published score remains one weighted consensus over all sources, and that has a consequence you should know when you read this site: 931 of the 961 published occupations carry at least one source from each family, and because the language family contributes most of the evidence, our published score mostly reflects the language question. Automation sources still count toward the consensus and appear on every occupation page. They inform the number. They do not dominate it.
That is a choice, not a law of nature, and our data page states it alongside the other limitations. A site with the opposite weighting would publish different numbers and would not be lying. It would be answering the other question.
How to read any exposure number
Ask which question it answered. A single exposure figure, in a study, an article or a ranking, is an answer to one of these two questions, and the label will not tell you which. The method section will. Task overlap with language models is the first question. Automation potential across all tasks is the second.
Cross family comparisons are the misleading ones. A chart that puts a language exposure score for writers next to an automation potential score for welders is comparing answers to different questions on one axis. Within a family, comparisons mostly hold. Across families they are close to meaningless, which the correlations above quantify.
Disagreement between families is information. A role that scores high on the language question and low on the machine question is a role where models change the workflow but a person stays in the loop. High on both is the profile that deserves the word risk. Our earlier post on what the research agrees on covers how much of the apparent chaos in this literature is really this one split, wearing a dozen different method sections.
Check the sources behind any number of ours. Every occupation page lists the studies behind its score, so you can see which questions your own number is built from.
Where does your job sit?
Every number above is an occupation average. Your own exposure depends on your seniority, your sector and how much of your day is routine. Answer a few questions and get your personal score free.
Get my personal scoreFree · No credit card · No signup
Figures come from the AI Job Risk Index and were current when this was published. Scores change as new research is ingested, so the index is always the live version. See how scoring works. Informational guidance based on published research, not professional career or financial advice.