Transparency
The data, and what it can and cannot tell you
Every score on this site comes from the same pipeline: take published research on AI and work, pull out the occupation level estimates, put them on one scale, and combine them. This page explains how that works and, more importantly, where it is weak.
We would rather you understood the limits than trusted the numbers more than they deserve. If something here is unclear or looks wrong, email willaitakemyjob@proton.me. Corrections are genuinely welcome.
What is in the dataset
961
Occupations published, each with 2 or more sources
20
Distinct publications the estimates come from
1,133
Occupations scored in total, including those held back
31
Mean exposure score across the index, 0 to 100
33 pts
Median gap between the lowest and highest estimate for the same occupation
568
Occupations where that gap is 30 points or more
How a score is built
- A published paper or report is ingested, and its occupation level estimates are extracted automatically by a language model, which also records how precise the source was.
- Each estimate is mapped onto a 0 to 100 exposure scale and matched to a canonical occupation from the O*NET classification.
- Estimates for the same occupation are combined as a weighted mean. More recent estimates and more precisely stated ones count for more.
- Occupations with fewer than two sources are excluded from everything public.
The full walkthrough, including the personal modifiers used in the assessment, is on the methodology page.
Where this is weak
Four limitations that matter more than anything else on this page. None of them are secret and none of them are fixed yet.
The sources do not all measure the same thing
A task overlap score, a wage weighted exposure estimate, a government employment projection and a survey of employer intentions are four different quantities. We place them on one scale so they can be combined. That is a modelling choice, not something the underlying research endorses, and it is the single biggest reason to treat a score as approximate. It also means the 33 point median gap above is an upper bound on disagreement rather than a clean measurement of it.
The deepest version of this split: the corpus contains two families of measurement, sources that ask how much of a job a language model can do and sources that ask how much of it machines could take over at all. Within a family, sources rank occupations almost identically. Across the families they barely agree, and we have measured how little. Because the language family contributes most of the evidence, the published score mostly reflects that question. Automation sources inform the consensus without dominating it, and we deliberately do not weight the two families equally, because coverage is so uneven that doing so would let a single source outvote several corroborating ones.
Coverage is concentrated
A small number of large datasets carry most of the index, and just over half of the published occupations rest on exactly two sources. For many of those the two are the same pair of broad automated mappings. Depth of corroboration varies enormously between an occupation with twelve independent estimates and one with two, which is why the source count appears on every occupation page. Read a two source score as a rough placeholder.
Extraction is automated, not hand coded
Estimates are pulled out of source documents by a language model rather than transcribed by a researcher, and the extraction records a confidence that is deliberately capped well below certainty. This adds error on top of whatever uncertainty the source already had. We have found and corrected wrong matches by hand, which is evidence the process makes mistakes, not evidence that it no longer does.
Nothing here is validated against outcomes
No part of this has been checked against what actually happened to employment in any occupation. It is an aggregation of what the literature says, not a tested predictor. There is no sensitivity analysis on the weighting scheme either, so we cannot tell you how much the numbers would move under different reasonable choices.
What it supports, and what it does not
Reasonable to use it for
- A rough sense of where an occupation sits relative to others
- The broad pattern that screen and text heavy work is more exposed than physical work, which is consistent across sources and the most robust thing here
- Finding the underlying studies, since each occupation page lists them
Not reasonable to use it for
- Predicting how many jobs will exist in an occupation, or when
- Reading a score as a probability, a percentage of tasks automated, or a timeline
- Claims about a specific employer, region or person
- Fine grained comparisons: the difference between a 41 and a 45 is well inside the noise
The most defensible way to read any score here is as a rough band, not a precise number.
The sources
20 distinct publications sit behind the index, ranging from peer reviewed academic work to government statistics and consultancy analysis. By number of occupations covered, the largest contributors are:
The concentration visible in this table is the limitation described above. Every occupation page lists its own sources.
What the numbers show
With all of the above in mind, the pattern in the aggregate is worth stating. Of 961 ranked occupations, 0 occupations clear the high risk threshold of 70. 281 sit in the transitioning band and 680 in the lower risk band. The mean is 31. The highest scoring occupation is Customer Service Representatives at 68.
The aggregate literature is markedly less alarming than the coverage of it. That is the finding we would stand behind, because it holds regardless of the weighting choices and does not depend on any single source. Read more in the research notes or browse the full index.
Using these figures
Quoting them is fine, with attribution and a link. Two requests, both about accuracy. Date the figures, because scores change as sources are added and the live index is always the current version. And describe a score as an aggregate estimate of an occupation's exposure to current AI capability, not as a probability that a job disappears.
If you want a cut of the data we do not publish, or want to check a specific occupation's sources before relying on it, email willaitakemyjob@proton.me.