How we built this
How we actually did this.
This page explains how the numbers on this site were made, in ordinary words. There is no maths to follow and nothing technical to understand. Just what we did, in the order we did it, including the parts that went wrong.
We used AI to do the work. That is the first thing most people want to know, and the rest of this page is about what we did to make sure that was not simply a machine guessing.
01
The question
Everybody is asking the same thing, in one form or another: is a machine going to do my job?
The answers on offer are almost all useless. You have seen the headlines: your job is 60% at risk, this profession is doomed, that one is safe. Nobody who publishes a number like that ever explains what it means. Sixty per cent of what? Risk by when? Decided by whom, using what?
It is not just vague. It is the wrong shape of answer. Nobody in the world does a job that is 60% gone and 40% left. What actually happens is that some of the things you do on a Tuesday get easier or get taken away, and other things you do do not change at all. Which is which is the only part that is any use to you.
So we tried to build the version of the answer we wanted for ourselves: specific, checkable, and honest about what it does not know.
02
The idea
A job is not one thing. It is a list of things.
A bookkeeper does not do “bookkeeping”. They match supplier statements against payments, chase unpaid invoices, prepare a tax return, and sit with a worried client who cannot see how they will make payroll this month. Those are four completely different jobs wearing one job title, and only one of them is really a spreadsheet.
So we do not score job titles. We score the individual things people do at work, one at a time, and then add them back up. That way a person can look at their own job and see which parts are moving and which parts are not, instead of being handed a single number about their whole life.
We do not score job titles. We score the individual things people do at work, one at a time.
03
Arguing with ourselves first
Before we scored anything, we wrote down the plan: what we would measure, where the facts would come from, what we would refuse to claim.
Then we attacked it. The plan was split into nine parts, and each part was written by one AI and then handed to a second AI whose only job was to tear it apart. The critic did not know who had written what it was reading. Each part was also judged against six real pieces of published work by other researchers: the ones we would have to be better than for this to be worth doing at all. A part was not finished until it beat them.
That process is why several things on this site are worded so cautiously. The critics kept catching claims we could not actually support, and the rule was that an unsupported claim gets cut rather than softened.
04
Where the facts come from
We did not sit around deciding what we thought people do at work. Two governments have already done that, carefully, by asking them.
In the United States, the Department of Labor keeps a public list of the tasks that make up each occupation, built from surveys of the people who actually do those jobs. In the United Kingdom, the Office for National Statistics defines the official list of occupations, and a published research project maps the tasks inside them. Between them, that covers 830 American occupations and all 412 British job groups.
Pay and employment numbers come from the same official sources: the government statistics agencies in each country, not our estimates and not a jobs board. Every source is named, dated, and linked on the method page, along with the licence that allows us to use it.
One consequence worth knowing: official task lists describe jobs as they were last surveyed, not as they are this month. Newer work, such as checking what an AI produced, mostly is not in them yet.
05
The five questions
Then we asked AI (Claude, made by Anthropic) the same five plain questions about every single task on those lists. The same five questions, in the same words, every time.
1Can a computer produce this?
Could today’s AI produce the actual thing this task produces, at a standard a competent professional would accept or lightly tidy up? Not one day. Today.
2Does it need a body in a room?
Some work cannot be done down a wire. Fitting a boiler, cutting hair, restraining a frightened animal, being on the ward.
3Does the law require a licensed person?
Some things are legally an act of a qualified human being: prescribing a medicine, signing off electrical work, representing someone in court. That is a wall, not a difficulty.
4Does it need live human trust?
Some work only functions because a real person is on the other side of it, in the moment: telling a family bad news, talking someone through a decision they are frightened of.
5Is the knowledge written down anywhere?
AI is good at things that have been documented and bad at things that live in someone’s head. Judging a crop by the feel of the soil is not in any book.
That is 36,664 separate task statements, each one read on its own, each one answered against the same five questions.
The AI never picks the score
This is the part that matters most, so here it is on its own.
The AI does not decide that a task is a 72. It cannot. All it does is answer the five questions above, each on a scale of nothing to a lot. A fixed sum turns those five answers into the number you see, and that sum was written down before any scoring started, published in full, and never changed for one job and not another.
So anyone who thinks a score is wrong does not have to argue with us about it. They can look at the five answers, look at the sum, and check the arithmetic on the back of an envelope. And if they think an answer is wrong, we publish the exact wording we sent, so they can ask the same question themselves and see where they land.
The AI never picks the final number. It answers five questions, and a fixed sum does the rest.
One rule inside that sum is worth spelling out, because it does a lot of work: if a task genuinely requires a human body in a particular place (laying a floor, lifting a patient, driving a lorry), the score goes to zero, no matter how well the paperwork around it could be written. We are measuring what software can do. We are not measuring robots.
The AI also wrote one short sentence about each task, saying which of the five things decided it. Those sentences are published too, next to the tasks, so you can see the reasoning rather than take the number on trust.
06
How we checked ourselves
A number is only worth as much as the effort spent trying to break it. We ran four checks. One of them found a real mistake in our own work. We have published all four, including the boring one.
Check one: does it agree with itself?
Twenty-five tasks were handed to every batch of scoring, so the same twenty-five sentences were scored over and over by different scorers who could not see each other’s answers, 135 times each in the end. That gives 226,125 head-to-head comparisons of two independent opinions about the same task.
They landed in the same place most of the time. Two independent scorers gave a task the same verdict about 85 times in every 100, and where they differed they were typically about four points apart out of a hundred. That is not perfect agreement, and we have never claimed it. It is roughly the level of agreement you would want from two experienced people marking the same paper.
Check two: we wrote down what we expected, before we looked
Before any of the scoring existed, we picked three jobs and wrote down the range we expected each one to land in. Paralegals, lorry drivers, nurses. The point of committing in advance is that you cannot quietly move the target afterwards.
Two came out inside their range. Paralegals did not: we had said 50 to 70, and the score came out at 46.
So we stopped and investigated, and the answer was uncomfortable. Our expectation had been wrong, not the score. When we wrote the prediction we had been imagining a list of paralegal tasks that included two things paralegals in the official data are not recorded as doing at all: managing lawyers’ calendars, and pulling case records. Both were near the top of what we had assumed. Worse, when we did the arithmetic properly, the top of the range we had committed to was not reachable at all. Half of a paralegal’s recorded work is not drafting: it is handling physical files, filing things at court, and sitting with clients and witnesses. Even if every single drafting and research task had scored the maximum the guide allows, the job could not have reached 60.
We replaced the expectation with one derived properly, we left the failed one on the record with the reasoning, and we published the whole thing. A prediction you are allowed to quietly delete is not a prediction.
Check three: does our British scoring actually know British law?
A real risk with this kind of work is that American assumptions get copied onto British jobs. Who is allowed to prescribe a medicine, sign off a gas installation, give financial advice or represent someone in court is not the same in the two countries, and getting it wrong would quietly distort hundreds of British scores.
So we tested it. We took 30 British tasks covering prescribing, legal work, regulated financial advice, gas and electrical work, drivers’ hours, social work and teaching, and scored them again, this time spelling out explicitly that these were British jobs under British rules. If the original scores had been quietly running on American assumptions, the answers should have moved.
They did not move meaningfully. The difference was no bigger than the ordinary disagreement between two scorers looking at the same task. We then read every one of the 21,084 published British explanations and counted the language in them: British legal and regulatory terms turned up about seventeen times more often in British rows than American ones, and there was not one genuine case of an American legal concept being applied to a British job.
So this check found nothing, and nothing was re-scored as a result. We are publishing it anyway. A test you only mention when it agrees with you is not a test.
Check four: the mistake we found
This one was real, and it was ours.
Many tasks appear word-for-word in more than one job. “Arrange appointments for clients and staff” sits in dozens of official job descriptions. To save repeating work, we scored each sentence once, and it got scored under whichever job happened to reach it first.
That is fine when the two jobs are similar. It is not fine when they are not. The example that made us stop was the task “interview patients to obtain their medical histories”, which had been scored while the scorer was looking at a description of engineering professionals. Whether a task is legally restricted to a licensed person often depends entirely on whose job it is: a farmer may lawfully vaccinate their own animals, a vet is doing something quite different when they do the same thing.
So we fixed it. We re-scored 5,968 job-and-task combinations from scratch, with the correct job named in front of the scorer. The important part: we did not let the scorer see the number we had already published. Then we compared.
The effect was real but smaller than we had feared. The published score changed for 194 of the 412 British job groups; sixteen of those moved far enough to change the label on the page; the largest single move was under eight points out of a hundred.
And then the genuinely awkward finding, which we could easily have kept quiet. When we looked closely at the movement, most of it was not the mistake at all. It was the ordinary disagreement you get any time a fresh scorer looks at the same sentence, the same disagreement measured in check one. The part actually caused by naming the right job was about half a step on one of the five questions. Announcing the full movement as “the size of our error” would have overstated it by roughly three and a half times: a more dramatic story, and a false one. So we published the smaller, duller, correct number, and the working behind it.
07
What we cannot tell you
Everything above is a judgement about what today’s AI can do. It is not a prediction of what employers will do. Those are completely different questions, and only one of them is answered on this site.
- It says nothing about whether your particular employer will change anything, or when. Most employers move far more slowly than the technology does.
- It says nothing about whether customers will accept it, whether your industry’s rules allow it, or whether it is worth the cost of changing how a place works.
- It cannot tell you that you are going to lose your job. Nobody can, and anyone selling you that certainty is selling you something.
- The scores are considered judgements, not measurements from an instrument. Two reasonable people could set the questions up slightly differently and get slightly different answers. That is exactly why we publish the questions.
- There is an obvious awkwardness in using an AI to assess what AI can do. We cannot remove it. We can only be transparent enough that someone else can check the result, and compare it against work built in completely different ways, which we will publish as it lands.
When AI takes over parts of a job, the usual result so far has not been that the job disappears. It is that the job changes shape. That is a much less exciting sentence than the headlines, and it is the one the evidence supports.
08
Why we made it
Collab365 is a small British company. For over a decade we have run training and communities for people learning new working skills, and lately that has mostly meant learning to work with AI. We built this because the people in those communities kept asking what was actually happening to their jobs, and we could not find an answer we trusted enough to repeat.
So, plainly: we sell something. Collab365 sells memberships to Collab365 Spaces: guided communities where people work through real problems with AI alongside others doing the same job. Some pages on this site link to one when it genuinely matches the work someone is looking at. That is the entire business model.
Here is what that does and does not mean.
- The scores, the reasoning behind every task, the questions we asked, the exact wording we sent and the raw data are all free, complete, and reusable by anyone, including people who want to use them to disagree with us. Nothing here sits behind an email address.
- No score was nudged. The scoring saw a task and five questions; it never saw anything about what we sell, and nobody looked at a result and asked whether it was good for business. The numbers existed before anyone thought about what they meant commercially.
- We are not asking you to take that on trust, because you should not have to. We publish the questions and name the exact AI that answered them, which means anyone can run the same thing and catch us if we put a thumb on the scale. That is a much stronger guarantee than a promise.
- The pages carrying the actual numbers (the method and the data) sell you nothing at all. No offer, no sign-up, nothing beside a figure.
We are not claiming to be neutral. We are claiming to be checkable, which is the more useful thing to be. If you find something wrong, tell us at hello@collab365.com. Corrections get published, dated, whether they came from us or from you.
