Futureproof

Open data

Take the whole thing.

Every score on this site, every task rationale behind it, the prompt that produced them and the exact upstream files they were built from: downloadable, in full, at a permanent address, with no signup and no email address.

The site reads the same files you are downloading. There is no private, better dataset.

Licence

Data, scores, rationales and the prompt: CC BY 4.0. Use it commercially, redistribute it, build on it. Just keep the attribution.

Permanence

Releases are immutable. A release you cite in March is byte-identical in November, at the same URL.

Verification

Every file carries its size and sha256 from the release manifest, so a mirror can be checked without trusting us or this domain.

Releases

Listed from the published data store.

  1. 2026-q4.1

    Current

    Published 2026-08-05T16:59:58.953Z · prompt task_scoring_v1.0 · model claude-opus-5 · method 2.0.0

    methodVersion 2.0.0: task scores for statements shared between several occupations are keyed to the occupation as well as the words, and every score row now names the occupation it was scored under.

    US occupations
    830
    UK unit groups
    412
    Tasks scored
    53,309
    From cache
    34,020
    Known gaps in this release (4)
    • 471 (task, occupation) placement(s) from the catalogue have no score yet and are omitted from this release rather than published with a zero. They are concentrated in 12 partially scored occupation(s), each of which states the shortfall in its own `gaps` array, and the exposure figures on those pages are computed only from the statements that are scored.
    • 47,341 of 53,309 score rows are still keyed on the task text alone. Where that text is shared between occupations the number was produced under one occupation's context and shown for all of them; ledger/OCCUPATION-CONTEXT-PROBE.md measures that at a mean 8.03 points and a 30% band-flip rate on the licensing dimension.
    • 4,711 score row(s) have no recorded originating occupation (scoredUnderOccupationId: null). Those pre-date the provenance field and could not be recovered from the batch ledger.
    • scores.json is keyed on (taskId, occupationId). Every task id in this release belongs to exactly one occupation (the id carries the occupation), so each of its 53,309 rows is that occupation's own published number and no page renders a row belonging to a different job. What is NOT closed is the scoring context: 12,353 (occupation, task) pair(s) across 1,021 occupation(s) publish a rating that was produced under a DIFFERENT occupation's context, because the task statement's wording is shared between jobs and has only been scored once so far. Each one names the occupation it was measured under in scoredUnderOccupationId, no page presents it as measured for the job on it, and they close as those (text, occupation) pairs are re-scored.
    Files in release 2026-q4.1, with size and sha256 for verification
    FileSizesha256
    US occupations (CSV)230.7 KBe2837e0717c0a284f5a6d3e5ed5c75a83214ef17e5ceb8d9329d9322324e71ff
    UK occupations (CSV)129.4 KB4dd14e0309ac79e2f16e9136526d304f4dbbad15afc8decb5f339d4b26b54591
    Task scores (CSV)21.4 MBbb163e3d834ddc77f01efe94112b2b0e607d3df07ca2eea10b62369137e22e50
    Occupations (JSON)4.0 MB63e229efd0733a6ace7f5caaea69fb6af0aa8b24644afcf57b946b136859d5ff

    6 more files: see the full list with hashes.

    Everything in release 2026-q4.1

  2. Published 2026-08-05T16:00:02.007Z · prompt task_scoring_v1.0 · model claude-opus-5 · method 2.0.0

    methodVersion 2.0.0: task scores for statements shared between several occupations are keyed to the occupation as well as the words, and every score row now names the occupation it was scored under.

    US occupations
    830
    UK unit groups
    412
    Tasks scored
    53,309
    From cache
    34,020
    Known gaps in this release (4)
    • 471 (task, occupation) placement(s) from the catalogue have no score yet and are omitted from this release rather than published with a zero. They are concentrated in 12 partially scored occupation(s), each of which states the shortfall in its own `gaps` array, and the exposure figures on those pages are computed only from the statements that are scored.
    • 47,341 of 53,309 score rows are still keyed on the task text alone. Where that text is shared between occupations the number was produced under one occupation's context and shown for all of them; ledger/OCCUPATION-CONTEXT-PROBE.md measures that at a mean 8.03 points and a 30% band-flip rate on the licensing dimension.
    • 4,711 score row(s) have no recorded originating occupation (scoredUnderOccupationId: null). Those pre-date the provenance field and could not be recovered from the batch ledger.
    • scores.json is keyed on (taskId, occupationId). Every task id in this release belongs to exactly one occupation (the id carries the occupation), so each of its 53,309 rows is that occupation's own published number and no page renders a row belonging to a different job. What is NOT closed is the scoring context: 12,353 (occupation, task) pair(s) across 1,021 occupation(s) publish a rating that was produced under a DIFFERENT occupation's context, because the task statement's wording is shared between jobs and has only been scored once so far. Each one names the occupation it was measured under in scoredUnderOccupationId, no page presents it as measured for the job on it, and they close as those (text, occupation) pairs are re-scored.
    Files in release 2026-q4, with size and sha256 for verification
    FileSizesha256
    US occupations (CSV)230.7 KBfac86cef5258dfee392d0d4b0d94150b807fc45dc07ebdb56c643e14c45b8018
    UK occupations (CSV)129.4 KBcaee16777e2b311f84919607fd99f21f58f298c273b388b298740f14f2c998f1
    Task scores (CSV)21.4 MBca6b4b45e194552c635d30ccc45250fa0788e42a8cab9ad8354f2a0f6220c464
    Occupations (JSON)4.0 MB52e9a162c78ccd1c045dad1dcb76509876cb0a15ead41c096af8dcb71010579d

    6 more files: see the full list with hashes.

    Everything in release 2026-q4

The /data/latest alias

/data/latest/<file> always serves the current release’s copy of that file, and the response carries an X-Release-Id header naming which release you actually got. It is a convenience for scripts, not a citable identifier: cite the dated URL, because that one cannot change under you.

The alias can never resolve to a retracted release. Retraction requires moving the pointer first, so “latest” and “retracted” are mutually exclusive by construction.

Cite us

Cite the release, not the site. Release 2026-q4.1 will still say exactly what it says today, at the same address, for as long as this project exists.

Plain text

Collab365 (2026). Collab365 Futureproof: task-level AI exposure for US and UK occupations, release 2026-q4.1 (methodVersion 2.0.0, promptVersion task_scoring_v1.0). https://futureproof.collab365.com/data/2026-q4.1. Licensed CC BY 4.0. Built with O*NET data (USDOL/ETA, CC BY 4.0); ONS data (Open Government Licence v3.0); GAISI task framework (arXiv:2507.22748, MIT); BLS data (public domain).

BibTeX

@misc{collab365futureproof2026q41,
  title        = {Collab365 Futureproof: task-level AI exposure for US and UK occupations, release 2026-q4.1},
  author       = {{Collab365}},
  year         = {2026},
  url          = {https://futureproof.collab365.com/data/2026-q4.1},
  note         = {Release 2026-q4.1, methodVersion 2.0.0, promptVersion task_scoring_v1.0, CC BY 4.0}
}

Found something wrong?

Then we would like to know, and we will publish the correction with a date on it. The fastest route is to email hello@collab365.com with the release id, the occupation or task id, and what you think is wrong. Corrections appear in the changelog on the method page; a release found to be wrong in itself is retracted rather than quietly edited.