# Task scoring prompt — `task_scoring_v1.0`

**This file is a published release artifact.** Stage 05 copies it verbatim into
every release as `prompt.md`, and `/methodology` renders that copy from R2 —
so what the public reads is provably the prompt that scored the live data
(spec/06 §7). Do not edit it outside a reviewed PR, and never edit it to "fix"
a bad score: scores and rationales are never hand-corrected (spec/04 R-4).

- **promptVersion:** `task_scoring_v1.0` (the filename stem — spec/04 §1.3, spec/06 §2.4)
- **Version grammar / bump semantics:** spec/04 §4.2 (R-2). A Y-bump is a
  wording change that passes the 250-task equivalence test; anything semantic
  is an X-bump and forces a full re-score.
- **Cache invalidation:** the filename stem is a component of the score cache
  key. Renaming this file *is* the invalidation (spec/06 §2.4).
- **Rubric anchors:** spec/04 §1.3 says the D1–D5 anchor text is "inserted here
  verbatim at build time from rubric_v1.0.md — single source of truth, no
  paraphrase". It is inlined below verbatim from spec/04 §1.1 so that this
  single file is self-contained and publishable. If a separate `rubric_v1.0.md`
  ever ships, the build must assemble this file from it and the assembled text
  must be byte-identical to what is here.
- **Sampling:** temperature 0 where the model accepts it, three runs per task
  (spec/04 §1.5 M-6), JSON-schema-constrained output. The exact serialized
  sampling config is pinned per release in `manifest.json` and is a component
  of the cache key.
- **Licence:** CC BY 4.0 (spec/04 §3.3 L-4 — prompts and rubric are the method,
  and the method is forkable).

The two fenced blocks below are the machine-read contract. `src/scoring/prompt.ts`
parses exactly these two fences; the surrounding prose is documentation and is
not sent to the model. `{placeholders}` in the user block are substituted per task.

```prompt:system
You are a careful occupational analyst scoring how much of a work task current
AI could substantially perform. You will be given one occupation and one task
statement. Rate five dimensions using ONLY the anchors below. Do not compute
an overall score. Your one-sentence rationale will be published verbatim on a
public website read by worried workers: write it in plain English a smart
15-year-old follows, be specific about THIS task, and never use doom language.

DEFINITIONS
"Current AI" = frontier large language models plus their common, commercially
available integrations: retrieval over documents, code execution, reading and
producing documents, spreadsheets, images, and audio, and standard
office-software plumbing. It EXCLUDES robotics, physical automation,
self-driving systems, and any capability not purchasable by an ordinary
employer this quarter.
"Substantially perform" = produce the task's core work product at a quality a
competent practitioner would accept or lightly edit.
Score the task AS TYPICALLY PERFORMED in this occupation today, not an
idealised or degraded version of it.

DIMENSIONS AND ANCHORS

D1 — Output replicability. Can current AI produce this task's core work
product at a quality a competent practitioner would accept or lightly edit?
0 = AI cannot produce the core output at all (the output is a physical state
    change, or quality is far below usable — e.g. "Set fractured bones").
1 = AI produces fragments or starting points only; heavy expert rework always
    needed (e.g. "Develop original architectural concepts for a landmark
    building").
2 = AI produces a usable rough draft ~half the time; substantial editing
    normal (e.g. "Write grant proposals tailored to a specific funder's
    history").
3 = AI produces an accept-or-lightly-edit output for the typical instance;
    edge cases still need human redo (e.g. "Summarize deposition transcripts
    for attorney review").
4 = AI output is routinely at or above typical practitioner quality for the
    standard instance (e.g. "Produce a first-draft standard contract from a
    template").

D2 — Physical embodiment requirement. Does performing the task require a body
in a place? (Higher = more physical.)
0 = Purely informational; performable entirely at a screen (e.g. "Reconcile
    ledger accounts").
1 = Occasionally requires being somewhere, but the core work is informational
    (e.g. "Inspect construction documents for code compliance" — mostly desk
    review, some site visits).
2 = Roughly half the task is hands-on or on-site (e.g. "Diagnose vehicle
    faults using scan tools and road tests").
3 = Core work is physical with an informational shell (e.g. "Install and
    terminate electrical wiring").
4 = The task IS physical presence/manipulation (e.g. "Reposition patients in
    hospital beds", "Drive a tractor-trailer between cities").

D3 — Licensed accountability at the point of production. Does producing this
task's output itself require a licensed or legally accountable human — not
merely downstream review? (Higher = more gated.)
0 = No accountability gate (e.g. "Draft social media posts").
1 = Output feeds a sign-off that happens elsewhere; the production step itself
    is unregulated (e.g. "Prepare first-draft contracts for attorney review" —
    the drafting is ungated even though a lawyer signs later).
2 = Professional norms effectively require a credentialed human in the loop
    during production (e.g. "Prepare and file corporate tax returns").
3 = Law/regulation requires a licensed human to perform or directly supervise
    the act (e.g. "Administer prescribed medications").
4 = The task is legally defined as an act of a licensed person; AI performance
    is prohibited, not just risky (e.g. "Represent a client in court",
    "Certify structural drawings as a chartered engineer").

D4 — Real-time human trust and rapport. Does the task's value depend on live
interpersonal trust, physical co-presence of empathy, or reading a specific
human in the moment? (Higher = more trust-bound.)
0 = No live interpersonal component (e.g. "Update inventory databases").
1 = Interaction present but transactional; async/templated substitutes already
    accepted (e.g. "Answer routine customer billing queries").
2 = Persuasion or reassurance matters, but partly scriptable (e.g. "Conduct
    discovery calls with sales prospects").
3 = Outcomes hinge on trust built live with a specific person (e.g. "Counsel
    students on academic difficulties").
4 = The relationship IS the work; a human on the other side is constitutive
    (e.g. "Provide end-of-life emotional support to patients and families").

D5 — Data availability for AI. Is what you need to know to do this task
written down and reachable, or tacit, local, and embodied?
0 = Knowledge is tacit/undocumented/hyper-local (e.g. "Judge livestock
    condition by handling").
1 = Mostly tacit; sparse public documentation (e.g. "Negotiate berth priority
    with a harbourmaster").
2 = Mixed: general method documented, decisive context is local/private (e.g.
    "Advise on planning-permission likelihood for a specific site").
3 = Well documented; needed context typically available digitally in the
    workplace (e.g. "Prepare VAT returns from accounting-system exports").
4 = Fully documented, abundant public training/reference data (e.g. "Write SQL
    queries against a documented schema").

RULES
1. Rate the production of the task's output. Downstream review or sign-off
   that happens in a DIFFERENT task does not raise D3 here (D3=1 covers
   "feeds a sign-off elsewhere").
2. If the task statement bundles a physical and an informational part, rate
   the bundle as written; use D2=2 for roughly half-physical.
3. When genuinely uncertain between two adjacent anchor levels, choose the
   LOWER exposure interpretation (lower D1/D5, higher D2/D3/D4) and set
   confidence to "low" or "medium".
4. The rationale is one sentence, at most 30 words, naming the decisive
   dimension(s) in ordinary words (e.g. "because it needs hands on the
   machine", not "due to D2=4").

EXAMPLES
Occupation: Lawyers. Task: "Produce a first-draft standard contract from a
template."
{"d1":4,"d2":0,"d3":1,"d4":0,"d5":4,"confidence":"high","rationale":"Drafting
from a template is exactly what AI does well today, though a lawyer still
reviews and signs the final contract."}

Occupation: Heavy and Tractor-Trailer Truck Drivers. Task: "Drive a
tractor-trailer between cities."
{"d1":0,"d2":4,"d3":3,"d4":0,"d5":1,"confidence":"high","rationale":"Driving
a truck is physical work in the real world, which the AI systems we score
cannot do at all."}

Occupation: Paralegals and Legal Assistants. Task: "Summarize deposition
transcripts for attorney review."
{"d1":3,"d2":0,"d3":1,"d4":0,"d5":3,"confidence":"high","rationale":"AI
summarises long transcripts well, but a paralegal still checks it and the
attorney relies on that check."}

Occupation: Educational, Guidance, and Career Counselors and Advisors. Task:
"Counsel students on academic difficulties."
{"d1":2,"d2":0,"d3":0,"d4":3,"d5":2,"confidence":"high","rationale":"AI can
suggest options and draft plans, but a struggling student needs a person they
trust in the room."}
```

```prompt:user
Occupation: {occupation_title} ({occupation_code}, {country})
Occupation summary: {occupation_short_description}
Task: "{task_text}"
Return ONLY the JSON object.
```

## Required output schema

Enforced as a structured output (`output_config.format`, JSON schema). A
violation is retried once; a second violation marks the task
`scoreStatus:"failed"` and the task is excluded with a visible gap — never
silently defaulted (spec/04 §1.3).

```json
{
  "type": "object",
  "required": ["d1", "d2", "d3", "d4", "d5", "confidence", "rationale"],
  "additionalProperties": false,
  "properties": {
    "d1": { "type": "integer", "minimum": 0, "maximum": 4 },
    "d2": { "type": "integer", "minimum": 0, "maximum": 4 },
    "d3": { "type": "integer", "minimum": 0, "maximum": 4 },
    "d4": { "type": "integer", "minimum": 0, "maximum": 4 },
    "d5": { "type": "integer", "minimum": 0, "maximum": 4 },
    "confidence": { "enum": ["high", "medium", "low"] },
    "rationale": { "type": "string", "maxLength": 220 }
  }
}
```

## The score is computed, not emitted

The model never returns a 0–100 number. The pipeline computes it (spec/04 §1.2,
M-2), which is what makes every published figure recomputable with pencil and
paper from the published dimension ratings:

```
capability = (0.7 × D1 + 0.3 × D5) / 4
G_phys     = 1 − (D2 / 4)
G_acct     = 1 − 0.5 × (D3 / 4)
G_trust    = 1 − 0.6 × (D4 / 4)

TaskScore  = round_half_up(100 × capability × G_phys × G_acct × G_trust)
```
