Case Study 02
The same architecture on a different kind of work, and a different result. Two frontier models emitted 110 diagnosis codes that do not exist. The compiled path emitted zero.
250 clinical notes. 2,842 physician-assigned diagnosis codes. Three systems, identical inputs.
| Coverage | Precision | Invalid codes emitted | Elusion | |
|---|---|---|---|---|
| Parity (on Haiku 4.5) | 16.9% | 65.2% | 0 of 804 | 22.0–28.7% |
| Claude Opus 5 | 26.6% | 59.3% | 20 of 1,361 | 29.2–37.9% |
| GPT-5.1 | 26.3% | 49.0% | 90 of 1,768 | 22.5–33.7% |
The frontier models code more of the chart. They also invent codes that do not exist. Parity emitted zero across 804 assignments, not because it was careful, but because it cannot: it binds each diagnosis to the CMS FY2026 tabular list, and anything it cannot bind is parked rather than guessed.
A code that does not exist is rejected by the payer and returns as a denial. A code that exists but is wrong is worse. It is paid, and it is a compliance exposure.
Scope
This is diagnosis coding only. No CPT, no E/M, no sequencing. That is why these numbers cannot be set against a vendor's "automation rate," which is blended across all of them.
Medical coding is revenue-cycle work: a coder reads a clinical note and assigns the diagnosis codes that justify payment. Rates vary by roughly an order of magnitude between a routine outpatient visit and a complex inpatient chart, and which of those this corpus resembles is settled below, and it is not the cheap one.
This is the second artifact scored against the eight-item disclosure schedule. As with the first, where we lose, we say so, and here we lose on the measure buyers ask about first.
CodiEsp (CLEF eHealth 2020): 250 held-out clinical notes, 2,842 diagnosis codes assigned by professional coders. Public, CC BY 4.0, no data-use agreement. Train and dev splits were never scored, so nothing here has been iterated against.
These are not outpatient office visits, and that governs every comparison. CodiEsp is published clinical case reports: 11.4 diagnosis codes per note (median 10, max 35), with 53% carrying ten or more and only 12% carrying the one-to-four a routine visit does. The chapter mix is neoplasm, circulatory, digestive and genitourinary: complex specialist cases with extensive comorbidity.
The input is machine-translated. CodiEsp is Spanish; the ICD-10-CM index is English. Translation artifacts are scored as our extraction errors.
9.2% of the gold is not valid in FY2026: CodiEsp uses the Spanish CIE-10-ES edition. Exact match is capped at 90.8% before anyone starts.
The gold is 10–20% noisy and was never adjudicated.
45% of what we "missed" was never in the note. See the elusion section. It applies to every arm equally.
| # | Measure | Parity | Denominator |
|---|---|---|---|
| 1 | Coverage | 16.9% | 2,842 gold codes |
| 2 | Selective accuracy (raw) | 65.2% | 804 emitted |
| 3 | Abstention | 54.5% | gold codes |
| 4 | Elusion | 22.0–28.7% | gold codes |
| 5 | Flag precision (proxy) | 0.62 | 2,272 parked rows |
| 6 | Attribution | 100.00% | 804 emitted |
| 7 | Repeatability | see below | — |
| 8 | Cost per accepted-correct code | $0.23–0.57 | accepted-correct |
| — | Code validity | 100.00% | 804 emitted |
On attribution: 100% is structural, not impressive. Every emitted code carries a verbatim quote that appears in the submitted note, because extraction drops any diagnosis it cannot ground before a code is assigned. Opus 5 reached 99.71% and GPT-5.1 98.47%.
Flag precision is a proxy and is labelled as one. It reports the share of parked rows corresponding to a gold code. A true flag precision asks whether auto-coding would have been wrong, which needs a counterfactual we do not have.
Same 250 notes, same rubric, abstention offered to every arm and stated as unpenalised.
| Parity | Opus 5 | GPT-5.1 | |
|---|---|---|---|
| Coverage | 16.9% | 26.6% | 26.3% |
| Precision (raw) | 65.2% | 59.3% | 49.0% |
| Abstention | 54.5% | 35.5% | 40.0% |
| Elusion | 22.0–28.7% | 29.2–37.9% | 22.5–33.7% |
| Code validity | 100.00% | 98.53% | 94.91% |
| Attribution | 100.00% | 99.71% | 98.47% |
| Latency / note | 13.4s | 54.6s | 11.2s |
Parity wins precision, elusion, validity and attribution. It loses coverage, by a lot. It surfaces less of the chart and is right more often about what it surfaces, a different result from Case Study 01, where the frontier models won completeness and lost containment. We are not going to force the two into one story.
Parity ran on Haiku 4.5, Anthropic's small model, against two frontier models, at a quarter of Opus 5's latency.
Parity's 2,272 parked rows carry a closed vocabulary: `see_also_fork` (748), `no_main_term` (540), `hint_unplaced` (284), `uncertain` (265), `neoplasm_site_unmapped` (199). Each names the missing dimension, so they aggregate, route, and become a physician query.
Opus 5's 1,628 parked rows carry free-text reasons, each one unique: "Unclear whether current use, dependence, or former use; nicotine dependence vs tobacco use status not documented." Better prose per item. Unusable across 1,628 of them.
That is the honest contrast, not "they give no reasons."
GPT-5.1 was run twice over the same 250 notes.
| GPT-5.1, run 1 vs run 2 | |
|---|---|
| Notes with an identical code set | 9.6% (24 of 250) |
| Mean code-set overlap | 52% |
| Codes differing per note | 3.8 |
Nearly four codes change per note between identical runs. For work that is billed and audited, a system whose output depends on which afternoon it ran is a different kind of risk from one that is merely imprecise, and it is invisible in any single-run benchmark.
Worth noting across the two studies
The same model was byte-identical on Case Study 01 and 9.6% repeatable here. Same model, opposite property, different work. A buyer cannot infer a system property from the model's identity. It has to be measured on the actual workflow.
Elusion is what nobody publishes: gold codes neither coded nor flagged. Of 997 on Parity:
| Share | What it is | |
|---|---|---|
| Concept present in the note | 55% (550) | a real recall gap, ours to fix |
| Concept absent from the note | 45% (447) | the gold names what the note does not contain |
The absent set is `R69 Illness, unspecified`, `F10.20 Alcohol dependence`, `E79.0 Hyperuricemia`: chart-context codes a coder assigns from the whole record, not from this note. No system reading only this note can recover them, and the same 45% is buried inside every arm's figure. Parity's attributable elusion is roughly 12–16%.
A term-overlap heuristic across 997 codes. Directionally solid; not to be quoted to two significant figures.
Parity's charge is $0.48 to $1.20 per note depending on daily volume, at 13.4 seconds each.
Why there is no vendor price beside that
Published rates of roughly $0.90 offshore and $3.50 domestic describe a one-to-four-code outpatient visit. This corpus averages 11.4 codes per note. Setting our cost against that rate would flatter us by comparing different work, the same error as setting an ICD-only result against a blended automation rate.
Should an automation vendor charge for the work it completes, or for the work it ingests?
We charge per note ingested. At 16.9% coverage the buyer pays full price for a chart and still codes most of it, and every parked row is something they pay us for and do themselves.
Charging for completed work returns value at essentially any coverage. Whatever is auto-coded is cheaper and faster than a human doing it; whatever is not costs the buyer nothing. The vendor carries a direct incentive to raise coverage, because unfinished work is unpaid work.
Charging for ingested work (our model) has a genuine break-even. Below some coverage, the fee plus the residual human time exceeds the human doing all of it. Coverage becomes the buyer's risk rather than the vendor's.
This is the same argument being had one layer down about model tokens: whether a customer should pay for tokens a system spent producing its own errors. A vendor billing on ingestion is insulated from its own miss rate. We are describing our own pricing model here, not somebody else's, and a buyer should ask it of any vendor including us.
We have not resolved it. The answer depends on a quantity this study did not measure: how long a coder takes on a parked row carrying an extracted diagnosis and a typed missing dimension, versus coding it from scratch. Review-minutes-per-parked-item is the most valuable measurement absent from this study, and no economic claim here should be believed without it.
Two items are partial and say so. Flag precision is a proxy. Repeatability is reported for GPT-5.1, where two independent runs exist; Parity's is not claimed here, and the structural argument that its coding step is deterministic is an argument, not a measurement.
Elusion required a judgment, and Case Study 01's did not. Parked rows carry free text, not codes, so linking one to a gold code cannot be done by string equality. CodiEsp ships 3,665 evidence spans that would make this objective, but they index the Spanish source and we ran the English translation. A blind judge did the linkage; cross-checked against a different model family on a seeded 60-decision sample it agreed 76.7%, with disagreements symmetric, so there is no bias toward crediting our own parks. Human coders agree on ICD-10-CM at κ=0.4857.
A precision measure we built was invalid and is withdrawn. The protocol defined a "forensic" precision crediting us where the gold was wrong. It counts a disagreement as a gold defect whenever the emitted code is valid and grounded, both of which hold by construction for a spine-bound system. It scored 524 matched, 279 gold-defect, exactly 1 wrong. It is not a measurement.
Diagnosis codes only. No ICD-10-PCS, CPT, E/M levelling, sequencing or present-on-admission. On an outpatient chart, E/M levelling is where most coding error is found.
One workflow, one corpus, one day. No platform-level claim is warranted. Cross-domain claims remain tentative, including in a direction that would refute them.
Every number is reproducible. The corpus is public and downloadable, the spine is versioned, the harness and all three arms' raw outputs are committed, and the judge's decisions are cached by content hash so scoring replays without re-judging.
Narrower audience
Addressed to the labs whose models are measured above. Everything is an inference from the numbers already published.
Case Study 01 found a small model inside a compiled architecture matching two frontier models on extraction while beating both on abstention and repeatability. This study finds something different, and the difference is the interesting part.
Here the frontier models are better at the task. They code more of the chart. What they cannot do is stay inside the codebook. GPT-5.1 emitted 90 codes that do not exist in FY2026; Opus 5 emitted 20; the compiled path emitted zero, structurally.
That is not a quality difference a stronger model closes. It is the difference between generating a code and navigating to one.
A model asked to output an ICD-10-CM code is being asked to recall a 74,719-entry codebook exactly. The same model asked to extract a documented condition, with the codebook consulted deterministically, cannot produce a code that does not exist.
The commercial reading: the work that is gated is not gated on model capability. It is gated on never emitting an unpayable code, on flagging what documentation cannot support, and on saying which items were missed. Supply those and regulated back-office work becomes addressable: work not currently done by a model at all, at any capability tier.
The honest limits. One workflow, one day, a machine-translated corpus, two partially-measured items, and coverage well below a human coder. The claim we will stand behind is narrow: on identical inputs, the compiled path emitted zero invalid codes where two frontier models emitted 110 between them, and was right more often about what it did emit.
On 250 clinical notes, a small model inside a compiled architecture coded less of the chart than two frontier models, was more accurate about what it coded, missed less, emitted zero codes that do not exist against their 110, and typed every one of its 2,272 abstentions with the documentation dimension that was missing.