NC

ncrofford.com

Case Study 01

Medical Chronology

A small model inside a compiled architecture, measured against two frontier models on the same 254 clinical encounters.

The result in twenty seconds

One 154-page medical record. The same document, three ways to produce the chronology.

TimeCostCompletenessKnows what it missed?
Human review vendor5–7 days~$200undisclosedundisclosed
Claude Opus 5309s$2.36100%no: 2.2% of flags warranted
GPT-5.1——could not process the recordno: zero flags, ever
Parity (on Haiku 4.5)146s$2.2599.2%yes: 100% of flags warranted

The frontier models are excellent at extraction and cannot tell you what they missed. That is the whole finding. A system that is 99.5% accurate but cannot isolate its 0.5% still gets every item reviewed; a system that is 99.2% accurate and reliably isolates the uncertain 0.8% does not. Those two have completely different labor economics.

On the cost figures

Cost is platform charge only and excludes reviewer labor. See measure 8 and the limitations. The ~$200 human figure is a published market price for the same deliverable and does include the labor, so the comparison favors us by construction until review minutes are measured.

What this is

A medical chronology is litigation work product: every dated clinical encounter in a medical record, attributed to its provider, ordered, with treatment gaps surfaced. A law firm or insurer buys it today from an offshore review vendor, an onshore nurse reviewer, or one of a dozen AI products.

This is the first artifact scored against the eight-item disclosure schedule published alongside it. Every number is measured on production code under a protocol fixed in advance. Where we lose, we say so.

The corpus

254 dated clinical encounters across 8 matters. 236 held-out, never seen during any hardening pass.

Clinical narrative is real and de-identified, verbatim from the published MTSamples corpus. The matter frame (provider identities, patient banner, dates of service, assembly order, Bates numbering, scan degradation) is synthetic and authored by us, which is why the answer key is known exactly rather than adjudicated. Elusion is therefore computed, not estimated from a sample.

MatterEncountersPagesImage-onlyStructural feature
0118544same-day multi-document; tuned, reported separately
02203910interleaved concurrent treatment
0314221no gap: negative control
0418214same-day multi-document encounters
05182079 providers, two near-identical name pairs
0618212no gap: second negative control
hard1825537 planted distractor dates, three formats
large13015426the record length a firm actually submits

Unit of analysis: the dated clinical encounter, one per source document. Two documents on one date are two encounters. A chronology reporting 18 dates where there are 18 documents on 16 dates has miscounted, and we score that as an error.

Results: the eight items

#MeasureResultDenominator
1Coverage99.2%252 of 254 encounters
2Selective accuracy100%auto-produced rows
2bRow precisionsee over-emission belowrows emitted
3Abstention: true encounters deferred0.0% (0 of 254)all encounters
3bPark volume: the reviewer's burden15 items, all on the distractor matteritems sent to review
4Elusion: silent misses0.79% (2 of 254); zero on 7 of 8 mattersall encounters
5Flag precision100%parked rows
6Attribution97.8%273 rows verified against source
7Repeatability1.000: byte-identical3 runs × 3 matters
8Cost per accepted-correct row$0.0174accepted-correct rows

Elusion is the number that matters and it is the one nobody publishes. Both misses fall on the 154-page record; the other seven matters are clean. This is an observed rate, not a bound. We are not claiming zero.

The denominator is all 254 true encounters, not the un-deferred subset. Electronic discovery measures elusion within the discard pile, which is a smaller denominator and an easier number, since defer more and it falls without fixing anything. We report against everything submitted, because the question a buyer is asking is "did anything vanish," not "how bad is the pile you discarded."

Flag precision at 100% is what makes abstention meaningful. On the distractor matter the system parked 15 dates: transcription stamps, records-release dates, a 2019 prior-treatment reference, and every one genuinely required judgment. Abstention that flags noise moves work rather than removing it; this does not.

The risk-coverage curve

A single operating point is gameable, so we swept the park threshold across its usable range. Elusion stays at 0.000 at every operating point of the sweep, while review burden falls as the threshold rises. The shipped setting is not perched on a knife-edge.

These are two different quantities

This 0.000 and the 0.79% above are not the same measurement. The sweep varies the park threshold, which acts on lines already extracted, with extraction held fixed. Both pooled misses were never extracted at all. They are extraction-stage failures, and no threshold setting could have recovered them. So the sweep shows that tuning review burden introduces no new silent misses; it says nothing about the two that occurred upstream of it. Read as "no misses at any threshold," the 0.000 would be wrong. The honest headline elusion is 0.79%.

What we compared against

Same corpus, same rubric, same scoring, including the abstention columns, which the model arms were explicitly offered in their task specification and told plainly they would not be penalised for using.

ParityClaude Opus 5GPT-5.1GPT-OSS-120BHuman review
Coverage99.2%100%100%62.7%undisclosed
Elusion0.79%0%0%37.3%undisclosed
Attribution97.8%100%100%78.4%undisclosed
Flag precision100%2.2%zero flags0%undisclosed
Repeatability1.0000.55–0.601.000—undisclosed
154-page record146s / $2.25309s / $2.36failed—5–7 days / ~$200

Where we lose, stated plainly

Opus 5 beat us on completeness. 130 of 130 encounters on the large record, where we missed one. GPT-5.1 also reached 100% coverage on every matter it could process. Our own prior published prediction, that frontier models fail under execution constraints, did not hold on extraction, and we report that rather than omit it.

GPT-5.1 is as repeatable as we are. Byte-identical across three runs. Any claim that compiled execution is uniquely reproducible is refuted by that column. What the column does show is that repeatability is model-dependent and unmanageable from outside: Opus 5 agreed with itself on 55–60% of rows between identical runs, and it rejects the temperature parameter, so it cannot be pinned even deliberately. A buyer cannot tell in advance which behavior they are getting.

We emit more rows than there are encounters: 1.33× to 2.22×, and 1.85× on the 154-page record against Opus 5's 1.02×. Every row is correct and correctly cited; the deliverable is repetitive, not wrong. The cause is understood: a multi-page document produces several extraction windows, and separating "two readings of one document" from "two real documents on one date at one provider" requires document segmentation we have built but not yet shipped, because a detector below spec would trade a visible duplicate for an invisible miss.

Where the models could not go

GPT-5.1 could not process the 154-page record at all: 41.5 MB of inline page images against a request-size ceiling. The workarounds (Files API, downsampling, chunking into roughly ten independent calls) all reintroduce cross-batch incoherence, since each batch decides on its own what counts as an encounter with no shared dedup.

GPT-OSS-120B lost 37.3% of encounters. It accepts no images, so image-only pages are unreadable to it. That is a capability limit, not a judgment failure, and we score it as such.

What nobody else publishes

We checked every AI chronology product we could find: Supio, EvenUp, Eve, Precedent, Parrot, Filevine, Alexi, Trellis.

Zero publish an accuracy, completeness, recall, or missed-item rate. Seven of eight publish no price.

The only numeric accuracy claim in the entire category is a human vendor's "99.8%", with no methodology, no denominator, and no definition of accuracy, on a site that states its own turnaround three contradictory ways across three pages.

And every AI vendor's verification claim is a user-interface affordance, not a measured rate: "verify in a click" · "Verify every fact, citation, and exhibit in one click with line-level citations" · "you can always check your facts and cite your sources". Each describes how quickly you can check the work. None discloses how often checking finds an error. That is review burden transferred to the buyer with the miss rate undisclosed, precisely the gap the disclosure schedule exists to close.

Declared limitations

Stated up front, because the schedule's whole point is that gaps get named rather than omitted.

We used a materially less capable model, and we mean capable, not cheap. Parity's extraction runs on Claude Haiku 4.5 (Anthropic's small model), measured against Opus 5 and GPT-5.1, both frontier. Anthropic would argue Opus 5 is the more capable model, and on general reasoning they would be right; that is what the tier means.

The finding is that it did not decide the delivered work. Haiku inside a compiled, evidenced architecture landed level with both frontier models on coverage and attribution, and ahead of both on the two properties a buyer of this work product actually needs: abstention that means something, and reproducibility that holds. Those two are not model capabilities that a stronger model would supply; Opus 5 is the stronger model and produced neither. Model capability was not the binding constraint on this work. Architecture was.

Abstention needs two numbers, and one of them is unflattering. Zero true encounters were deferred, but the system did send 15 items to review, every one a distractor date on the hard matter. A single rate hides this: put parked distractors over a denominator of true encounters and it reads 1.778, which is nonsense. 15 items against 18 encounters on that matter is a real review burden, and reporting only the 100% flag precision would describe it as a triumph while hiding the cost.

One workflow, measured over roughly one day. Repeatability across three runs in an afternoon says nothing about week- or month-scale variance: provider-side model updates, silent behavioral drift, which is where determinism actually earns its keep. No platform-level claim is warranted from a single workflow, and this one is our most inference-heavy, so it understates the economics.

The reference standard is author-constructed. No inter-rater agreement, no κ. We built the test and the answer key. A pilot on a buyer's own records, scored by their own reviewers, is strictly stronger evidence than this document.

Component notes were selected to contain no full dates, so assigned dates of service are the only encounter-shaped dates the narrative contributes. That makes date extraction easier than reality. The hard matter exists to counter it: 37 planted competing dates in three formats, and results are reported with and without that advantage.

Blind scoring was not performed. Scoring is deterministic against an exact key, so labeler bias is not the live risk, but the protocol called for it and it was not done.

Cross-domain claims are pending. Everything here describes one workflow. Whether the pattern generalises is an open question these results cannot settle. Three further studies are planned against this same schedule; any positioning that depends on the pattern holding across domains is tentative until those complete, including in a direction that would refute it.

What this implies for model vendors

Narrower audience

This section is addressed to the labs whose models are measured above. Everything in it is an inference from the numbers already published; where it depends on something we cannot observe, it says so.

Two constraints define a frontier model business right now: gross margin is thin relative to software, and compute is the binding supply limit on how much work can be served at all. Almost every lever available attacks one and worsens the other. Cheaper models raise margin and lower quality. Bigger models raise quality and consume the scarce input.

The architecture measured here moves both in the same direction, and the reason is structural rather than clever.

1. Same delivered work, materially less compute. The two paths produce the same chronology from the same 154 pages. One runs frontier-model inference over every page. The other runs a small model on the extraction leg and a compiled register with no model call at execution at all. Token volumes are broadly comparable. Our windowed extraction actually emits more output tokens than Opus 5 does, so the difference is not that we send less. It is which model serves it: a tier priced 5× lower on input and output. For a vendor whose ceiling is capacity rather than demand, that ratio is the business.

2. Recurring inference becomes one-time inference. A compiled workflow calls the model once, at authoring, to turn a description into a specification. Every execution after that is deterministic replay. Run it once or ten thousand times, the compute is the same. Token-metered revenue falls per execution; revenue per unit of compute rises, because the same fleet now serves a far larger volume of executed work.

3. The addressable surface expands rather than shifts. Work that cannot tolerate an answer that varies between runs, cannot be cited back to a source, and cannot say what it missed is not currently being done by a model at all. It is being done by people, slowly, at ~$200 and 5–7 days per record. Opus 5 flagged nothing useful and disagreed with itself 40% of the time between identical runs; GPT-5.1 emitted zero flags across every matter despite being offered the mechanism and told plainly it would not be penalised for using it. Those are not quality failures. The extraction was excellent. They are the reason the work stays manual. Supplying repeatability, attribution, and honest abstention converts a non-customer into a customer. That is net-new demand for inference, not a reallocation of existing demand.

The honest limits of this argument. We have not measured any vendor's cost of serving, and no vendor publishes one, so nothing here is an observation about anyone else's economics. Compute consumption is inferred from model tier and token counts, not from FLOPs we can see. And this is one workflow, our most inference-heavy, over roughly one day. The claim we will stand behind is narrow and measured: at the same delivered work and nearly the same price, the compiled path used a model tier priced 5× lower, ran in half the wall-clock, and produced the two properties the frontier models did not.

Method, in brief

Documents are dropped through the production front door, the same path a customer uses, never an engine called directly. Extraction is windowed, governed, and confidence-gated; anything unresolved is parked with a reason code and a citation rather than dropped. The compiled register that assembles the chronology runs deterministically with no model call at execution.

Scoring matches each true encounter to output rows by exact date and contained provider name, consumes rows so duplicates cannot inflate a match, and classifies every true encounter as auto-produced, parked, or eluded.

Every number above is reproducible. The corpus builder, the scoring harness, and the raw arm outputs behind each figure are committed. The corpus regenerates deterministically from a seed: real MTSamples clinical narrative, public and requiring no data-use agreement, and re-scoring the published outputs reproduces the table exactly. The matter PDFs themselves are deliberately not published: they contain the assembled records, and we retain nothing we do not have to.

We mention this because the alternative is the norm. A benchmark whose numbers cannot be recomputed by the reader is a claim, not a measurement, and the schedule this study is scored against exists to make that distinction enforceable.

The one-sentence version

On this workflow, a small model inside a compiled, evidenced architecture matched two frontier models on extraction, beat both on the two things a buyer of this work product needs (abstention that means something and results that repeat) ran a 154-page record in 146 seconds for $2.25 against 5–7 days and ~$200 of human review, and is the only participant in the category that will tell you what it missed.