Evidence
Measurement work on systems that process documents and records.
A buyer evaluating one of these systems is normally given an accuracy figure and no denominator: no sample size, no interval, and no statement of what the system declined to attempt. The schedule below sets out what should be published instead.
None of the eight quantities is new. Each has an established name in the literature, some of them for decades. What is missing from the market is not the mathematics but the obligation to publish it.
Standard
Eight numbers, all of which already have names, none of which any framework requires anyone to publish.
Case Study
A small model inside a compiled architecture, measured against two frontier models on the same 254 clinical encounters.
Case Study
The same architecture on a different kind of work, and a different result. Two frontier models emitted 110 diagnosis codes that do not exist. The compiled path emitted zero.
Case Study
A pre-registered study whose headline hypothesis failed. What it found instead: two frontier models, built by different vendors, buy near-perfect precision by declining ~40% of the consequential work.