NC

ncrofford.com

Standard

What an Automated Work System Should Disclose

Eight numbers, all of which already have names, none of which any framework requires anyone to publish.

Why this exists

A buyer evaluating a system that processes documents or records is told an accuracy figure. That figure almost never comes with a denominator, a sample size, an interval, or any statement of what the system declined to attempt.

Without those, it does not predict the thing the buyer is accountable for. A system reporting 97% accuracy has not said which three items in a hundred are wrong, so a reviewer still opens all one hundred. The accuracy number moved; the labor did not.

This document is not a new measurement method. The quantities below are long-established and are named here by their established names. What is missing from the market is not the mathematics. It is the obligation to publish it.

These quantities already have names

The framework for systems that can decline to answer is selective prediction, sometimes called classification with a reject option. It dates to Chow's error-reject tradeoff (1957, 1970) and was formalised by El-Yaniv and Wiener (2010), who define coverage (the fraction of items the system attempts) and selective risk (the error rate on that fraction). The pair is normally reported as a risk-coverage curve rather than a single point.

Measuring what a system missed on the items it did not flag has a name in electronic discovery: elusion, defined in the Grossman-Cormack Glossary (2013) as "the fraction of Documents identified as Non-Relevant by a search or review effort that are in fact Relevant," estimated by drawing a random sample from the null set. Detecting a system's own errors is a studied task with a standard evaluation convention (Hendrycks and Gimpel, 2017).

Whether a cited source actually supports the claim it is attached to is attribution (Rashkin et al. 2021; Gao et al. 2023). Same-input-same-output is repeatability under ISO 5725-1:1994. Applying the rule in force at the time of an item is as-of or bitemporal correctness (Snodgrass 1999; SQL:2011). Bounding a rate from zero observed events is the rule of three (Hanley and Lippman-Hand, 1983).

Reasoning about whether the reference standard is itself trustworthy is the imperfect reference standard problem, handled systematically in diagnostic-accuracy work by QUADAS-2 (2011) and STARD (2015).

Nothing below is new. Where this document has a contribution it is stated plainly in the next two sections.

The gap

No framework currently in force requires a system to publish the abstention rate, or to report accuracy conditional on non-abstention.

EU AI Act (Regulation 2024/1689) Article 15(3) requires that "the levels of accuracy and the relevant accuracy metrics of high-risk AI systems shall be declared in the accompanying instructions of use." It is silent on the denominator, the sample size, the confidence interval, and whether the system declined to answer.

NIST AI RMF 1.0 (2023) asks that accuracy measurements be paired with representative test sets and documented methodology. It is voluntary and prescribes no format.

ISO/IEC 42001:2023 is a management-system standard: the wrapper, not the measurement.

The legislature has said as much itself. Article 15(2) directs the Commission to encourage "the development of benchmarks and measurement methodologies" in cooperation with metrology and benchmarking authorities. That is a statute conceding the instrument does not yet exist.

So the contribution here is narrow and specific: treat the selective-prediction quantities as a disclosure obligation, and connect one of them to cost. The observation that review labor is a function of error containment rather than of abstention rate appears in no evaluation framework found, and it is where a calibration property meets a labor line. Everything else is packaging.

The disclosure schedule

Eight things a vendor can either state or cannot. Each names its denominator.

01

Coverage

The fraction of submitted items the system produced final output for with no human involvement. State what was excluded before submission (eligibility filters, unsupported formats, out-of-scope classes), because items filtered out are in nobody's denominator.

02

Selective accuracy

Of the covered fraction, the share that is correct. Not "precision," which in information retrieval and in the Grossman-Cormack Glossary means positive predictive value.

03

Abstention rate

The fraction deferred to a human, with a reason attached to each. Coverage and abstention sum to one.

04

Elusion

Of the items not deferred, the share that was wrong, estimated by sampling the unflagged set, not inferred from the flagged one. This is the number nobody publishes and the only failure that costs you after you have decided to trust the system.

05

Flag precision

Of the items that were deferred, the share that was actually wrong. Without this, abstention cannot be distinguished from noise.

06

Attribution rate

The share of sampled output elements whose cited source was verified to support the claim. Verified, not merely present: a precise citation to the wrong page resolves and is still wrong.

07

Repeatability

The share of items producing identical output across repeated runs, with the equivalence rule stated and caching disabled.

08

Cost per accepted-correct item, including review labor

Whose reviewers, at what loaded rate, over what volume and period, and whether one-time integration is amortised.

Alongside any of these: the unit of analysis (an item is a page, a claim, a line, a file, and it changes every rate by an order of magnitude), the sample size and interval method, and whether correctness was judged against ground truth or against the system being replaced.

How to read the answers

A single operating point is gameable; ask for the curve.

Flag everything and elusion goes to zero. Flag nothing and abstention looks excellent. Items 1, 3, 4 and 5 are one point on a tradeoff, which is why the literature reports risk-coverage curves. The Grossman-Cormack Glossary's own counterexample: in a million documents of which 10,000 are relevant, a review returning 1,000 documents none of which are relevant shows 1.001% elusion, "belying the failure of the search."

Abstention without flag precision is not a finding.

A system that defers 30% with perfect elusion, where only 2% of deferred items were actually wrong, has moved work rather than removed it, particularly if the deferred items route back to the buyer's own staff while cost per item is quoted on the vendor's.

Recall measured against the incumbent is a substitution test.

If a system is scored on the population its predecessor already surfaced, the result establishes that it is a safe replacement for something already failing. It says nothing about what neither system ever surfaced. Say so when that is what was done.

A zero is a bound, not a fact.

Zero events in n trials places the rate below roughly 3/n at 95% one-sided confidence: under 3% at n=100, under 0.3% at n=1,000. Report the bound.

Where trained people disagree, ask what "correct" was.

Inter-rater agreement in specialist classification work is often only moderate: Stausberg et al. (2008) found coding specialists at κ = 0.42, corresponding to 46.8% raw agreement. Note that κ is chance-corrected and is not an agreement rate; the two differ by tens of points. The right question is not whether the system beat κ, but how the reference standard was constructed, by whom, and whether the adjudicators were blind to which system produced the output.

And ask what was excluded before submission.

This is the most common way a good scorecard describes a narrow system. A monitoring system at a large US bank produced excellent numbers on the alerts it generated; when the bank sampled transactions the system had deliberately not flagged, roughly half warranted a regulatory filing, against about 4% of the alerts it did flag. The bank discontinued the sampling (FinCEN, U.S. Bank National Association, 15 February 2018).

Four questions that end an evaluation

01

What is your elusion rate, and how did you sample the unflagged set?

02

Of the items you flag, what share were actually wrong?

03

What is your cost per accepted-correct item, including whose reviewers?

04

Show me the source for these outputs, and let me check three of them.

Disclosure

This document is written by someone who builds systems in this category. That is a conflict and it is stated rather than managed.

It also has a shape. Several of these items are structurally cheaper for some architectures than others. A system that executes deterministically finds item 7 nearly free, and a system built with an abstain mechanism finds items 3, 4 and 5 far easier to produce than one without. That tilt is real, and a reader who assumed neutrality would be wrong.

The defence is not neutrality. It is that each item is derived from what a buyer needs the work product to prove, and that each is a quantity the field already defined for its own reasons long before this document. The items are offered severally, not as a package. A buyer who does not need repeatable output should ignore item 7 and say so.

Sources

Chow, IRE Trans. Electronic Computers EC-6(4):247–254 (1957); IEEE Trans. Inf. Theory 16(1):41–46 (1970) · El-Yaniv & Wiener, JMLR 11:1605–1641 (2010) · Geifman & El-Yaniv, NeurIPS 2017 (arXiv:1705.08500) · Hendrycks & Gimpel, ICLR 2017 (arXiv:1610.02136) · Grossman & Cormack, The Grossman-Cormack Glossary of Technology-Assisted Review, 7 Fed. Cts. L. Rev. 1 (2013) · Grossman & Cormack, 17 Rich. J.L. & Tech. 11 (2011) · Hedin, Tomlinson, Baron & Oard, Overview of the TREC 2009 Legal Track, NIST · Rashkin et al. (arXiv:2112.12870) · Gao et al., EMNLP 2023 (arXiv:2305.14627) · ISO 5725-1:1994 · Snodgrass, Developing Time-Oriented Database Applications in SQL (1999); SQL:2011, SIGMOD Record 41(3):34–43 (2012) · Hanley & Lippman-Hand, JAMA 249(13):1743–1745 (1983); Eypasch et al., BMJ 311:619 (1995) · Whiting et al., QUADAS-2, Ann. Intern. Med. 155(8):529–536 (2011) · Bossuyt et al., STARD 2015, BMJ 351:h5527 · Cohen, Educ. Psychol. Meas. 20(1):37–46 (1960) · Stausberg et al., Int. J. Med. Inform. 77(1):50–57 (2008) · NIST AI 100-1 (2023), DOI 10.6028/NIST.AI.100-1 · NIST AI 600-1 (2024) · ISO/IEC 42001:2023 · Regulation (EU) 2024/1689 · Raghu et al. (arXiv:1903.12220) · Mitchell et al., Model Cards (arXiv:1810.03993) · FinCEN, Assessment of Civil Money Penalty, U.S. Bank National Association (15 February 2018).

Not yet checked

ISO/IEC 25059:2023 and ISO/IEC TS 25058:2024 are the only ISO artifacts in this measurement lane and are paywalled. If either already specifies an abstention-rate disclosure, the gap claimed above narrows and this document should say so.