NC

ncrofford.com

Case Study 03

Vendor Insurance Compliance

A pre-registered study whose headline hypothesis failed. What it found instead: two frontier models, built by different vendors, buy near-perfect precision by declining ~40% of the consequential work.

Read this first: the registered hypothesis failed

This study pre-registered its primary measure before any arm was run: the silence-as-denial rate: of the requirements where a certificate says nothing at all, how many does the system wrongly report as a deficiency?

Silence-as-denialRate
Parity0 / 25
Claude Opus 51 / 25
GPT-5.60 / 25

No difference. All three arms handle silence correctly, and the hypothesis this study was designed around is not supported by its own data.

Everything below is secondary. It comes from a measure that was registered as a refutation criterion but not as the headline, and it is not being promoted to headline after the fact. A reader should treat the primary result of this study as null.

What this is

A vendor sends a certificate of insurance. Someone has to check it against what the contract actually required: every coverage line, every limit, every endorsement, and produce a register of what is missing. It is high-frequency, consequential in both directions, and largely manual.

Wrong in one direction blocks a vendor from working over a deficiency that is not real. Wrong in the other leaves an uninsured party on site.

Both halves of every package are real documents. 50 insurance-requirements articles from 44 public agencies across 8 agency types, and 14 certificates transcribed from public municipal records. Only the pairing between them is constructed: reconciling a given article against a given certificate is exactly the operation a reviewer performs, but we cannot and do not claim that any particular vendor satisfied any particular contract.

The secondary finding

AccuracyFlag precisionElusionDeferred a real deficiency
Parity (on Haiku 4.5)73.8%100%3.8%11.5%
Claude Opus 537.7%97.6%25.4%36.9%
GPT-5.639.2%100%20.8%40.0%

Both frontier models achieve near-perfect flag precision by declining to flag. They defer on real deficiencies at 36.9% and 40.0%: two different vendors, two different architectures, the same operating point.

A deficiency register that stays silent on roughly 40% of real gaps has returned the work to the reviewer. Being right about what it says is not the same as having done the job.

Why the measurement standard earned its keep here

GPT-5.6 scored 100% flag precision: identical to Parity. An evaluation that reported flag precision alone would have called this a tie, and it would have been wrong.

The disclosure schedule requires abstention and elusion to be read together, on the explicit grounds that a single operating point is gameable: flag everything and elusion falls while precision collapses; flag nothing and precision holds while elusion explodes. Neither can be optimised without surrendering the other.

This study's own registered criterion required both to be matched simultaneously. On elusion the arms separate by a factor of five, and the operational difference becomes visible.

An earlier version of this protocol named flag precision alone as decisive. Under that criterion, this study would have reported no difference. That is the strongest argument we have for the standard, and it comes from a study where our own hypothesis failed.

The rubric is not ours

A certificate of insurance is evidence of intent to cover. It is not proof that coverage exists, a point four appellate courts state in near-identical language.

Corter-Longwell v. Juliano (NY App Div, 2021): "it is well established that a certificate of insurance, by itself, does not confer insurance coverage"

Landsman Development v. RLI (NY App Div, 2017): "not conclusive proof, standing alone, that such a contract exists"

Shala v. Park Regis (NY App Div, 2021): "merely evidence of a contract rather than conclusive proof that coverage was procured"

Molinar v. 21st Century (Cal Ct App, 2024): "it does not create additional rights or obligations beyond the policy"

The certificate form prints the same rule on its own face: "A statement on this certificate does not confer rights to the certificate holder."

So a system reporting "requirement satisfied" from a certificate alone is making a claim courts have repeatedly declined to make. This is the one element of the study that is external to us: the corpus is ours, but what a certificate is permitted to prove is not.

Where we lost

Both models beat Parity on the rows requiring judgment. GPT-5.6 read 88% of interpretive rows correctly against Parity's 52%, and both scored 100% on true negatives against Parity's 80%. Parity over-parks: it defers on rows the models resolve correctly.

Parity's own accuracy is not good. 73.8% end-to-end; 85.4% on the verdict alone, so roughly twelve points of the gap is the wording of the reason rather than a wrong call.

Eleven of Parity's errors are its own extraction, not its reconciliation logic. It failed to read a limit that is present on the document. The protocol pre-registered that this class of error be attributed to extraction quality rather than to the compiled path, and it is.

GPT-5.6 was both cheaper and faster than Opus 5. Any cost claim has to name the model.

Limitations

The registered primary measure returned null. Stated first, above.

The failure taxonomy has not converged. 28 of 30 failure families found in the last six certificates were new. This corpus demonstrates failure modes; it does not cover them, and no coverage claim is made.

Pairings are constructed. No claim about any actual vendor's actual compliance.

Gold is single-read, by one reader, with no independent double-read, therefore no inter-rater agreement figure.

Form-version diversity is effectively absent. Excluding one certificate on privacy grounds removed the only example of the older form revision.

5 of 20 certificate pages were unusable, degraded precisely on the fields this study measures. Excluded rather than guessed.

No court in the harvested set rules on what a description-of-operations entry proves, and that is where endorsement claims actually appear. The rubric leans on it.

One document population: municipal agenda packets. Findings may not transfer to broker-issued certificates.

Parity's runtime cost was not captured, so cost per accepted-correct item is reported for the model arms only and is incomplete for the study as a whole.

Protocol deviation, disclosed: the GPT-5.6 arm was added after registration. Adding an arm can only make the comparison harder, but it is a deviation and is recorded as one.

The one-sentence version

On 130 requirement rows drawn from real contracts and real certificates, this study's registered hypothesis failed, and the measure that did separate the arms showed two frontier models, built by different vendors, buying near-perfect precision by declining roughly 40% of the consequential work.