NC

ncrofford.com

Paper

[O-01] When Language Models Fail as Execution Systems: A Study of Exactness, Conformance, and Determinism

March 2026

Summary

A 1,260-run study showing that language models fail as execution systems under exact-match constraints, while deterministic systems remain stable.

Key findings

  • No evaluated language model exceeded a 25% exact-match pass rate under execution-oriented constraints.
  • A large share of failures were execution fidelity failures: correct substance, invalid structured artifact.
  • Identical inputs produced non-identical outputs across repeated runs, revealing non-determinism at the execution boundary.

Abstract

Large language models are increasingly proposed not merely as interfaces for generating text, but as systems capable of participating directly in operational workflows. In these settings, the relevant question is not whether a model can produce a plausible answer, but whether it can reliably produce the exact required output, under fixed constraints, in a form that downstream systems can safely consume. This paper evaluates that question through a controlled study of 1,260 live executions across five frontier language models and one deterministic execution system, using twelve policy scenarios derived from four operational workflows: invoice discrepancy, refund policy, conditional credit, and insurance adjudication. Across all evaluated language models, exact-match reliability remains low, with no system exceeding a 25% pass rate. The dominant failure mode is execution non-conformance: models often produce outputs that are semantically reasonable or partially correct while still failing to emit valid structured artifacts. In addition, repeated executions on identical inputs yield non-identical outputs, indicating systematic non-determinism at the system boundary.