NC

ncrofford.com

Paper

[O-02] The Case for Compiled Execution: Execution Properties in High-Stakes AI Systems

March 2026

Summary

A 3,900-run study demonstrating that inference-based systems fail to satisfy core execution properties (correctness, determinism, temporal fidelity, and governance) while compiled execution achieves perfect reliability by removing inference from the runtime path.

Key findings

  • Compiled execution achieves 100% correctness, 0% temporal misbinding, and perfect determinism across all evaluated scenarios.
  • Inference-based systems exhibit persistent failure modes, including incorrect decisions, temporal misbinding, and output instability, even with structured output and retries.
  • Determinism alone is insufficient: frozen artifacts recover correctness, but lack structural governance capabilities.
  • Execution failures are architectural, arising from the use of probabilistic inference at runtime rather than model capability gaps.

Abstract

As large language models move from conversational interfaces into operational infrastructure, a structural question emerges: can runtime inference serve as the execution substrate for systems that must be correct, deterministic, temporally faithful, and governable? This paper examines that question through a controlled comparison of 3,900 executions across five system architectures, three operational domains, and four execution properties. The central finding is not simply that one system performs better on average. It is that the relevant execution properties do not emerge reliably under runtime inference. Compiled execution (Opus) achieves 100% decision correctness, 0% temporal misbinding, perfect determinism, and zero runtime-induced failures across 780 cells. No inference-based system, including a structured-output stack with validation and retries, multi-role agent orchestration, and agent- generated deterministic artifacts, matches these properties on any axis. The contribution of this paper is therefore architectural as much as empirical. The observed failures are not best understood as bugs awaiting stronger models or more careful prompting. They arise from using probabilistic inference in a role that demands deterministic execution. The results suggest that as AI systems move deeper into governed operational settings, they increasingly benefit from an execution layer with different properties from the reasoning layer above it.