Paper
March 2026
A 3,900-run study demonstrating that inference-based systems fail to satisfy core execution properties (correctness, determinism, temporal fidelity, and governance) while compiled execution achieves perfect reliability by removing inference from the runtime path.
As large language models move from conversational interfaces into operational infrastructure, a structural question emerges: can runtime inference serve as the execution substrate for systems that must be correct, deterministic, temporally faithful, and governable? This paper examines that question through a controlled comparison of 3,900 executions across five system architectures, three operational domains, and four execution properties. The central finding is not simply that one system performs better on average. It is that the relevant execution properties do not emerge reliably under runtime inference. Compiled execution (Opus) achieves 100% decision correctness, 0% temporal misbinding, perfect determinism, and zero runtime-induced failures across 780 cells. No inference-based system, including a structured-output stack with validation and retries, multi-role agent orchestration, and agent- generated deterministic artifacts, matches these properties on any axis. The contribution of this paper is therefore architectural as much as empirical. The observed failures are not best understood as bugs awaiting stronger models or more careful prompting. They arise from using probabilistic inference in a role that demands deterministic execution. The results suggest that as AI systems move deeper into governed operational settings, they increasingly benefit from an execution layer with different properties from the reasoning layer above it.