Paper
March 2026
An 850-run study showing that no existing system architecture (LLMs, structured systems, or agents) produces reliable execution, and that external governance, applied as a wrapper, enforces constraints but renders systems non-functional.
As AI systems move into autonomous execution, a critical question emerges: are observed failures a property of models or of the systems that host them? This paper evaluates 850 executions across five system architectures, two models, and three workflow complexity tiers. The central finding is architectural: no system type achieves reliable execution, with only 3.4% of runs passing and 0% success beyond simple workflows. More importantly, external governance—implemented as a policy enforcement layer outside the system—creates a structural paradox. Ungoverned systems are unsafe, freely violating constraints, while governed systems are safe but non-functional, with blocked actions preventing task completion entirely. The results demonstrate that the governance gap cannot be closed through better models, prompting strategies, or agent frameworks. It is a missing infrastructure layer. Reliable autonomous execution requires governance to be integrated into the system itself as a runtime control plane, not applied externally as a wrapper.