Research
A growing archive of empirical work on reliable AI systems.
March 2026
An 850-run study showing that outcome-based evaluation misrepresents agent behavior in workflow-realistic settings, failing to distinguish true failure, invalid execution, and correct remediation.
Embedded preview
March 2026
A 3,900-run study demonstrating that inference-based systems fail to satisfy core execution properties (correctness, determinism, temporal fidelity, and governance) while compiled execution achieves perfect reliability by removing inference from the runtime path.
Embedded preview
March 2026
An 850-run study showing that no existing system architecture (LLMs, structured systems, or agents) produces reliable execution, and that external governance, applied as a wrapper, enforces constraints but renders systems non-functional.
Embedded preview
March 2026
A 750-run study demonstrating that runtime governance can operate as domain-invariant infrastructure, producing consistent enforcement, tamper-evident records, and fully reconstructable decision trails across policy domains and model providers.
Embedded preview
April 2026
A 120-workflow, 12-domain study showing that compiled execution preserves determinism, complete policy-change visibility, and bit-identical replay at enterprise scale.
Embedded preview
April 2026
A 39-workflow study showing that natural-language enterprise workflow descriptions can be compiled into deterministic, governed execution artifacts with 100% compilation, execution, and replay determinism.
Embedded preview