NC

ncrofford.com

Research

Papers, experiments, and supporting material

A growing archive of empirical work on reliable AI systems.

March 2026

[P-01] Runtime Verification for Autonomous AI Agents: Toward Decision Evidence Infrastructure for Reliable Agent Systems

A 540-run study showing that systems can reach correct outcomes while violating constraints during execution.

Embedded preview

March 2026

[O-01] When Language Models Fail as Execution Systems: A Study of Exactness, Conformance, and Determinism

A 1,260-run study showing that language models fail as execution systems under exact-match constraints, while deterministic systems remain stable.

Embedded preview

March 2026

[P-02] Runtime Governance for Workflow-Realistic AI Agents: Validity, Remediation, and the Limits of Outcome-Based Evaluation

An 850-run study showing that outcome-based evaluation misrepresents agent behavior in workflow-realistic settings, failing to distinguish true failure, invalid execution, and correct remediation.

Embedded preview

March 2026

[O-02] The Case for Compiled Execution: Execution Properties in High-Stakes AI Systems

A 3,900-run study demonstrating that inference-based systems fail to satisfy core execution properties (correctness, determinism, temporal fidelity, and governance) while compiled execution achieves perfect reliability by removing inference from the runtime path.

Embedded preview

March 2026

[P-03] The Governance Void Is Architectural: System Architecture, External Governance, and the Case for Runtime Infrastructure

An 850-run study showing that no existing system architecture (LLMs, structured systems, or agents) produces reliable execution, and that external governance, applied as a wrapper, enforces constraints but renders systems non-functional.

Embedded preview

March 2026

[P-04] Runtime Governance as Invariant Infrastructure: Domain-Invariant, Model-Agnostic Enforcement with Reconstructable Evidence Chains

A 750-run study demonstrating that runtime governance can operate as domain-invariant infrastructure, producing consistent enforcement, tamper-evident records, and fully reconstructable decision trails across policy domains and model providers.

Embedded preview

April 2026

[O-03] Governance at Scale Through Compiled Execution: Structural Guarantees Across 120 Workflows and 12 Domains

A 120-workflow, 12-domain study showing that compiled execution preserves determinism, complete policy-change visibility, and bit-identical replay at enterprise scale.

Embedded preview

April 2026

[O-04] From Intent to Infrastructure: Compiling Natural-Language Workflow Descriptions into Governed, Deterministic Execution Artifacts at Scale

A 39-workflow study showing that natural-language enterprise workflow descriptions can be compiled into deterministic, governed execution artifacts with 100% compilation, execution, and replay determinism.

Embedded preview