NC

ncrofford.com

Nesean Crofford

Work that runs without a person in it, and proof that it ran correctly.

Operator background, production systems, and the research that came out of building them.

I spent seven years inside PE-backed portfolio companies, mostly operations. I've built and shipped production platforms since. Most operational cost turns out to be people reading documents, and that is the part I work on.

What I'm working onMore about me

How I think about the work

Four things I'd argue for.

Most software makes you choose between powerful and simple. That trade is a design failure, not a law of nature.

Intelligence you rent gets more expensive the more you use it. Intelligence you own gets cheaper.

The same input should produce the same output. Every time, provably, years later.

A system that guesses when it is unsure is worse than one that stops and asks.

Current focus

What I'm building

Two layers that make operational work run unattended and prove afterward that it ran correctly. Together they are the foundation of Parity.

Featured research

Current papers and empirical studies

Empirical work examining failure modes and reliability gaps in autonomous systems. As AI systems move from generating outputs to taking actions, reliability becomes a systems problem, not a model problem.

March 2026

[P-01] Runtime Verification for Autonomous AI Agents: Toward Decision Evidence Infrastructure for Reliable Agent Systems

A 540-run study showing that systems can reach correct outcomes while violating constraints during execution.

March 2026

[O-01] When Language Models Fail as Execution Systems: A Study of Exactness, Conformance, and Determinism

A 1,260-run study showing that language models fail as execution systems under exact-match constraints, while deterministic systems remain stable.

March 2026

[P-02] Runtime Governance for Workflow-Realistic AI Agents: Validity, Remediation, and the Limits of Outcome-Based Evaluation

An 850-run study showing that outcome-based evaluation misrepresents agent behavior in workflow-realistic settings, failing to distinguish true failure, invalid execution, and correct remediation.

March 2026

[O-02] The Case for Compiled Execution: Execution Properties in High-Stakes AI Systems

A 3,900-run study demonstrating that inference-based systems fail to satisfy core execution properties (correctness, determinism, temporal fidelity, and governance) while compiled execution achieves perfect reliability by removing inference from the runtime path.

March 2026

[P-03] The Governance Void Is Architectural: System Architecture, External Governance, and the Case for Runtime Infrastructure

An 850-run study showing that no existing system architecture (LLMs, structured systems, or agents) produces reliable execution, and that external governance, applied as a wrapper, enforces constraints but renders systems non-functional.

March 2026

[P-04] Runtime Governance as Invariant Infrastructure: Domain-Invariant, Model-Agnostic Enforcement with Reconstructable Evidence Chains

A 750-run study demonstrating that runtime governance can operate as domain-invariant infrastructure, producing consistent enforcement, tamper-evident records, and fully reconstructable decision trails across policy domains and model providers.

April 2026

[O-03] Governance at Scale Through Compiled Execution: Structural Guarantees Across 120 Workflows and 12 Domains

A 120-workflow, 12-domain study showing that compiled execution preserves determinism, complete policy-change visibility, and bit-identical replay at enterprise scale.

April 2026

[O-04] From Intent to Infrastructure: Compiling Natural-Language Workflow Descriptions into Governed, Deterministic Execution Artifacts at Scale

A 39-workflow study showing that natural-language enterprise workflow descriptions can be compiled into deterministic, governed execution artifacts with 100% compilation, execution, and replay determinism.