NC

ncrofford.com

Paper

[P-01] Runtime Verification for Autonomous AI Agents: Toward Decision Evidence Infrastructure for Reliable Agent Systems

March 2026

Summary

A 540-run study showing that systems can reach correct outcomes while violating constraints during execution.

Key findings

  • Agents can achieve correct outcomes while still exhibiting flawed execution logic.
  • Failure signatures are classifiable and model-specific.
  • Trace-based runtime verification surfaces process failures that outcome-only evaluation misses.

Abstract

Autonomous AI systems are increasingly executing real operational actions such as modifying source code, invoking APIs, provisioning infrastructure, and orchestrating multi-step workflows. Despite this shift from conversational interfaces to operational agents, system reliability is still primarily evaluated using outcome-based metrics: whether the task ultimately appears to succeed. However, outcome evaluation obscures an important dimension of reliability. In a controlled experiment consisting of 540 autonomous agent executions across five frontier models, we observe repeated cases where agents achieve correct outcomes while misinterpreting verification signals or violating intermediate workflow constraints during execution. To address this gap, we introduce a runtime verification architecture for autonomous AI agents that captures execution traces, classifies behavioral failure signatures, evaluates policy compliance, and identifies intervention opportunities before task completion.