NC

ncrofford.com

Paper

[P-02] Runtime Governance for Workflow-Realistic AI Agents: Validity, Remediation, and the Limits of Outcome-Based Evaluation

March 2026

Summary

An 850-run study showing that outcome-based evaluation misrepresents agent behavior in workflow-realistic settings, failing to distinguish true failure, invalid execution, and correct remediation.

Key findings

  • Outcome-based evaluation overstates failure: fewer than half of labeled failures are genuine decision errors.
  • 15.8% of runs represent valid remediation: operationally correct behavior misclassified as failure under outcome-only scoring.
  • Structural workflow validity and decision correctness are separable dimensions of agent performance.
  • Runtime governance resolves all ambiguous outcomes, enabling full classification of agent behavior at execution time.

Abstract

Autonomous AI systems are increasingly being positioned inside operational settings such as incident triage, deployment management, migration handling, release gating, code review, and compliance workflows. This shift is not hypothetical. Agents are already being tested and deployed in internal automation, CI/CD-adjacent processes, support operations, and infrastructure-facing workflows. In these environments, failures are not merely benchmark misses; they can produce regressions, outages, and compliance risk. The relevant question is therefore not simply whether a task appears to succeed. It is whether the system gathered the required evidence, interpreted that evidence correctly, followed the required sequence of steps, chose the appropriate consequence action when success became impossible, and terminated the workflow in a manner consistent with operational constraints. This paper argues that workflow-realistic agent evaluation requires a different lens. We present results from Probity Experiment P-02, a controlled runtime governance study consisting of 850 autonomous agent executions across five language models and ten workflow-realistic scenarios spanning incident response, deployment and migration workflows, approval-gated changes, release validation, review workflows, and compliance evidence collection. Each run is evaluated through three complementary layers: validity (did the model execute a structurally sound workflow?), a governance-aware view (did the model make the correct operational decision, including remediation?), and a legacy outcome-based view retained as a contrastive baseline. The legacy outcome-based layer reports only 84/850 task successes (9.9%) and labels 766/850 runs as failures. The governance-aware view reveals a materially different distribution: 357 runs (42.0%) are genuine model decision failures, 134 runs (15.8%) are valid remediation (correct detection and handling of failure conditions), 219 runs (25.8%) are structurally invalid workflows, and 65 runs (7.6%) are system failures. All 242 runs classified as unknown_failure_mode by the legacy outcome-based layer are resolved by the governance-aware view, with zero unresolved cases remaining. These results show that outcome-based evaluation systematically overstates failure and obscures the distinction between true failure, invalid execution, and operationally correct remediation. More broadly, they suggest that there is no reliable path to deploying autonomous AI systems in workflow-realistic environments without infrastructure capable of observing and governing runtime decision processes, not merely scoring outcomes.