Reliability & Assurance

No Task Fails Every Time: Why One-Shot Audits Are Structurally Blind to Agent Damage

arXiv cs.AI Surfaced Read the original

Background research did not complete for this item, so it carries the headline and the summary only.

A pre-registered, multi-model study shows that single-run audits reliably miss stochastic, irreversible-action damage in agentic systems, challenging the validity of one-shot evaluation as a safety assurance method.

rss · arXiv cs.AI · Aug 18, 04:00

Tags: #agent reliability, #evaluation methodology, #AI audits, #multi-agent systems, #benchmark design