No Task Fails Every Time: Why One-Shot Audits Are Structurally Blind to Agent Damage
arXiv cs.AI Surfaced Read the original
Background research did not complete for this item, so it carries the headline and the summary only.
A pre-registered, multi-model study shows that single-run audits reliably miss stochastic, irreversible-action damage in agentic systems, challenging the validity of one-shot evaluation as a safety assurance method.
rss · arXiv cs.AI · Aug 18, 04:00
Tags: #agent reliability, #evaluation methodology, #AI audits, #multi-agent systems, #benchmark design