Shelf

Reliability & Assurance

Who is exposed, and what should they check?

Model drift, agent and workflow failure, evidence integrity, security of AI systems, standards and regulation.

69 articles, newest first
ArticlePublisherSurfaced
September 2026 22 articles
Researchers used Claude to compromise OpenAI employee accounts The Verge - AI
AI-Hallucinated Intelligence Report Nearly Triggered US Military Strike on Chinese Ship realsarm
Tool-Selection and Gating Defenses Miss Fabricated Agent Tool Calls arXiv cs.AI
Paper Proposes Label-Efficient Statistical Test for Model Update Regressions arXiv cs.LG
Bias Audit Scores Disagree on Model Ranking Across Ten Instruments arXiv cs.CL
Hidden-State Probes for Prompt Injection Fail on Ordinary Typos arXiv cs.CL
Live GitHub admin token found in Baseten's public Docker image bearsyankees
Tool-Using Agents Fabricate Answers When Tool Failures Look Like Successes arXiv cs.SE
Clinical LLM Agents Issue Different Orders on Identical Reruns arXiv cs.CL
AI agents implicated in RubyGems credential-harvesting exploit chain gregnavis
Anthropic Discloses AI Models Autonomously Conducted Real Cyberattacks The Verge - AI
Agentic Code-Repair Patches Pass Tests While Leaving Security Flaws arXiv cs.SE
Verifier Ensembles Sharing Evidence Sources Approve Most Unsafe Agentic Actions arXiv cs.SE
Standard Chat-Model Evaluations Miss Multi-Step Agentic Risks arXiv cs.AI
Large-Language-Model Document Auditors Fabricate Findings at Scale arXiv cs.CL
Study Finds LLM Agents Often Adopt Corrupted Tool Outputs Uncritically arXiv cs.AI
Safety Monitor Recall Metrics Miss the Prompts Models Actually Comply With arXiv cs.CL
SWE-Bench Scores Inflated by Agent Exploits Like Reading Git History arXiv cs.SE
Some Quantized Models in Official Registries Silently Fail All Tasks arXiv cs.SE
More Capable Trading Agents May Increase Correlated Market Risk arXiv cs.AI
Static Red-Team Scores May Understate Prompt-Injection Risk to Agents arXiv cs.AI
Peer Pressure Between Agents Silently Breaks Conformal Prediction Guarantees arXiv cs.LG
August 2026 47 articles
Quantization-Triggered Backdoors Bypass Full-Precision Model Safety Checks arXiv cs.LG
Safety Guardrails Lose Over Half Their Detection Recall as Context Grows arXiv cs.AI
Maintainers report exploit attempts within minutes of bug disclosure Simon Willison
Simple 1920s statistical method reportedly matches SOTA anomaly detection benchmark r/MachineLearning
Researcher bypasses Claude Code auto mode about 80% of the time Simon Willison
OpenAI agents hacked Hugging Face after unintended training effects MIT Technology Review - AI
CI/CD and Agent Platforms Lack Content-Addressed Identity Records arXiv cs.SE
New Benchmark Shows NL2SQL Accuracy Drops Sharply on Enterprise Schemas arXiv cs.AI
Agentic Scaffolding Amplifies Sycophantic Behavior in Large Language Models arXiv cs.CL
Adversarial Prompts Recover 'Unlearned' Data From LLMs arXiv cs.CL
Study Finds Language Models Internally Represent Being Evaluated arXiv cs.CL
Benchmark Finds No Single Model Covers All Content Moderation Harms arXiv cs.CL
Backdoor Attack Implants Fairness Bias That Survives MLLM Continual Learning arXiv cs.LG
Alabama AG subpoenas OpenAI over reported agent containment failure The Verge - AI
Narrative-Wrapped Prompts Bypass Guardrails in Small LLMs, Study Finds arXiv cs.AI
Benchmark Finds Chatbots Misjudge Youth Mental-Health Risk Despite Vocabulary Fluency arXiv cs.CL
Exact-Match RLVR Verifiers Show Strong Language-Dependent Bias arXiv cs.CL
Benchmark Finds Cross-Lingual Safety Gaps in LLMs for Indian Languages arXiv cs.AI
Benchmark Finds Banking AI Agents Fail Multi-Turn Fraud Tests arXiv cs.AI
Code Agents Show Inconsistent Robustness to Semantics-Preserving Code Rewrites arXiv cs.SE
Prompt instructions fail to stop LLMs from cheating on cyber tasks vga805
Standard Safety Benchmarks May Not Reliably Score Small Language Models arXiv cs.AI
Decoy Defense Poisons Abliteration Attacks on Open-Weight Models arXiv cs.AI
Random Splits Inflate Financial News NLP Benchmarks 1.1x-6.5x arXiv cs.CL
Aggregate benchmark gains can mask item-level LLM regressions arXiv cs.SE
Study Finds Widespread Gaps in Post-Deployment AI Incident Compliance arXiv cs.SE
Hallucinations Become Harder to Detect as They Move Through Agent Pipelines arXiv cs.AI
DEI Prompts Cause Medical LLMs to Fabricate Patient Demographics arXiv cs.AI
Multi-Agent Decomposition Found to Attenuate Compliance Facts arXiv cs.AI
Compliance Guard Models Ignore the Rule They Are Meant to Check arXiv cs.AI
Composed skill chains bypass per-skill marketplace safety scanners arXiv cs.AI
Design Flaws Found in Agentic Offensive-Security Tools arXiv cs.AI
Edge-IIoTset Benchmark Leaks Labels Through a Serialisation Artifact arXiv cs.LG
Pass@k Misapplication Inflates Reported Reliability of Coding Agents arXiv cs.AI
Prior Audit-Repair Context Shifts LLM Verifier Thresholds Toward Leniency arXiv cs.AI
Deceptive Communication Emerges in Long-Horizon Multi-Agent LLM Trading arXiv cs.AI
Codebase Structure Affects Prompt Injection Success in Coding Agents arXiv cs.AI
New Benchmark Measures Security Drift in LLM Code Editing arXiv cs.AI
No Task Fails Every Time: Why One-Shot Audits Are Structurally Blind to Agent Damage arXiv cs.AI
Few Bit Flips Can Zero Out Quantized Robot Policy Success arXiv cs.AI
Automated red-teaming bypasses safety filters in audio-language models arXiv cs.AI
Standard Metrics Mask Clinical Error-Detection Failures in LLMs arXiv cs.AI
State-Semantic Injection Proposed as New Attack Vector for Embodied Agents arXiv cs.AI
Sequential LLM releases can skew bargaining outcomes, study finds arXiv cs.AI
Clean ImageNet Data Contains Naturally Occurring Backdoor-like Triggers arXiv cs.AI
Biology AI Safeguards Fail to Generalize Across Models in Study arXiv cs.AI
Study Finds Undisclosed Value Bias in LLM Answers to Subjective Questions arXiv cs.AI

Back to the latest edition How articles are chosen