2026-06-15
AI Agents
Today's work zeroes in on how agents fail. We're seeing a wave of new benchmarks designed to find subtle but critical vulnerabilities: decomposition attacks, social engineering in PRs, latent planning errors, and even denial-of-service against the guardrails themselves. This shift from capability to reliability is driving more robust architectural patterns, from privacy-preserving UI brokers to specialized sub-agents for efficiency.
Papers
15Bayesian-Calibrated Detection of Hallucinated Package Imports in AI-Assisted Code
Bayesian-Calibrated Detection of Hallucinated Package Imports in AI-Assisted Code
Closing the Reflection Gap: A Free Calibration Bonus for Agentic RL
Closing the Reflection Gap: A Free Calibration Bonus for Agentic RL
Hidden in Plain Sight: Benchmarking Agent Safety Against Decomposition Attacks with DECOMPBENCH
Hidden in Plain Sight: Benchmarking Agent Safety Against Decomposition Attacks with DECOMPBENCH
SEVRA-BENCH: Social Engineering of Vulnerabilities in Review Agents
SEVRA-BENCH: Social Engineering of Vulnerabilities in Review Agents
Towards Direct Latent-Space Synthesis for Parallel Branches in LLM-Agent Workflows
Towards Direct Latent-Space Synthesis for Parallel Branches in LLM-Agent Workflows
StreamMemBench: Streaming Evaluation of Agent Memory for Future-Oriented Assistance
StreamMemBench: Streaming Evaluation of Agent Memory for Future-Oriented Assistance
SIMMER: Benchmarking Latent Failures in LLM Executable Planning with a World Model
SIMMER: Benchmarking Latent Failures in LLM Executable Planning with a World Model
Minim: Privacy-Aware Minimal View for Agents via Trusted Local Sanitization
Minim: Privacy-Aware Minimal View for Agents via Trusted Local Sanitization
When the Tool Decides: LLM Agents Defer Blindly to Graph Neural Network Tools, and Stronger Backbones Defer More
When the Tool Decides: LLM Agents Defer Blindly to Graph Neural Network Tools, and Stronger Backbones Defer More
From Shield to Target: Denial-of-Service Attacks on LLM-Based Agent Guardrails
From Shield to Target: Denial-of-Service Attacks on LLM-Based Agent Guardrails
HarnessX: A Composable, Adaptive, and Evolvable Agent Harness Foundry
HarnessX: A Composable, Adaptive, and Evolvable Agent Harness Foundry
AgentCyberRange: Benchmarking Frontier AI Systems in Realistic Cyber Ranges
AgentCyberRange: Benchmarking Frontier AI Systems in Realistic Cyber Ranges
FastContext: Training Efficient Repository Explorer for Coding Agents
FastContext: Training Efficient Repository Explorer for Coding Agents
SkillAudit: Ground-Truth-Free Skill Evolution via Paired Trajectory Auditing
SkillAudit: Ground-Truth-Free Skill Evolution via Paired Trajectory Auditing
tap: A File-Based Protocol for Heterogeneous LLM Agent Collaboration
tap: A File-Based Protocol for Heterogeneous LLM Agent Collaboration
News
2Why Is Claude Turning into an a**Hole?
A practitioner's first-hand account of Claude models becoming overly-censorious and unhelpful, even for benign prompts. This isn't just a complaint; it's a documented case of model degradation or a safety-alignment change with real-world consequences for developers relying on model stability. The post catalogs specific examples of previously working prompts now failing, highlighting the operational risk of undocumented behavioral drift in foundation models.
Ponytail – make your AI agent think like the laziest senior dev in the room
Ponytail – make your AI agent think like the laziest senior dev in the room
8 more items the ranker flagged but didn't feature
Papers
- Formalizing Numerical Analysis: An Agent Pipeline and Quality Audit Beyond Kernel Acceptancearxiv:arxiv-agents
- VeriGeo: Controllable Geometry Question Generation with Numerical and Analytical Verificationarxiv:arxiv-agents
- Security in a Workflow: Exploring Role-Based Agentic Architectures for Vulnerability Handlingarxiv:arxiv-software-eng
- LLM Agents Can See Code Repositoriesarxiv:arxiv-software-eng
- When Plausible Is Not Realistic: Evaluating Human Mobility in LLM-Based Urban Simulationarxiv:arxiv-agents
- No Accidental Software Agent First Canonical Code for Human Code Entropy Reduction and 30 to 500 times Lower Frontier Model Requirementsarxiv:arxiv-agents
- From Chatbot to Digital Colleague: The Paradigm Shift Toward Persistent Autonomous AIarxiv:arxiv-agents
- A Benchmark and Framework for Evaluating Next Action Predictions in Spreadsheetsarxiv:arxiv-software-eng