1. FRI · MAY 29

    Agent capability is accelerating, but the security and reliability surface area is expanding faster. Today's items highlight the tension between shipping autonomous workflows and defending against prompt injection, data hallucination, and supply chain attacks.

    6 news · 3 blogs · 5 appendix
    code-agentsmemorydevops-agentscost-latencyevalsframeworks
    Read issue →
  2. WED · MAY 27

    Production agent capabilities are outpacing safety guardrails and operational support. Today's items contrast new benchmarks for long-horizon coding and desktop automation against concrete data exfiltration risks and the overwhelming burden AI-generated security reports place on maintainers.

    3 news · 2 blogs · 2 appendix
    safetytool-usecode-agentsinfraevalsdevops-agents
    Read issue →
  3. MON · MAY 25

    This week's coverage converges on agent trustworthiness in production: evidence verification, contractual boundaries, failure taxonomies, and the human cost of AI-generated slop. The field is moving from capability demos to auditable, reliable systems.

    5 papers · 3 news · 3 blogs · 85 appendix
    code-agentstool-useinfracost-latencysafetyevals
    Read issue →
  4. FRI · MAY 22

    Today's coverage highlights the gritty work of productionizing agent tooling. While Simon Willison's Datasette Agent ecosystem iterates rapidly on permissions, SQL observability, and sandboxing plugins, broader infrastructure plays like Daytona and Runtime focus on environment snapshots and team-wide guardrails. The common thread is moving beyond chat interfaces to secure, auditable action.

    3 news · 6 blogs · 2 appendix
    frameworkstool-usecode-agentsinfraevalssafety
    Read issue →
  5. THU · MAY 21

    Today's picks converge on the operational reality of agents: reliability layers that boost small models, observability tools that auto-remediate, and cloud infrastructure built for agent spend rather than human users. The focus is less on capability and more on containment and cost.

    2 news · 2 blogs · 4 appendix
    cost-latencycode-agentssafetytool-useevalsobservability
    Read issue →
  6. WED · MAY 20

    Today's papers converge on architecting around stochasticity in production systems. From database tuning to runtime patterns and binary remediation, the focus shifts from capability to the boundary between agent output and system action.

    5 papers · 2 blogs · 19 appendix
    cost-latencytool-useinfraevalsdevops-agentsmemory
    Read issue →
  7. TUE · MAY 19

    This week's papers converge on a single production reality: agent safety failures aren't exotic—they're systematic. Overeager actions, untrusted tool feedback, and observation contract violations all emerge from the same gap between capability and authorization. The evals and infrastructure pieces show how to measure and constrain this.

    4 papers · 3 news · 44 appendix
    devops-agentsinfracode-agentssafetytool-useevals
    Read issue →
  8. MON · MAY 18

    Today's items cluster around two live threats to agent reliability: memory as an attack surface, and the token waste hiding inside standard tool-use patterns. Several items offer concrete countermeasures with measured results.

    5 papers · 2 news · 1 blog · 34 appendix
    evalstool-usecode-agentscost-latencyplanningsafety
    Read issue →
  9. FRI · MAY 15

    The harness is the product. Across today's featured items — from Hamel's eval skills to the grep-vs-vector retrieval study to Mozilla's 14-20x bug-finding jump — the consistent finding is that the scaffolding around the model matters more than the model itself.

    5 papers · 3 news · 6 blogs · 96 appendix
    evalscode-agentsobservabilitysafetymemorytool-use
    Read issue →
  10. THU · MAY 14

    Safety failures are concrete and measured this week: 29.9% of deployed agent skills have exploitable guardrail violations, 40.64% of high-risk OS operations execute through "aligned" models, and a single injected sentence flips flagship models to 91-98% unsafe. The week's production news responds in kind.