AI Agent Weekly Overview W31: Capability Becomes Reliability Engineering (Industry × Research × Open Source)

AI Agent Weekly Overview W31: Capability Becomes Reliability Engineering (Industry × Research × Open Source)

中文 EN

This overview covers July 27–August 2, 2026 in Asia/Taipei. W31's common signal was that capability competition continues, but value is moving toward reliability engineering. “Can the model do it?” is giving way to whether the system can constrain, observe, verify, and recover the work.

This week's three deep dives

What the three tracks imply together

In industry, cyber-evaluation incidents, a multi-vendor security alliance, and realtime voice deployment forced identity, credentials, egress, human takeover, and incident reporting into the product specification. An agent is not a model endpoint; it is a continuously executing service that may change external state.

In research, HANDBOOK.md, ORCA-bench, and PatientAgentBench expose gaps in policy compliance, on-call root-cause analysis, and high-stakes patient dialogue. Their common message is that success rate is insufficient. Evaluations need violations, wrong evidence, irreversible actions, calibration, and intervention.

In open source, Flue, Headroom, Open Code Review, and MAI-UI turn sandboxing, the context data plane, hybrid rules, and GUI grounding into replaceable layers. This accelerates composition while leaving security, licensing, and trace compatibility with adopters.

Action for engineering teams

Classify agent actions as read-only, reversible, or irreversible, then assign permissions, approvals, traces, and recovery to each class. Evaluate task success, policy violations, takeover rate, and full cost per successful task together. Preserve raw tool output and replayability so compression or summarization cannot erase incident evidence.

Watchlist

  • Complete root-cause reports for real evaluation incidents.
  • Policy following, abstention, and on-call diagnosis in standard vendor evals.
  • Portable trace, policy, and sandbox interfaces across harness projects.
  • Published failure and takeover rates for voice, health, and GUI agents.
  • Cost per successful task replacing token price as the buying metric.