AI Agent Frontier Research Weekly W33: Interfaces and Trajectories Are the Real Variables (SWE-Bench ProMax, Recursive Synthesis, TraceSafe)

AI Agent Frontier Research Weekly W33: Interfaces and Trajectories Are the Real Variables (SWE-Bench ProMax, Recursive Synthesis, TraceSafe)

中文 EN

This edition covers August 10–16, 2026 in Asia/Taipei. W33's shared research lesson is that agent performance does not belong to the model alone. Tool interfaces, decomposition, and the full execution trajectory can change the conclusion.

From patches to large refactors

SWE-Bench ProMax pushes coding-agent evaluation toward large-scale multilingual refactoring. That broader task radius matters, but repository sampling, test coverage, contamination, and cost remain open checks. A high pass rate obtained through heavy retries is not yet predictable production capability.

The interface is an experimental condition

The Devil Is in the Interface studies how tool architecture shapes coding-agent behavior. The implication for benchmarks is direct: the same model can look different with another file API, shell feedback loop, or error format. Harnesses and tool contracts must be versioned before models are compared.

Recursive Synthesis for Long-Horizon Terminal Tasks decomposes long tasks into composable subproblems. Its practical promise is reducing single-path failure, but independent reruns across environments and models—under fixed budgets—are needed to separate algorithmic gains from extra token spend.

Guardrails must inspect trajectories

TraceSafe evaluates guardrails over multi-step tool-calling trajectories. Single-turn refusals miss sequences where each step appears acceptable but the composition exceeds authority. Useful safety evaluation should report dangerous-action interception, false positives on legitimate work, latency, and recovery.

Weekly take

The next unit of agent evaluation is not an answer; it is a complete trajectory in a versioned environment. Capability claims should ship with the harness, tool schema, budget, retry policy, and failure taxonomy.

Watchlist

  • ProMax contamination audits, complete cost accounting, and per-repository results.
  • Interface effects replicated across models and coding harnesses.
  • Recursive Synthesis under fixed-cost comparisons.
  • TraceSafe extensions to prompt injection, inter-agent authority, and real SaaS tools.