When Anthropic's engineering team needed to scale Claude's agent capabilities, they didn't invent yet another AI framework. They reached for the oldest wisdom in computer science: operating system design. This article dissects the Managed Agents architecture published in February 2026, examining how OS virtualization principles solve the hardest engineering problems in AI agent systems.
From Pets to Cattle: The Infrastructure Paradigm Shift
Understanding Managed Agents requires understanding the problem it solves.
Anthropic's original agent architecture coupled all components — inference engine (Brain), execution environment (Hands), and state management (Session) — into a single container. This created the classic distributed systems anti-pattern: pet servers.
The "pets vs cattle" metaphor originates from Microsoft engineer Bill Baker (2011-2012), popularized by Randy Bias at Cloudscaling. Pets are named, hand-tended, irreplaceable servers — "Bob the mail server goes down, it's all hands on deck." Cattle are numbered, automated, disposable — "one goes down, it's taken out back, shot, and replaced on the line."
Anthropic's agent containers were textbook pets. Container failure meant total session loss, requiring engineers to manually debug via WebSocket event streams. Worse, untrusted LLM-generated code ran alongside authentication credentials in the same container — a successful prompt injection only had to convince Claude to read its own environment variables to exfiltrate every access token.
The Three-Layer Decoupled Architecture
Managed Agents virtualizes an agent into three independent components:
Session: The Immutable Event Log
The Session is an append-only event log, durably stored outside any single service. It exposes three key interfaces:
emitEvent(id, event)— write eventsgetSession(id)— retrieve the complete event loggetEvents()— flexible positional queries
This directly implements the event sourcing pattern from distributed systems. In traditional event sourcing, an append-only store captures all state-changing events as the single source of truth. State is reconstructed by replaying events. Materialized views provide optimized read projections.
The mapping is precise: the session event log is the event store; getEvents() with positional slicing enables temporal queries; event transformation before context injection creates materialized views.
The critical advantage is reversibility. Other frameworks force irreversible decisions about what to retain — sliding windows discard history, summarization loses fidelity. Event sourcing preserves the complete log and defers the "what to show" decision to query time.
Brain: The Stateless Inference Engine
The Brain is Claude plus its Harness (the control loop). The key design decision: it is completely stateless.
When a Brain crashes, no session data is lost. A new Brain boots via wake(sessionId), reconstructs context from the event log, and resumes from the last recorded event — exactly like an OS restoring process state from saved registers.
This transforms Brains from pets into cattle. Any Brain instance is replaceable because state doesn't live there.
Hands: Isolated Execution Environments
Hands are sandboxed execution environments exposed through a universal interface:
execute(name, input) → string
This interface mirrors Unix's read() syscall. read() has remained unchanged from 1970s disk packs to modern SSDs — it's agnostic to the underlying implementation. Similarly, execute() abstracts containers, phones, and emulators behind the same interface. The harness doesn't need to know whether it's talking to a Docker container or an Android emulator.
The OS Metaphor Is Not Just a Metaphor
Anthropic explicitly references OS design principles: "Operating systems solved this problem by virtualizing hardware into abstractions — process, file — general enough for programs that didn't exist yet."
This isn't marketing. Academia has been systematically exploring this correspondence:
The AIOS paper from Rutgers University formally defines "LLM system calls" with a complete taxonomy: agent scheduling (FIFO, Round Robin), context switching via beam-search-tree snapshots, memory management (short-term runtime logs vs. long-term persistent storage), and access control via privilege groups.
The Agent-Kernel project from Zhejiang University adopts an explicit microkernel architecture: a minimal core handles plugin registration, communication, and verification, while specialized modules run as plugins — directly mirroring how microkernels delegate drivers and filesystems to user-space servers. The system demonstrated management of 10,000 concurrent agents.
| OS Concept | Managed Agents Equivalent | Explanation |
|---|---|---|
| Process | Agent Session | Execution unit with its own lifecycle |
| Process Scheduler | Harness orchestration loop | Decides what to execute next |
| System Call | execute(name, input) → string | Stable interface between Brain and Sandbox |
| Virtual File System | getEvents() session interface | Abstraction layer over the event log |
| Context Switch | Stateless Brain re-hydration from session log | Restoring execution from saved state |
| Kernel/Userspace boundary | Brain-Sandbox separation via MCP | Credentials never enter the Sandbox |
Security Boundaries: Solving the "Lethal Trifecta"
Simon Willison identifies the "lethal trifecta" of AI agent security: an agent that (1) accesses untrusted content, (2) has tools with side effects, and (3) holds privileged credentials — all in the same execution context.
Anthropic's coupled design embodied this problem perfectly. The decoupled architecture solves it with two patterns:
Git access: Access tokens are wired into git remotes during sandbox initialization. The agent uses git push/pull without ever handling tokens directly.
Custom tools: OAuth tokens reside in a secure vault outside the sandbox. Claude calls MCP tools through a dedicated proxy that fetches credentials from the vault before making external calls. Tokens are never reachable from within the sandbox.
This isn't a theoretical risk. Real-world agent security incidents in 2025-2026 demonstrate the urgency:
- CVE-2026-25049 (n8n): CVSS 10.0 — arbitrary command execution and decryption of all stored credentials
- EchoLeak (2025): Zero-click prompt injection in Microsoft Copilot exfiltrating data from OneDrive, SharePoint, and Teams
- Perplexity Comet: Hidden Reddit comments triggered AI to log into users' email and transmit credentials within 150 seconds
- AWS AgentCore sandbox bypass: DNS tunneling circumvented "complete isolation" to access IAM role credentials
According to Obsidian Security, 73% of production AI deployments encountered prompt injection attacks in 2025.
Context Management: Why Event Sourcing Wins
Long-horizon tasks inevitably exceed Claude's context window. Frameworks handle this differently:
LangGraph uses checkpointing with trimMessages() to truncate history, or inserts summary nodes for compression. CrewAI provides a unified Memory class backed by ChromaDB (short-term) and SQLite (long-term). Manus treats the filesystem as "the ultimate context — unlimited in size, persistent by nature, directly operable by the agent," with aggressive tool output pruning.
JetBrains research provides compelling evidence: on SWE-bench, observation masking — replacing older tool outputs with placeholders while preserving reasoning/actions — outperformed LLM summarization in 4 of 5 scenarios. Summarization caused agents to run ~15% longer and consumed 7%+ of total costs for summary generation itself.
Anthropic's event sourcing approach occupies a unique position on this spectrum. Rather than irreversible compression, it preserves the complete log and lets the harness decide how to transform events before context injection — achieving "context engineering for a high prompt cache hit rate."
Production data from the Manus team supports this direction: KV-cache hit rate is "the single most important metric for a production-stage AI agent." On Claude Sonnet, the cost difference between cached (3/MTok) tokens is 10x.
Many Brains, Many Hands: Scaling to N-by-M
The original architecture required one container per Brain, creating startup delays. After decoupling:
- Inference starts immediately: Brains pull pending events from the session log without waiting for container provisioning
- p50 time-to-first-token dropped ~60%
- p95 TTFT dropped >90%
More importantly, multiple Brains can share multiple Hands, and agents can delegate work to other agents. Each Hand is an MCP tool exposed through execute(name, input) → string. MCP (Model Context Protocol) solves the N x M integration problem: instead of N apps each building M tool integrations (N x M total), you need only N + M integrations.
As of early 2026, MCP has reached 97M+ monthly SDK downloads, 10,000+ active servers, first-class support in Claude, ChatGPT, Cursor, Gemini, and VS Code, and was donated to the Linux Foundation in December 2025.
How It Compares to the Orchestration Landscape
Managed Agents occupies a fundamentally different layer than mainstream agent frameworks:
| Dimension | Managed Agents | LangGraph | CrewAI | OpenAI Agents SDK |
|---|---|---|---|---|
| Architecture | Brain/Hands/Session decoupling | Directed graph + conditional edges | Role-based crews | Handoff-based delegation |
| State Management | External append-only event log | Checkpointing + time travel | Sequential task outputs | Session variables (ephemeral) |
| Scaling | Stateless harness, independent component scaling | Graph-level parallelism | Limited at scale | Lightweight, low overhead |
| Security Model | Vault + Proxy, credentials outside sandbox | Application-level | Application-level | Guardrails (input/output validation) |
LangGraph shows the lowest latency and token consumption in benchmarks. CrewAI claims 60%+ Fortune 500 adoption but teams often migrate to LangGraph at scale. OpenAI Agents SDK is lightweight but only supports OpenAI models.
The key distinction: these frameworks are orchestration layers — they define how agents coordinate work. Managed Agents is an infrastructure layer — it defines how computation, state, and execution run reliably in production. They operate at different levels of the stack and can be complementary.
Design Philosophy: Stability of Interfaces
Anthropic's team summarizes their philosophy: "We're opinionated about the shape of these interfaces, not about what runs behind them."
This echoes OS design perfectly. VFS is opinionated about read()/write(), not about whether the backend is ext4 or NFS. Managed Agents is opinionated about state manipulation (via Session), computation (via Sandbox), and many-to-many scaling, but unopinionated about specific harness implementations.
As models evolve rapidly, the key architectural question shifts from "How do I control the model?" to "What can I stop doing?" Claude Sonnet 4.5's "context anxiety" — wrapping tasks early near token limits — disappeared entirely on Claude Opus 4.5, turning harness patches for that behavior into dead weight.
Stable interfaces let implementations change with model evolution without rewriting the system. This is exactly why read() has survived since the 1970s.
Takeaways for Agent Builders
The core insight of Managed Agents isn't a technical innovation — it's the migration of engineering wisdom. The abstractions that operating systems spent decades building — Process, Syscall, VFS, Kernel/Userspace boundary — find precise correspondences in AI agent systems.
Three actionable principles for engineers building agent systems:
- Decouple state: Separate agent state from the execution environment. Use event sourcing over sliding windows or summarization
- Isolate credentials: Never let agent-generated code access authentication tokens. Use the proxy pattern to inject credentials
- Design stable interfaces: Invest in abstractions that won't become obsolete as model capabilities change.
execute(name, input) → stringis more durable than any specific prompt engineering technique
In the AI agent battlefield, the most durable weapon isn't the latest model — it's the most stable abstraction.
References
- Scaling Managed Agents: Decoupling the brain from the hands — Anthropic engineering blog, the original Managed Agents architecture article
- AIOS: LLM Agent Operating System — Rutgers University paper formalizing LLM system call taxonomy
- Agent-Kernel: A MicroKernel Multi-Agent System Framework — Zhejiang University paper applying microkernel architecture to multi-agent systems
- The History of Pets vs Cattle — Randy Bias's definitive account of the Pets vs Cattle metaphor origin
- Context Engineering for AI Agents: Lessons from Building Manus — Manus team's production context engineering insights
- JetBrains Research: Efficient Context Management — Observation masking vs summarization benchmarks on SWE-bench
- Event Sourcing: The Backbone of Agentic AI — Event sourcing pattern applied to agent architectures
- Securely deploying AI agents — Anthropic's secure deployment guide with proxy pattern and isolation techniques
- How to sandbox AI agents — AI agent sandboxing technology comparison (containers, microVMs, gVisor)
- A Year of MCP: From Internal Experiment to Industry Standard — MCP ecosystem growth review
- OWASP Top 10 for LLM Applications 2025 — LLM application security risk taxonomy
- AgentFold: Long-Horizon Web Agents — Proactive context management with 92% token reduction


