AI Agent Open-Source Weekly W26: Terminal Agents, Computer Use, and World-Model Repos Accelerate Together
This report covers 2026-06-22 to 2026-06-28 in Asia/Taipei time. The open-source story this week split into three accelerating layers: terminal coding agents (Codex, Claude Code, Gemini CLI, Goose), computer-use infrastructure (CUA, browser-use, Playwright MCP), and agent training/evaluation environments (Qwen-AgentWorld). Stars are only weak evidence, so this review emphasizes creation dates, pushes, releases, PR activity, and license.
1. Qwen-AgentWorld: the cleanest new repo signal of the week
QwenLM/Qwen-AgentWorld was the clearest new project signal. GitHub metadata shows it was created on 2026-06-22, describes itself as "Language World Models for General Agents," and uses Apache-2.0. At review time it had about 626 stars / 57 forks, with week-window commits updating the demo link, README, and a vLLM serving example.
The engineering implication is that agent training and evaluation need reusable environments, not just chat wrappers. If Qwen-AgentWorld connects to real tool traces, browser tasks, or desktop workflows, it could become useful infrastructure for agent evaluation and curriculum generation.
The risk is maturity. Star growth is fast, but the issue and PR ecosystem is still early. Treat it as research/prototyping infrastructure rather than a production dependency.
2. Codex, Claude Code, and Gemini CLI: terminal-agent competition is now release velocity
Three terminal coding-agent repositories showed strong activity during the week:
- openai/codex: about 94k stars / 13.9k forks, Apache-2.0. Week-window commits included TUI safety buffering, remote plugins enabled by default, and timeout work. Releases included
rust-v0.143.0-alpha.27through.29; updated PR count was about 641. - anthropics/claude-code: about 134k stars / 21.8k forks. Week-window releases included
v2.1.191,v2.1.193, andv2.1.195; updated PR count was about 17. - google-gemini/gemini-cli: about 105k stars / 14.2k forks, Apache-2.0. Releases included
v0.50.0-preview.1andv0.51.0-nightly; updated PR count was about 155.
The competition is no longer just prompt quality. It is plugin capability, remote tools, safety gates, IDE companions, patch approval, and team governance. Adoption decisions should compare workflow controls and ecosystem, not just demos.
3. CUA and browser-use keep building the computer-use layer
trycua/cua had about 19k stars / 1.2k forks, MIT license, and describes itself as open-source infrastructure for computer-use agents. Week-window releases included cloud-v0.1.1, train-v0.1.2, and sandbox-v0.1.17. The updated PR count was about 83, with activity around Linux/macOS/Windows drivers, interactive test harnesses, and an experimental MCP OAuth front door.
browser-use/browser-use had about 101k stars / 11.2k forks, MIT license, and week-window activity around browser event-bus reliability, Gmail integration, config IO, and Element click error handling. The updated PR count was about 54.
This layer matters because it moves agents beyond IDEs into browser and desktop control. It also carries the highest operational risk: OAuth tokens, desktop click scope, cross-OS drivers, prompt injection, and secret handling. Teams should pair these tools with sandboxes, credential isolation, and audit logs.
4. Goose, OpenAI Agents SDK, and Playwright MCP point to composable runtimes
aaif-goose/goose had about 50k stars / 5.4k forks, Apache-2.0, and shipped v1.39.0 during the week. Updated PR count was about 191, with hooks docs, CLI history, and desktop locale work.
openai/openai-agents-python had about 27k stars / 4.2k forks, MIT license, and released v0.17.7. Commits touched realtime validation log redaction, session iterator cancellation, and token usage handling.
microsoft/playwright-mcp had about 34k stars / 2.9k forks, Apache-2.0, and week-window docs/dependency updates. Its value is to expose browser automation as an MCP server so each agent does not need to rebuild browser control.
Together, these projects show the agent runtime becoming modular: LLM SDK, MCP tool server, browser automation, hooks, logs, approval gates, and runtime policy can be composed rather than bundled into one framework.
5. Magentic-UI remains a useful human-agent UI reference
microsoft/magentic-ui had about 9.9k stars / 994 forks, MIT license, and describes MagenticLite as working across browser and local file system. Week-window commits were mostly dependency and frontend fixes, with about 13 updated PRs. It is not moving at the same velocity as the terminal agents, but it remains useful for studying human-agent UI patterns and browser-local workflows.
This week's engineering judgment
Open-source agent infrastructure is becoming a stack rather than a single framework. Terminal agents own coding workflows. Computer-use repositories own browser and desktop control. MCP and SDKs own composition. World-model repositories may become the training and evaluation layer. Builders should select by layer instead of betting on one universal agent.
Watchlist
- Whether Qwen-AgentWorld releases benchmarks, datasets, a paper, or integrations with other agent projects.
- Whether Codex, Claude Code, and Gemini CLI converge on comparable plugin, remote-tool, and approval-gate models.
- Whether CUA productionizes sandboxing, credential isolation, and cross-OS parity.
- Whether browser-use reliability and credential/OAuth fixes lower long-task automation risk.
- Whether OpenAI Agents SDK, Playwright MCP, and Goose form stable MCP tool stacks.


