Home / State of AI / Agents & Tooling

Agents & Tooling

AI agents, tool use, MCP protocol, and autonomous systems

Living document tracking AI agents, tool use, MCP protocol, and autonomous systems. Newest entries appear at the top.

Key Areas to Watch

  • Agent frameworks and architectures
  • Model Context Protocol (MCP) ecosystem
  • Tool use and function calling
  • Computer use and browser automation
  • Multi-agent systems and orchestration

2026-W39

  • RecreationWorld — Five-platform environments test agents that interleave GUI exploration, coding, execution, and visual verification. [arXiv]
  • AutoViewMem — Self-configuring memory views reduce interference among preferences, events, constraints, and temporal updates. [arXiv]
  • An Interpretable Memory Decision Controller for LLM Agents — Explicitly decides whether retrieved memory should be trusted when stored evidence conflicts. [arXiv]
  • How V7 gives AI agents institutional memory — Source-linked organizational context shows a concrete retrieval and provenance pattern for long-lived agents. [OpenAI]
  • MCP was always a bad idea? — MCP remains useful when agents need curated capabilities instead of unrestricted terminal and network access. [Simon Willison]

2026-W38

  • OpenAI Agents API — Scaled agent runtime exposed as an API, useful for teams moving beyond chat-only workflows. [OpenAI]
  • ToolGrad — Generates better tool-use data with textual gradients, relevant to agent training and evaluation loops. [Google Research]
  • T1 terminal agent — RL-trained MoE terminal agent running real shell tasks for long horizons, pushing evals closer to coding work. [x.com]
  • EvoSafeHarness — Evolves model/domain-specific safety harnesses around frozen agents, treating scaffolding as a safety layer. [x.com]

2026-W37

  • CUA-Universe — Scalable environment for hybrid GUI+CLI agents, useful for realistic computer-use evaluation. [arXiv]
  • Agent memory portability — Tests whether long-lived agent memory survives model upgrades. [arXiv]
  • Terminal-Universe — Reconstructs realistic terminal environments from agent trajectories for training and evaluation. [x.com]
  • StarHarness — Searches for better agent scaffolds around fixed models, highlighting harness design as an optimization target. [x.com]

2026-W36

2026-W32

  • Orchard — Open framework for training and evaluating agents across task types. [Microsoft Research]
  • Echoverse — Evolving environments target more realistic computer-use agent training. [Microsoft Research]
  • EvoLib — Turns agent experience into reusable, evolving knowledge across tasks. [Microsoft Research]
  • Real-Time Detection and Repair of LLM Agent Failures — Runtime telemetry monitors catch loops, drift, fabricated results, and tool cascades during agent episodes. [arXiv]

2026-W28

  • SkillOpt — Agent skills are being treated as trainable parameters rather than hand-edited prompts. [Microsoft Research]
  • Memora — Agent memory design moved toward balancing abstraction with specific conversational detail. [Microsoft Research]
  • Distributed Attacks in Persistent-State AI Control — Persistent codebase state creates a new attack surface for multi-PR agent behavior. [arXiv]

2026-W27

2026-W26

  • LedgerAgent — Structured state for policy-adherent tool-calling agents, especially customer-service style workflows with domain rules. [arXiv]
  • MosaicLeaks — Benchmark for whether research agents can keep secrets while working with sensitive context. [Hugging Face]
  • UltraQuant — 4-bit KV caching aimed at context-heavy agents with reused long prefixes and high concurrency. [arXiv]
  • Agentic Resource Discovery — Hugging Face adds agent-oriented resource search over Hub artifacts. [Hugging Face]

2026-W25

2026-W24

2026-W23

2026-W22

  • SkillOpt — Treats reusable agent skills as optimizable external state, pointing toward versioned procedures, evals, and regression-tested agent capabilities. [arXiv]
  • MagenticLite, MagenticBrain, Fara 1.5 — Microsoft Research pushes agentic workflows toward smaller, cheaper models via tighter orchestration. [Microsoft Research]
  • Forge — Guardrails reportedly move an 8B model from 53% to 99% on agentic tasks, underscoring how much agent reliability lives in scaffolding. [HN]

2026-W20

  • OpenAI Codex on mobile — Live agent environments become reviewable and approvable from a phone, pushing async/remote agent operation toward the mainstream. [TechCrunch]
  • SkillGenBench — A benchmark for how well agents synthesize reusable skills, relevant as skill/tool generation becomes a core agent capability. [arXiv]
  • PopPy: Exploiting Parallelism in Python Compound AI Applications — A runtime that auto-parallelizes chained LLM/tool pipelines written in plain Python, targeting latency in multi-step agent apps. [arXiv]
  • Public agent-trace repository push — Mario Zechner published his coding-agent sessions and called for a free, shared repository of real-world agent traces so smaller labs aren’t locked out of training data. [x.com]

2026-W19


2026-W18


2026-W17


2026-W16


2026-W15


2026-W14


2026-W13


2026-W12


2026-W11


2026-W10


2026-W09


2026-W08


2026-W07


2026-W06


2026-W05


2026-W04


2026-W03


2026-W02

  • NVIDIA Cosmos Reason 2 — Open reasoning VLM for embodied agents; top of Physical AI Bench; 2B/8B. [NVIDIA / Hugging Face]
  • NVIDIA Alpamayo — ‘First thinking, reasoning autonomous-vehicle AI’ trained end-to-end camera-to-actuation. [NVIDIA]
  • DGX Spark + Reachy Mini — Embodied-agent reference stack. [NVIDIA / Hugging Face]

2026-W01

  • Claude with FreeTaxUSA — Consumer agentic-browser + tax-form workflow on the base Claude product; 109 upvotes. [r/ClaudeAI]

2025-W52

  • MiniMax M2.1 VIBE benchmark — Introduces Visual & Interactive Benchmark for Execution — checks whether generated apps actually run correctly. [MiniMax]
  • AprielGuard — Guardrail model for safety and adversarial robustness in modern LLM systems. [ServiceNow-AI / Hugging Face]

2025-W51


2025-W50


2025-W49


2025-W48


2025-W47

  • Google Antigravity — Agent-first IDE; agents get editor/terminal/browser access. [Google Developers]
  • Apriel-H1 — ServiceNow’s distillation approach for efficient reasoning models. [Hugging Face]

2025-W46

  • SIMA 2 — Gemini-powered generalist 3D-world agent; open-ended self-improvement. [DeepMind]

2025-W45


2025-W44


2025-W43

  • ChatGPT Atlas Agent Mode — AI browser whose Agent Mode chains actions across multiple websites to complete a high-level goal. [OpenAI]

2025-W42

  • Claude Skills — Composable, persistent AI capability framework; changes the developer model from ephemeral MCP connections to reusable structured toolkits; 816 HN points. [Anthropic]
  • Gemma model helps discover cancer therapy pathway — Concrete AI-in-science contribution; Gemma model identifies a new potential therapeutic target; 225 HN points. [Google]

2025-W41


2025-W39


2025-W37


2025-W36


2025-W35


2025-W34


2025-W33


2025-W32


2025-W31


2025-W30


2025-W29


2025-W28


2025-W27


2025-W26


2025-W25


2025-W24


2025-W23


2025-W22


2025-W21


2025-W20


2025-W19


2025-W18


2025-W17


2025-W16


2025-W15

  • Google embraces MCP — Google commits to MCP, effectively making it the industry standard; 268 HN points. [HN]
  • Browser MCP — MCP server enabling browser automation from AI agents; 616 HN points. [HN]

2025-W14


2025-W13


2025-W12


2025-W11


2025-W10

  • Manus AI agent launches — Butterfly Effect’s general-purpose autonomous agent scored 86.5% on GAIA benchmark at launch (March 6); one of the highest published agent benchmark scores to date. [Wikipedia]
  • Salesforce Agentforce 2dx — Salesforce pushes autonomous agents deeper into enterprise workflows with proactive, background execution; significant enterprise agent platform move. [VentureBeat]
  • Amazon Alexa+ with enhanced AI — Amazon’s biggest Alexa upgrade since launch adds agentic capabilities including task chaining and third-party integrations. [ZDNet]

Baseline (through 2025-Q1)

Model Context Protocol (MCP)

  • MCP Launch (Nov 2024) — Anthropic released Model Context Protocol as an open standard for connecting LLMs to external tools and data sources. USB-C analogy for AI integrations. JSON-RPC based, supports resources, tools, and prompts.
  • MCP Adoption — Rapidly adopted by Cursor, Zed, Sourcegraph, Replit, and others. Became the de facto standard for LLM tool integrations. Community-built servers for databases, APIs, file systems, and more.
  • MCP Specification — Open-source spec on GitHub. Supports stdio and SSE transports. TypeScript and Python SDKs provided.

Computer Use and Browser Automation

Agent Frameworks

  • LangGraph (LangChain) — Graph-based framework for building stateful, multi-step agents. Supports cycles, persistence, and human-in-the-loop. Became the most popular agent framework.
  • OpenAI Agents SDK (Mar 2025) — OpenAI’s production agent framework (successor to Swarm). Features handoffs between agents, guardrails, and tracing. Python-first.
  • CrewAI — Multi-agent orchestration framework. Agents with roles, goals, and backstories collaborate on tasks. Popular for workflow automation.
  • AutoGen (Microsoft) — Multi-agent conversation framework. Agents communicate via messages. Strong in research and enterprise contexts.
  • Smolagents (Hugging Face) — Lightweight agent framework focused on code-based tool calling (agents write Python to use tools rather than JSON function calls).

Tool Use and Function Calling

  • Structured Outputs — OpenAI, Anthropic, and Google all improved function calling with guaranteed JSON schema adherence. Reduced tool-use failures significantly.
  • Parallel tool calling — Most frontier models now support calling multiple tools in a single turn, enabling more efficient agent loops.
  • Tool use benchmarks — Berkeley Function Calling Leaderboard (BFCL) became the standard for evaluating tool use. Claude 3.5 Sonnet and GPT-4o consistently topped it.
  • Agent reliability remains the core challenge. Success rates on complex multi-step tasks are still 30-60% for most benchmarks. Reliability improves with better prompting, retries, and human-in-the-loop.
  • “Agentic coding” emerged as the highest-value agent use case. Coding agents have clearer success criteria (tests pass) than general-purpose agents.
  • Enterprise agent adoption is cautious. Most production deployments use tightly scoped agents with human approval gates, not fully autonomous systems.