Click to Play Episode
What an AI agent actually is, why coding agents got good first, how memory really works, what MCP and A2A standardize, which SDKs are alive, how to evaluate on trajectories, and where the products stand after browser agents contracted.
First of two episodes on AI agents. This one is the architecture: the loop, tools and verifiable feedback, memory, the protocols (MCP, A2A, computer use), the SDK landscape, evaluation and observability, the product map, and when multiple agents help. The next episode, OpenClaw and the Personal Agent, applies it to one always-on assistant with security as the centerpiece. Coding-agent products and mechanics live in the vibe coding sequence starting at MLA 22.
A chat model returns a message; a workflow is your code calling a model at fixed steps; an agent is a model that owns the control flow, choosing its next action from what it observes. That puts systems on a spectrum (chat, chat plus tools, workflows, agents) rather than in a binary, the framing Anthropic's Building Effective Agents uses. The loop itself is ReAct (Yao et al.): thought, action, observation, repeat, with the reasoning trace letting the model track and update a plan. What changed by 2026 is not the loop but the infrastructure around it, and every part of that infrastructure is an attack on per-step error compounding.
Function calling: you describe tools as schemas, the model emits a structured call, your code executes it and returns the observation. The model never runs anything itself, which is the security model. Writing effective tools for agents gives the practical rules: few high-impact tools, clear namespaces, meaningful identifiers, token-efficient responses, descriptions treated as prompt engineering. Effective context engineering for AI agents adds the overlap test: if a human cannot say which tool applies, neither can the agent. The central principle: coding agents got good first because tests and compilers give verifiable feedback that catches a bad step inside the same loop that made it. Find or manufacture the verifier before writing the prompt.
"Memory" means four things: the context window (the only memory the model has), retrieval from an external store, files on disk, and episodic records of prior sessions. Most agent memory is files. Anthropic's memory tool is a client-side file protocol (view, create, replace, insert, delete) against storage you own. The hard part is context management, and both labs converged on the same three mechanisms: context editing to clear stale tool results, compaction to summarize near the limit, and notes written to files before summarization. OpenAI's Responses API conversation state has the same shape with a compaction threshold and compact endpoint. Third-party layers Mem0, Letta (from MemGPT), and Zep (temporal knowledge graph) now compete with first-party primitives. Multi-session patterns: Effective harnesses for long-running agents.
Model Context Protocol is the agent-to-tool standard, now a Linux Foundation project with individual-maintainer governance. 2026 additions: elicitation (server asks the user mid-operation), an extensions mechanism, and the async Tasks extension for long-running tools. Every major SDK below consumes it; its cost is the context each connected server's tool list occupies. A2A is the agent-to-agent standard, Google-built, Linux Foundation-hosted, at v1.0 with a steering committee spanning AWS, Cisco, Google, IBM, Microsoft, Salesforce, SAP and ServiceNow. Strong governance, weak observed consumption; worth knowing, not yet worth building on for small teams. Computer use is the universal fallback: Anthropic's computer use tool (GA toolset with zoom and an automatic injection classifier), Google's Gemini computer use, open-source Browser Use, and Playwright MCP, which drives the accessibility tree instead of screenshots. Prefer API, then accessibility tree, then screenshots.
Both labs advise starting without a framework: Building Effective Agents and OpenAI's A Practical Guide to Building Agents. The 2026 SDKs have converged on that critique as thin harnesses around a loop.
create_agent as a minimal middleware harness. LangSmith is the separate tracing product.Decision rule: machine-operating agent fast, Claude Agent SDK; lightweight handoffs, OpenAI; durable human-in-the-loop state, LangGraph; inside Google or Microsoft, their kit; to understand what you run, write the loop yourself first.
Agents are evaluated on trajectories, not answers: traces, task evals, cost per task. Traces follow the OpenTelemetry GenAI semantic conventions; products include LangSmith, Langfuse (open source, acquired by ClickHouse), Arize Phoenix, Braintrust, W&B Weave, and Helicone. Public benchmarks show the shape of a task eval: SWE-bench Verified, which OpenAI stopped reporting citing contamination; SWE-bench Pro; tau2-bench; Terminal-Bench 2.0; OSWorld-Verified; GDPval. Cost and reliability: Princeton's Holistic Agent Leaderboard (paper) and its reliability dashboard separate pass@k capability from pass^k reliability; METR time horizons with their own limitations note. Guardrails: OpenAI agent safety, NeMo Guardrails, Guardrails AI; prompt injection framed by Simon Willison's lethal trifecta and Google's CaMeL architectural defense.
Claude Cowork: "Claude Code for everyone," a sandboxed desktop agent with open-sourced plugins. ChatGPT agent remains; the Atlas browser was retired within a year, folded into ChatGPT and Codex. Google discontinued Project Mariner and moved the capability into Gemini and Antigravity, which absorbed Gemini CLI. Standalone browser agents contracted; the capability moved into models and existing apps. Still shipping: Perplexity Comet (free), Manus (ownership contested this year; check before building on it), Devin, Copilot Studio, Agentforce. Glue: n8n (AI Agent node inside a drawn workflow, MCP server trigger) and Zapier Agents with Zapier MCP. Browser-agent prompt injection is the documented security problem: the PleaseFix research note.
Two essays a day apart: How we built our multi-agent research system (orchestrator plus parallel subagents beat a single agent on research at roughly 15x the tokens; token usage explained most of the variance) and Cognition's Don't Build Multi-Agents (dispersed decisions and unshared context make it fragile). The disagreement is task shape. Parallelize independent, read-mostly work; keep stateful, sequential work single-threaded; prefer a small hierarchy where workers return findings rather than decisions.
Companion show: Agentic Business on Gnothi follows one business as agents take on research, software, sales and operations.
Before we get into it, this episode came out of Gnothi. It's a tool I built that researches a topic, writes it up as chapters, and narrates them. What you'll hear here is the short version. The dedicated Gnothi show, Agentic Business, follows one business as agents take on research, software, sales, and recurring operations. O C devel dot com slash agents. That's O C D E V E L dot com, slash agents.
This is the first of two episodes on AI agents. This one is the architecture: what an agent actually is, why coding agents got good before everything else, how memory really works, what the protocols standardize, how you build one with today's SDKs, how you evaluate it, and where the products stand. The next episode, OpenClaw and the Personal Agent, takes all of that and applies it to one worked example, an always-on assistant that lives in your messaging app, with security as the centerpiece. If you came here for coding agents specifically, the vibe coding sequence that starts with Vibe Coding in twenty twenty-six covers the products, the mechanics, and the engineering practice, and I'll point at those episodes rather than repeat them.
Let me start with a definition, because the word agent has been stretched to cover everything from a chatbot with a search button to a fleet of processes rewriting a codebase overnight. I want a definition that's precise enough to be useful and loose enough to survive the next product launch.
Here is the version I use. A chat model takes a message and returns a message. A workflow is code you wrote that calls a model at certain steps, and the code decides what happens next. An agent is a model that decides what happens next. The model gets a goal, it looks at the situation, it picks an action, it sees the result of that action, and it picks the next one. The loop runs until the model decides it's done or something stops it. The difference from a workflow is who owns the control flow. In a workflow, you do. In an agent, the model does.
That definition puts things on a spectrum rather than in a box, and the spectrum is the useful part. On one end is plain chat, no tools, no autonomy. One step over is chat with tools, where the human drives and the model occasionally calls a search or runs a snippet. Then workflows, where you've designed a fixed path and the model supplies intelligence at specific nodes. Then agents, where the model chooses its own path and its own tools based on what it observes. Anthropic's essay Building Effective Agents draws this line the same way, and it's the framing I find survives contact with real products. Most things sold as agents are somewhere in the middle, and knowing where a given product sits tells you more than its marketing does.
The loop itself has a name and a paper. In twenty twenty-two, Yao and colleagues published ReAct, which stands for reasoning plus acting. Before that paper, the research community treated reasoning, meaning chain of thought, and acting, meaning tool use, as separate tricks. ReAct interleaved them. The model writes a thought about what it needs, takes an action, gets an observation back, and writes the next thought informed by that observation. Thought, action, observation, repeat. Every agent you'll touch this year is running some elaboration of that cycle. Claude Code does it. A LangGraph graph does it inside its nodes. The AI agent node in n eight n does it inside one box on a canvas.
ReAct showed that the reasoning traces aren't decoration. They let the model track and update a plan and handle exceptions, while the actions let it gather information it didn't already have. Either alone is weaker. This matters in twenty twenty-six because the loop hasn't changed much, but everything around it has. The models got much better at deciding, the tools got standardized, the context got managed, and the harnesses got engineered. So when someone asks what is new about agents, the answer is mostly not the loop. It's the infrastructure around the loop, which is the rest of this episode.
One more framing before we move on, because it will come up in every section. A loop that decides its own next step can also decide wrong, and errors in a loop compound. If each step is right ninety-five percent of the time, a twenty-step task succeeds a bit over a third of the time. Everything good in agent engineering is, in some form, an attack on that compounding: better models per step, feedback that catches a bad step before the next one, memory that keeps the model from re-deciding what it already knew, and evals that measure the whole trajectory rather than the individual answers.
So that's the loop. Now what the loop can touch, which is what makes it useful.
An agent without tools is just a model talking to itself. Tools are how the loop reaches the world, and the mechanism is function calling. You describe a function to the model in a schema: its name, what it does, what arguments it takes. The model, when it decides that function is the right next action, emits a structured call instead of prose. Your code runs the function and hands the result back as the observation. The model never executes anything itself. It requests, you execute, you report. That indirection is the whole security model of agents, and it's also why the quality of your tool descriptions matters as much as the quality of your prompt.
Anthropic wrote a piece last year called Writing Effective Tools for Agents that I think of as the practical manual here. The principles are short. Build a few high-impact tools, not a wrapper around every API endpoint. Give tools clear namespaces so a model with forty tools can tell them apart. Return meaningful identifiers rather than opaque UUIDs. Make responses token efficient, paginated and filtered, because every byte a tool returns lands in the context window. And treat the tool description as prompt engineering, because that's literally what it is. Their context engineering post adds a test I use constantly: if a human engineer can't say definitively which of two tools should be used in a situation, an agent can't be expected to do better. Overlapping tools are a bug.
Now the principle that explains the most. Coding agents got good two or three years before every other kind of agent, and the reason isn't that code is easy. It's that code gives verifiable feedback. When an agent edits a file and runs the tests, it gets an unambiguous observation: pass or fail, with a specific error message. The compiler says exactly which line is wrong. The type checker says exactly which type is wrong. That signal feeds straight back into the next thought, so a bad step gets caught in the same loop that made it, before it compounds. Compare a browser agent asked to book a hotel. How does it know it picked the right one? The feedback is a page of HTML and a vague sense that something happened. The loop is the same. The observation is worthless.
This gives you a design rule that transfers to any domain. When you build an agent, your first job isn't the prompt. It's to find or build the verifier. If the task has a test suite, a schema validator, a linter, a numerical check, a diff against a known-good output, anything deterministic that says yes or no, wire that in as a tool and make the agent use it. Domains where a verifier exists will get agents that work. Domains where success is a matter of taste will get agents that need a human in the loop, and no amount of prompting changes that. This is also, I think, the plain explanation for why the products that actually work in twenty twenty-six are still mostly coding products, and why the general-purpose assistants keep getting relaunched.
There's a second-order version of the principle too. Sometimes you can manufacture a verifier where none exists. Ask the agent to write the acceptance test before it does the work. Ask it to produce a structured output that a schema can check. Ask a second model to grade the output against a rubric. None of these are as good as a real test suite, but each one turns a fuzzy observation into a sharper one, and sharpening the observation is where most of the gain is.
That's tools. Next, memory, a word that gets used far more often than it gets defined.
When people say an agent has memory, they usually mean four different things, and it helps to separate them. The first is the context window, which is the only memory the model actually has. Everything the model knows about the current task is either in its weights or in the tokens currently in front of it. Nothing else exists. The second is retrieval, where you keep a store outside the model, embeddings or a database or a search index, and pull relevant pieces into the context when they are needed. The third is files, plain text the agent reads and writes on disk. The fourth is what researchers call episodic memory, a record of what happened in earlier sessions, which in practice is just some combination of the other three.
Most agent memory is files, which isn't how the memory products lead. Claude Code's project instructions are a markdown file. The memory folder it keeps is markdown files. OpenClaw's soul and memory documents, which the next episode covers in depth, are markdown files. When Anthropic shipped a memory tool in their API, the design is that the model issues commands like view, create, replace, insert, and delete against a directory, and your application executes them against storage you own. The model reads its notes at the start, writes notes as it goes, and the docs literally instruct it to assume it may be interrupted at any time. That's a file system with discipline. It's not a new kind of cognition, and I mean that as a compliment, because files are inspectable, versionable, and editable by a human when the agent gets something wrong.
Storage isn't what makes memory hard. The context window is, because it's finite and expensive. A long task fills it with tool results, most of which are stale by the time the task ends. So the real memory engineering is context management, and both major labs converged on the same three mechanisms within a year of each other. One, editing: automatically clearing old tool results out of the context when it gets large, keeping the last few. Two, compaction: when the conversation approaches the limit, summarize the whole thing into a shorter version and continue from that. Three, notes: because compaction is lossy, the agent writes durable facts to files before they get summarized away. Anthropic exposes these as context editing, compaction, and the memory tool. OpenAI's Responses API has conversation state, a compaction threshold, and a compact endpoint. Same shape, different names. If you read Anthropic's post on effective context engineering, the thesis is that you want the smallest set of high-signal tokens that maximize the chance of the outcome, and every one of these mechanisms is a way to get stale tokens out.
Then there are the third-party memory layers: Mem zero, Letta, which grew out of the MemGPT paper, and Zep, which builds memory as a temporal knowledge graph rather than flat retrieval. As of this recording they are all maintained, and they are all now competing with a first-party primitive from both major labs, which is the actual story. My advice is to start with files and the built-in mechanisms, and reach for a graph or a vector store only when you can name the query that files can't answer. A lot of memory infrastructure gets bought to solve a problem that a well-named markdown file would have solved.
One thing to carry into the next episode. The mechanics of memory in coding agents, meaning the instruction files, compaction, and auto memory inside a tool like Claude Code, are covered in Inside a Coding Agent. What I have given you here is the architecture those mechanics implement.
So that's memory. Now the protocols, where the plumbing between vendors is being decided.
The Model Context Protocol, MCP, is the standard for connecting agents to tools. Anthropic published it in late twenty twenty-four, and it has since moved into the Linux Foundation, with a governance structure of individual maintainers rather than company seats. The idea is a client-server contract. A server exposes tools, resources, and prompts. A client, meaning an agent harness, connects to any server and gets those capabilities without custom integration code. Every major agent framework I'll name later supports it, from Anthropic's own to OpenAI's to Google's to Microsoft's. If you write a tool once as an MCP server, it works in all of them. That's the value, and it's what turned a vendor project into an industry standard.
The spec kept moving through twenty twenty-six. The additions worth knowing are elicitation, where a server can ask the user for input mid-operation, so a tool can say I need you to confirm this before I proceed. An extensions mechanism, so optional features live outside the core. And an async tasks extension, where a long-running tool returns a task identifier instead of blocking, and the agent polls for completion or gets notified. That last one matters because real work takes minutes, not seconds, and a blocking tool call was one of the sharpest edges in the original design. There's also a working group on skills over MCP, which points at where this is going: not just tools but packaged procedures that an agent loads on demand.
One caution. MCP has a context cost. Every server you connect advertises its tools into the model's context, and a dozen servers with twenty tools each is a lot of tokens spent before the agent has done anything. Inside a Coding Agent covers how to manage that inside a harness. Here I'll just say that the protocol standardizes the plumbing, not the discipline. You still have to choose few, well-described tools.
The second protocol is A two A, agent to agent, which Google published and has since donated to the Linux Foundation, where it reached version one point oh. The pitch is complementary to MCP: MCP connects an agent to tools, A two A connects an agent to other agents. An agent publishes a card describing what it can do, another agent discovers it and sends it a task, and they exchange messages and artifacts over a standard transport. The steering committee includes AWS, Cisco, Google, IBM, Microsoft, Salesforce, SAP, and ServiceNow, and there are SDKs in six languages. So it's a real standard with real weight behind it.
My read, and I'll label it as opinion. The governance signal is strong and the consumption signal is weak. Every framework that says it supports A two A supports it as an optional transport, and as of this recording I couldn't find evidence of anyone depending on cross-vendor agent discovery in production. Anthropic and OpenAI aren't on that committee, and both put their interop energy elsewhere, Anthropic into MCP and OpenAI into hosted orchestration. For an individual or a small team, A two A is a standard to know about, not one to build on yet. MCP is the one that pays rent today. Enterprises with agents from three vendors that need to talk to each other are the audience for A two A, and if you are one of those, it's worth watching.
The third area isn't a protocol but a capability that behaves like one: computer use and browser agents. Computer use means the model takes screenshots, reasons about what it sees, and emits mouse and keyboard actions. Anthropic's version is now a general availability toolset with a couple dozen actions, including a zoom action for inspecting a region at full resolution, and it runs an automatic prompt injection classifier on what it sees. Google ships computer-use models in the Gemini API. On the open source side, Browser Use is an actively maintained library for driving a browser from a model, and Playwright MCP exposes a browser through the accessibility tree rather than screenshots, which is often the cheaper choice because a tree of labeled elements is far fewer tokens than an image.
The reason to bring computer use into a protocols section is that it's the universal fallback tool. Anything without an API becomes usable through a screen. It's also the slowest, most expensive, and least verifiable way to do anything, for exactly the reason I gave earlier: the observation is a picture. When you have a choice between an API, an accessibility tree, and a screenshot, take them in that order. And in the products section I'll come back to why browser agents specifically have had a rough year on security.
That's the protocol layer. Now, how you actually build one.
Before I name a single SDK, here is the advice from both major labs, and it surprised people when it landed. Anthropic's Building Effective Agents says the most successful implementations they saw were not using complex frameworks, and it recommends starting with direct API calls and composable patterns, adding a framework only when it earns its place, because frameworks obscure the prompts and make debugging harder. OpenAI's practical guide to building agents is gentler but says the same thing: start with a single agent and a minimal loop, and split into multiple agents only when complexity forces you. Both of those companies now sell the frameworks their own guidance tells you to defer. That tension is real, and the way to resolve it's to notice that the twenty twenty-six frameworks have converged on the critique. They are now thin harnesses around a loop you could have written yourself, and their value is a loop that has already been run hard in production, not one that has been abstracted away.
With that said, here are the options as of this recording, sorted by how much control you want.
The Claude Agent SDK, in TypeScript and Python, is the harness-as-SDK option. It's literally Claude Code's agent loop exposed as a library. You get the built-in file, shell, and web tools, subagents, lifecycle hooks at tool events, MCP servers, permission modes, and sessions with automatic compaction. If you want a coding-shaped agent that operates on a machine and you want to skip writing any of that infrastructure, this is the shortest path. The cost is that you are buying Anthropic's opinions about the loop, and it's still on a pre one point oh release line with frequent updates.
The OpenAI Agents SDK, also Python and TypeScript, is built around handoffs. You define small agents with tools and guardrails, and control passes from one agent to another inside a single run. It has sessions and tracing built in and rides on the Responses API. The notable twenty twenty-six change is that OpenAI deprecated its visual no-code Agent Builder and pointed people at this SDK instead, which tells you something about where they think serious agent work happens.
LangGraph is the graph option. You define state, nodes, and edges, and the runtime gives you checkpointing, interrupts for human approval, and durable execution across failures. LangChain rebuilt its one point oh line on top of it with a deliberately minimal create agent function, composed from model, tools, prompt, and middleware. If your agent needs to pause for a human, resume a day later, and survive a crash in the middle, the graph is the right abstraction. LangSmith is the separate commercial tracing and eval product; you don't need it to run LangGraph.
Google's Agent Development Kit is code-first hierarchical composition, with explicit agent trees, sequential and parallel and loop workflow agents, and native A two A so a subagent can be a remote agent. It deploys to a managed runtime on Vertex, but it doesn't require it. Microsoft's Agent Framework is the merger of AutoGen and Semantic Kernel, now generally available with an LTS commitment, and it's what you use if your organization runs on dot net and Azure. AutoGen itself is in maintenance mode, so if you picked it two years ago, you have a migration.
CrewAI is role-based: you declare agents by role, goal, and backstory and compose them into a crew with a sequential or hierarchical process. It's past one point oh and has a commercial management platform on top. Hugging Face's smolagents is the code-agent idea, where instead of emitting JSON tool calls the model writes Python that calls your tools, run in a sandbox, in about a thousand lines of framework. It's the one on this list with the weakest maintenance signal as of this recording, fine for research and prototypes. Two more worth a sentence: Pydantic AI, for validation-first typed agents in Python, and Vercel's AI SDK, whose version seven turned its loop control into a real agent abstraction for TypeScript apps.
The decision rule. If you want a machine-operating agent fast, Claude Agent SDK. If you want lightweight multi-agent with handoffs, OpenAI Agents SDK. If you need durable, resumable, human-in-the-loop state, LangGraph. If you are inside Google Cloud or Microsoft, take their kit. If you want to understand what you are running, write the loop yourself first. It's a while loop with a function-calling model inside it, and I promise the first version is under a hundred lines.
That's building. Now the part that separates a demo from a product: evaluation and observability.
An agent that works in the demo and fails in production is the default outcome, and the reason is that agents are evaluated on trajectories, not answers. A chat model gives one answer and you grade it. An agent takes thirty steps, and the question is whether the whole path arrived, how much it cost, and whether it did anything it shouldn't have along the way. That means three things you need: traces, task evals, and a cost per task.
Traces first. A trace is the full record of one run: every model call, every tool call, every tool result, with timing and tokens. You can't debug an agent without one, because the failure is usually a bad decision on step eleven caused by a misleading observation on step nine, and you won't find that from the final output. The OpenTelemetry project has generative AI semantic conventions for this, with client spans stable and agent spans still experimental as of this recording. Products in the space include LangSmith, Langfuse, which is open source and was acquired by ClickHouse this year while keeping its license and self-hosting, Arize Phoenix, Braintrust, Weights and Biases Weave, and Helicone. They differ on whether they are eval first or observability first and whether you self-host, and any of them beats none of them.
Task evals second. Build a set of tasks with known correct outcomes, run the agent on all of them, and score the trajectory. The public benchmarks show what this looks like. SWE-bench Verified was the standard for coding agents until OpenAI publicly stopped reporting on it, citing contamination, and Scale's SWE-bench Pro replaced it with harder, multi-file tasks held out from training data. Tau bench and tau two simulate a customer and a policy document and measure whether the agent follows the rules while helping. Terminal Bench two point oh is eighty-nine hard command-line tasks graded on end state. OSWorld Verified is real desktop tasks in a VM. GDPval is real occupational deliverables graded by experts. I'm not going to read scores, because scores are stale by the time a recording ships and half the numbers floating around are on aggregator sites naming models that don't exist. What I want you to take from the list is the shape: each one is a set of tasks with a verifier, which is the same verifiable-feedback principle from earlier turned into a measuring instrument. Your private eval should look like a tiny one of these for your domain.
Cost per task third. Princeton's Holistic Agent Leaderboard is the cleanest example of taking this seriously. It runs the same models and scaffolds across many benchmarks and maps accuracy against dollars, and the finding that stuck with me is how often the top of the accuracy chart costs orders of magnitude more than a point just below it. Their reliability dashboard makes a related point: the gap between whether a model can do a task at least once in k tries and whether it does it every one of k tries is large across every model. Capability and reliability are different numbers, and production cares about the second one. METR's time-horizon work is the other reference here, measuring the length of task an agent completes half the time, with a doubling period of a few months; their own follow-up note warns that doubling the horizon doesn't double what you can automate, which is the right caveat.
Guardrails last. There are two layers. Deterministic checks on inputs and outputs, which the OpenAI SDK ships as guardrails and NVIDIA's NeMo Guardrails and Guardrails AI offer as libraries. And architectural defense against prompt injection, where the threat is that an agent reading untrusted content, a web page or an email, gets instructions from it. Simon Willison's lethal trifecta names the condition: an agent with access to private data, exposure to untrusted content, and a way to communicate externally can be made to exfiltrate. Remove any one leg and the attack fails. Google's CaMeL paper proposes the strongest architectural fix, a policy engine outside the model that authorizes actions, and as of this recording nobody ships it in a mainstream harness. So in practice, guardrails means: minimize what the agent can reach, verify what it produces, and log everything. The next episode spends its longest section on exactly this, because a personal agent that reads your inbox is the trifecta by construction.
So that's measurement. Now the product tour, which I'm keeping to one section because the products change faster than the ideas.
Anthropic's Claude Cowork is the pitch of Claude Code for everyone, a desktop agent that works on files and documents in a sandbox with plugins bundling skills, connectors, subagents, and slash commands. Anthropic open-sourced a set of those plugins for sales, finance, legal, data, and so on. It's the consumer-facing shape of the same harness as the SDK. OpenAI's ChatGPT agent mode is the general assistant that browses and produces files inside ChatGPT. The interesting story there's the Atlas browser, which OpenAI launched in late twenty twenty-five and retired less than a year later, folding browser-based agent work back into the ChatGPT desktop app and Codex. Google similarly discontinued Project Mariner, its standalone browser agent, and moved the capability into the Gemini models and into Antigravity, which is now its agent platform for developers and absorbed the Gemini CLI. Read those three moves together and the pattern is clear: standalone browser agents as products didn't hold, while the capability got pushed into the model and the existing apps.
Perplexity's Comet is the counterexample, a browser with an agent in it that's still shipping and, as of this recording, free. Manus is the general agent that made a splash in early twenty twenty-five; its ownership has been the subject of a contested acquisition this year, and rather than narrate a corporate story that may change before you hear this, I'll just say the product is still operating and you should check its status before you build on it. Devin is Cognition's autonomous coding agent for ticket-sized work. Microsoft's Copilot Studio and Salesforce's Agentforce are the enterprise agent builders for organizations already inside those estates.
Then the glue layer. n eight n and Zapier are workflow tools that grew agent nodes, and I want to place them precisely on the spectrum from the start of the episode. The AI agent node in n eight n runs a real ReAct loop that picks its own tools, but the node sits inside a workflow you drew. That's agentic behavior inside a deterministic container, and for a lot of business automation it's exactly the right amount of autonomy. Both also speak MCP now, n eight n by exposing any workflow as a tool a coding agent can call, Zapier by exposing its integrations to whatever agent you run. If you already have a Zap or an n eight n flow, wrapping it as a tool for an agent is often a better move than rebuilding it as an agent.
The one product story I wouldn't skip is browser agent security. In the space of about six months, researchers demonstrated prompt injection attacks on Comet, on Atlas, on Gemini via a calendar invite, and on a couple of smaller agentic browsers, in some cases hidden in screenshots or page content the agent read. OpenAI has said prompt injection is unlikely to ever be fully solved. That's the lethal trifecta showing up in shipping products, and it's part of why the standalone browser agent category contracted. It's also the setup for the next episode.
Last section, short, because it's a judgment call: multiple agents.
Two essays landed one day apart in June twenty twenty-five arguing opposite conclusions. Anthropic's How We Built Our Multi-Agent Research System described an orchestrator with parallel subagents that outperformed a single agent by a wide margin on research tasks, at roughly fifteen times the tokens of a chat. Cognition's Don't Build Multi-Agents argued that multi-agent systems are fragile because decisions get dispersed and context isn't shared, and that a single-threaded agent with good context compression beats them. Both are right, and the disagreement is about task shape. Anthropic's case was read-only parallel research, where subagents explore independently and the only thing that has to merge is findings. Cognition's case was stateful code editing, where two agents making decisions about the same files without seeing each other's context will conflict.
So the rule I use is this. Parallelize when the subtasks are independent and read-mostly, and when each one would otherwise flood a single context with tool results that the final answer doesn't need. Search, research, reviewing many files, running many evals. Don't parallelize when the subtasks share mutable state or when each decision depends on the last, because then you've multiplied the error compounding I described at the start by the number of agents. The most common productive pattern isn't a team of peers but a small hierarchy: one agent that holds the goal and delegates bounded slices, and workers that return findings rather than making decisions. Even then, the token bill is the real cost, and the Anthropic post found that token usage alone explained most of the performance variance. Multi-agent is a way to spend more tokens on a problem. Make sure the problem is worth it.
Let me recap. An agent is a model that owns its own control flow, running a thought, action, observation loop that hasn't changed much since ReAct. Tools are how the loop reaches the world, and the tools that give verifiable feedback are why coding agents got good first, so find the verifier before you write the prompt. Memory is mostly files plus context management, and both labs converged on editing, compaction, and notes. MCP is the tool standard that pays rent today, A two A is the agent-to-agent standard to watch, and computer use is the universal but expensive fallback. Build with the SDK that matches how much control you need, after writing the loop yourself once. Evaluate on trajectories with traces, task evals with verifiers, and cost per task. And treat multiple agents as a way to spend tokens on independent work, not as a default architecture.
Next episode, OpenClaw and the Personal Agent, takes one always-on assistant apart, from its memory files and heartbeats to the security posture you need before it's allowed anywhere near your email.