The short version
AI agents evolved from fragile language-model loops into systems that can inspect files, use tools, run tests, and return work for human review. Better models mattered, but the surrounding harness made their output more useful.
- ReAct supplied a useful reason-act-observe pattern in 2022, but early agents lacked reliable tools, strong models, and ways to check their work.
- The useful coding-agent stack combines models with tool interfaces, sandboxes, tests, diffs, checkpoints, and human approval.
- Use agents first where mistakes are visible and recoverable, and measure results on real work rather than relying on benchmark gains alone.
You can now describe a software task in plain English, leave a coding agent alone for a while, and return to changed files, a running app, or a pull request. Sometimes it really is that good. Three years ago, an agent might have spent the same time inventing subtasks and burning through tokens.
So what actually changed? I have wanted to put the whole timeline in one place, because the answer is more useful than “the models got better.” This story runs from the ReAct paper and ChatGPT through BabyAGI, AutoGPT, Devin, Claude Code, Codex, Cowork, and OpenClaw. It also includes the results that should make us cautious about declaring victory.
I am Louis-François, CTO and co-founder at Towards AI. Here is how I understand the shift, and where I think agents are useful now.
The loop existed before ChatGPT
The ReAct paper was submitted in October 2022, a month before ChatGPT launched. Its central idea was to interleave reasoning with actions and observations. An agent could form a next step, use an external source or tool, see what happened, and update its plan. That pattern is still recognizable in today’s products.
But “agent” meant a very different thing in 2023. Often it was a chat loop with a goal pasted on top. A modern coding agent can work inside an isolated environment, read and edit a repository, call tools, run tests, show a diff, checkpoint its progress, and hand a change back for approval. The loop survived. The surrounding system grew up.

2023: the agent boom ran ahead of the tools
GPT-4 gave developers a much stronger model in March 2023. ChatGPT plugins then made the shift from “what can the model say?” to “what can a system connected to the model do?” feel tangible.
A few days later came BabyAGI and Auto-GPT. The promise was exciting: give a loop a goal, let it decompose the work, and maybe it would keep going until it finished. In practice, early agents could lose the goal, hallucinate a subtask, repeat an action, or spend a surprising amount on API calls. The demo looked like autonomy; the finished work often did not.
The less flashy progress was plumbing. OpenAI’s function calling API made it easier to specify tools and receive structured arguments. It did not solve reasoning or reliability, but it removed one source of brittle prompt parsing. Later, Structured Outputs tightened schema adherence for supported JSON schemas. These are incremental engineering improvements, and they matter when an agent has to perform hundreds of actions.
The field also started testing end-to-end completion. In WebArena, the best GPT-4-based agent in the original paper completed 14.41% of the tasks, compared with 78.24% for humans. SWE-bench used real GitHub issues and tests to assess whether models could produce working patches. Its early results were similarly humbling. These benchmarks did not make agents reliable; they gave us a clearer view of how far there was to go.
Microsoft’s AutoGen explored multiple agents with specialized roles and conversations. A planner, coder, critic, and tester can be a useful design. But more agents did not automatically repair weak models or unreliable tools. Coordination is another thing that can fail.
2024: a real environment begins to take shape
LangGraph made cycles, state, and human checkpoints easier to build into agent workflows. Devin made the coding-agent idea concrete for a wider audience with a sandbox, editor, shell, and browser. The launch attracted attention because an agent could now operate in something closer to a developer’s working environment. The distance between a selected benchmark task and dependable day-to-day delivery remained large.
Computer use exposed that distance even more clearly. Anthropic’s October 2024 computer-use release let Claude interact with a desktop through screenshots and actions, while calling it experimental and error-prone. A misplaced click or a misread field is trivial in a demo and costly in a live workflow.
Then Anthropic introduced the Model Context Protocol, or MCP, in November 2024. It offered a common way to connect assistants to data and tools. It was not magic interoperability overnight, but it reduced the amount of custom integration each agent product had to invent. By the end of the year, the components were more recognizable: a model, a tool interface, a persistent workflow, and places for a human to intervene.

2025: experiments become products
OpenAI released Operator in January 2025 as a browser-using research preview. The company was explicit about its limits and the need to hand sensitive steps back to a person. In February, deep research showed a more useful shape for an agent: spend time finding and synthesizing information, then give the reader sources to inspect. Search and citation make the work easier to check than a long chain of clicks across live accounts.
That same month, Anthropic introduced Claude 3.7 Sonnet and Claude Code in research preview. The terminal was a natural place for an agent. Code lives in files. Tests can pass or fail. A diff shows what changed. If the agent goes wrong, a developer can reject or revert the patch.
In May, OpenAI introduced Codex as a cloud software-engineering agent. Its jobs ran in sandboxes and could return changes for review. Codex CLI is the separate local terminal interface, so the distinction matters when comparing products. Anthropic also made Claude Code broadly available alongside Claude 4. This is the point where “ask an agent for a change, then review it” started to feel like a real engineering workflow.

Then came an important reality check. In a randomized METR study, 16 experienced open-source developers worked on 246 tasks in repositories they knew. With the early-2025 tools in that study, allowing AI made them 19% slower. Beforehand they expected to be 24% faster; afterward they still believed they had been 20% faster. That result is scoped to those developers, tasks, and tools. It does not prove that every current agent slows everyone down. It does show how badly demos, benchmarks, and even our own feelings can mismeasure productivity.

In the second half of 2025, ChatGPT agent, Claude for Chrome, improved coding tools, skills, and more mature harnesses extended the product surface. Risks extended with it. Browsers, inboxes, and external documents give an agent useful context, but they also give untrusted content a route into the agent’s next decision.
The public mood was changing too. In an October 2025 interview, Andrej Karpathy described coding agents as unhelpful on his own project and argued that the field might need a decade. By February 2026, he was describing a very different daily workflow, assigning coding tasks and reviewing the result. One person’s change of mind is not a controlled study, but the reversal captures how quickly the tools felt different to experienced users.
Late 2025 added another layer to the agent stack. Anthropic brought reusable instructions into its agent workflow through skills, MCP moved into a broader open governance effort, and OpenAI released GPT-5.2-Codex for longer coding tasks. I read those as parts of the same shift: better models inside more mature environments, rather than a brand-new agent loop.
2026: coding and computer use converge
Anthropic’s Cowork brought the agent pattern to more general computer work. OpenClaw showed how much demand exists for a local, messaging-connected agent. The appeal is obvious: an agent that can use your own machine and services can do genuinely useful work. The security question is just as obvious. Every message, webpage, and document it reads can be untrusted input, while its available credentials may be powerful. Its project introduction is a starting point for the design, but I would not treat viral growth as proof of safety or utility.

Meanwhile, coding agents became more capable and more integrated with computer use. Cursor described cloud agents controlling their own computers, and OpenAI’s Codex expansion connected coding work to desktop interaction and broader tools. Those launches do not mean agents can be left alone with every task. They mean the product boundary between code, browser, files, and desktop is getting thinner.
The agent architecture is not a single model thinking harder. It is a model in a system that can act, observe, recover, and show its work. Better models improve every step; good boundaries stop a bad step from becoming a costly one.
Where agents work best now
Here is the rule I would actually use: give agents work where failure is cheap, visible, and reversible. A code change with tests and a reviewed pull request often fits. A research brief with sources to inspect may fit. A broad instruction to operate inside your email, financial accounts, or production systems requires far stronger controls.
That is why I keep returning to coding agents in this history. They have a relatively good feedback loop: tests fail, diffs are visible, and branches can be discarded. None of that makes code review optional. It makes useful autonomy easier to build and measure.
The original ReAct pattern is still there. What changed from 2022 to 2026 is the quality of the model, the quality of its tools, and our ability to inspect and contain the result. The next step should be measured by how much real work survives that inspection, not by how convincingly an agent appears to be busy.
What agent release or failure changed your own view? I would love to hear which part of this timeline you think mattered most.
FAQ
What is an AI agent?
Here, an agent is a system that uses a model to choose actions, interact with tools or an environment, inspect the result, and repeat toward a goal. Modern agents also need boundaries, checks, and human review.
Why did early agents such as AutoGPT struggle?
Early loops had weaker models, brittle tool connections, high costs, limited isolation, and few reliable ways to verify progress. They could drift or repeat actions without completing the user's goal.
What changed for coding agents in 2025 and 2026?
Stronger coding models arrived inside better harnesses with repository access, sandboxes, tests, diffs, checkpoints, and reviewable pull requests.
Did MCP make agents reliable?
MCP gave tools and data sources a common interface, reducing custom integration work. It did not verify an agent's decisions; permissions, tests, and human review still matter.
Do benchmark gains prove that agents make developers faster?
No. Benchmarks measure defined tasks and conditions. In a 2025 METR trial, 16 experienced open-source developers were 19% slower on 246 tasks with early-2025 AI tools, despite expecting and perceiving speedups.
Where should teams use agents first?
Start with bounded work where failures are visible, inexpensive to undo, and easy to test. Keep approval gates before costly or irreversible actions.

