orwelllab
AI agents / Practical guide

How do AI agents work? Tools, memory and reasoning explained

What actually happens inside an AI agent on each turn: who runs the tools, what its memory really is, and why long tasks fail more often than short ones.

In brief
  • The language model inside an agent never runs a tool itself: it writes a structured request, and ordinary code carries it out and hands back the result.
  • If each step of a task goes right 19 times in 20, a 10-step task finishes cleanly only about 60% of the time.
  • In Sierra's τ-bench, GPT-4o solved about 61% of retail customer tasks on one try, but only about 25% when the same task had to succeed on all 8 tries.

An AI agent works by running a language model in a loop. On each turn the model reads its instructions, the goal and everything gathered so far, then picks one next action: answer, or ask for a tool such as an order lookup, a calendar check or an email send. Ordinary software carries out that request and feeds the result back, and the loop repeats until the job is done, a step limit is reached or a person is needed. What people call memory is whatever the software chooses to send back in on the next turn, because the model itself keeps nothing between calls.

That design fact explains most of the odd behaviour owners report: agents that forget, repeat themselves, or get nine steps right and the tenth wrong. Our guide to agentic AI covers what the category is and where it pays. This page opens the bonnet.

One turn of the loop

Take a customer email asking to move a delivery to Friday. Here is what happens inside a typical agent, in order.

  1. The software assembles a request: instructions, the email, and a menu of tools, each with a name, a description and the inputs it expects.
  2. The model replies with a decision. Either it writes an answer, or it returns a tool request, for example find_order with the customer's email address filled in.
  3. Your code runs the tool. It checks the request is allowed and queries the order system.
  4. The result goes back in, added to the conversation, and the whole lot is sent to the model again.
  5. Repeat. Next might come check_delivery_slots, then draft_reply, until the model gives a final answer or a limit stops it.

Anthropic, which makes the Claude models, puts the first point bluntly in its developer documentation: "The model never executes anything on its own." Its engineers describe agents as "LLMs using tools based on environmental feedback in a loop". Both lines are a vendor describing its own product. The pattern itself, where the model asks and your code acts, is what developers call function calling.

Tools are where the permissions live

Because your code runs every tool, that code is where limits go. The model can ask for a refund. Whether one happens depends on the tool: a cap, a check against the order value, or a pause for a manager.

So split tools by consequence. Reading a record is cheap to get wrong. Drafting a reply is cheap too, because nobody has seen it yet. Sending, paying and deleting are not, and those are the tools that sit behind an approval gate. Our page on AI agents for business sets out how to scope that access role by role.

Memory is four different things

Vendors use one word for four mechanisms that fail in different ways. Ask which one a product means.

Kind of memoryWhat it actually isHow long it lastsWhat goes wrong
Context windowEverything sent to the model on this call: instructions, messages, tool resultsOne callLong inputs bury details in the middle
Task historyThe running transcript the software re-sends on every turnOne taskEach turn costs more; an early mistake is carried forward
Retrieved knowledgeDocuments or records found by a search and pasted into the contextAs long as the source existsThe search fetches an outdated or wrong document
Stored notesFacts the system writes to a database to reuse on later tasksUntil someone deletes themA wrong note keeps resurfacing; personal data rules apply

The first row has research behind it. A Stanford-led study published in 2024 found that model performance "significantly degrades when models must access relevant information in the middle of long contexts", and that this held even for models built for long inputs. A bigger context window isn't a cure. Sending less, and putting what matters near the start or end, works better.

Reasoning happens one step at a time

An agent doesn't hold a plan in its head. On each turn it predicts the most sensible next action from the text in front of it. Asking the model to write a plan first helps, but that plan is just more text in the context. When a tool returns an error, a well-built agent tries something else. A badly built one asks again.

That's why stopping rules belong in the software, not the prompt. Anthropic's own guidance recommends stopping conditions such as a maximum number of iterations. Set a cap on turns, on tool calls and on spend, and route anything that hits a cap to a named person with the transcript attached.

The reliability maths every buyer should run

Each turn is a fresh chance to go wrong, and the chances multiply. The table assumes every step succeeds independently at the same rate. Real systems don't quite, but the shape holds.

Each step rightChance a 10-step task finishes cleanlyChance a 27-step task finishes cleanly
19 in 20about 60%about 25%
99 in 100about 90%about 76%
999 in 1,000about 99%about 97%

Twenty-seven steps isn't an arbitrary column. In TheAgentCompany benchmark from Carnegie Mellon and Duke, the strongest agent took almost 27 steps per task on average and completed 30% of the tasks on its own. Our AI employees page breaks those results down by job.

Consistency is the other half. Sierra's τ-bench, a peer-reviewed benchmark of agents serving simulated customers under company policies, gave GPT-4o about 61% on retail tasks and about 35% on airline tasks. Asked to get the same retail task right on all 8 attempts, it managed about 25%. Sierra sells customer service agents, though these figures come from its peer-reviewed paper, not its marketing.

Models are improving fast. METR, an AI evaluation group, found the length of task frontier models can finish half the time has doubled about every 7 months. For o3 that horizon was software work taking a person around 110 minutes. Half the time is a coin toss, not an operating standard for invoices or customer replies.

What this means when you design one

  • Make the model decide fewer steps. Anything predictable (send the confirmation, update the status field) belongs in fixed code. The difference is laid out in our note on agents versus RPA and the agentic workflows guide.
  • Check after each action. Read the record back after writing it; a tool reporting success isn't proof.
  • Treat every tool result as untrusted text. An email or web page fed into the context can contain instructions aimed at the model. The UK's National Cyber Security Centre warned in December 2025 that prompt injection "may never be totally mitigated in the way SQL injection attacks can be", so limit what any one tool can do.
  • Run it in shadow mode first, scoring its output against staff before it sends anything.

Common questions

Do AI agents learn from every task?

Not by themselves. The model's weights don't change when it finishes a job. An agent appears to learn only if its builder saves notes, corrected examples or updated instructions and feeds them into later tasks. That stored material needs an owner, because a wrong lesson gets repeated just as faithfully as a right one.

What is a tool call?

A tool call is the model asking software to do something on its behalf. It names a tool from the menu it was given and fills in the inputs as structured data. The application runs the request, such as a database search, and sends the result back for the model to read before choosing its next move.

How do AI agents remember past conversations?

The software saves the conversation, or a summary of it, and pastes the relevant parts into the next request. The model has no memory of its own between calls. That's why a long-running agent can lose track of early details, and why stored summaries need checking like any other business record.

Why do AI agents get stuck repeating the same step?

Usually because a tool returned an error the model couldn't interpret, so it asks again with the same inputs. Clear error messages from each tool help the model try something different. A hard cap on turns and tool calls, set in the software, stops a stuck loop from running up costs or spamming a customer.

Is it safe to connect an AI agent to company systems?

It can be, if access is narrow. Give it read access first, separate drafting from sending, and require a person's approval for payments, deletions and anything a customer will see. Assume any document or email it reads could carry hidden instructions, and make sure no single tool can do serious damage alone.

Pick one process and write down every step a person takes today, including the lookups. If the list runs past ten decisions, that number tells you how much of it should be fixed code and how much needs a model. Send the list, with roughly how many cases arrive each month, through the contact form under Business automation; after an informal scoping chat we'll come back with a price range for the build.

Sources

Every figure in this article links back to the source below it was checked against.

  1. Orwell Lab calculation: chance a multi-step task completes cleanly at a fixed per-step success rate Our calculation from official rates · checked 9 October 2026
  2. Sierra: τ-bench, a benchmark for tool-agent-user interaction (ICLR 2025) Research · checked 9 October 2026
  3. Anthropic (vendor): How tool use works, Claude platform documentation Vendor estimate · checked 9 October 2026
  4. Anthropic (vendor): Building effective agents (December 2024) Vendor estimate · checked 9 October 2026
  5. Stanford University and others: Lost in the Middle (TACL, 2024) Research · checked 9 October 2026
  6. Carnegie Mellon University and Duke University: TheAgentCompany benchmark (NeurIPS 2025) Research · checked 5 October 2026
  7. METR: Measuring AI ability to complete long software tasks (arXiv 2503.14499, v3 February 2026) Research · checked 9 October 2026
  8. NCSC: Mistaking AI vulnerability could lead to large-scale breaches (December 2025) Official source · checked 9 October 2026
About this guide

Written for Orwell Lab’s practical guide collection. Benchmarks, papers and official guidance checked on 9 October 2026.

Discuss your own requirements ↗

A useful next step.

Tell us what you’re building, what could work better and where you want to take the business.

Book a free discovery call