An AI agent is software in which a language model decides, step by step, what to do next to reach a goal — using tools it has been given, information it can read, and limits someone has set. It isn't a single product or a single technology. It is an arrangement of parts, and most of what makes an agent useful, or risky, comes from how those parts fit together.
- An agent is a language model run in a loop. It looks at its situation, chooses an action, carries it out through a tool, reads the result and decides whether to continue.
- The model is only one part. The instructions, the data it can see, the tools it can call, the state it keeps and the limits around it matter as much.
- "Reasoning" is a function, not a guarantee. It means the model producing a next step from its context. That step can be wrong, and it is only as good as the information and tools available.
- An agent's autonomy is bounded by design — by its tools, what those tools may change, when it must stop, and which actions need a person.
- You know an agent works by observing and evaluating it, not by trusting it.
Many AI features call a model once: summarise this, classify that, draft a reply. The software around the model decides what happens before and after. An agent differs in one respect: the model decides what happens next. It chooses an action, takes it through a tool, looks at what came back and chooses again — until it judges the goal met or reaches a limit.
Anthropic's guide to building agents describes them as "typically just LLMs using tools based on environmental feedback in a loop". Three things in that sentence define the category: a model; tools that let it inspect or change something outside itself; and a loop in which each result feeds the next decision.
That is a narrower definition than the word suggests in marketing. A chat interface isn't an agent because it converses, and a workflow isn't an agent because one of its steps calls a model. If every step is fixed in code and the model only fills some of them in, it is a workflow that uses AI.
There is no single standard architecture, and frameworks name and divide these parts differently. Most agents are still built from the same ones.
The language model reads text — sometimes images or other inputs too — and produces text, including, in an agent, structured requests to call tools. It brings general language ability and patterns learned in training. It doesn't bring knowledge of your business, access to your systems, or any guarantee of being right. Its output can vary from one run to the next, and it changes when the model version changes.
Instructions tell the model its role, what it is trying to achieve, what it must and must not do, and how to format what it returns. The goal for a particular run — "resolve this support request", "prepare this month's supplier report" — usually arrives with the task. Clear instructions narrow what the model is likely to do, but they don't enforce anything. Whatever must never happen needs a control outside the model.
Context is everything the model can read when it makes a decision: the instructions, the task, the exchange so far, tool results and any documents passed in. The model can only work with what is in its context at that moment. If the order record isn't there, it can't reason about the order — although it may still produce a confident answer.
Tools are the functions an agent can call: search a knowledge base, query a database, read a file, call an API, create a ticket, send a message. Each has a name, a description and defined parameters. The model asks for a call; the software around it executes the call and returns the result.
Tools are how an agent reaches anything beyond its own text, and they set the limits of what it can do. Anthropic's guide recommends investing as much effort in these agent-computer interfaces as in human-computer interfaces: clearly named, well-documented tools reduce the model's mistakes.
Within a run, the agent's state is what has happened so far: the steps taken, the results received, what is still pending. Across runs, some agents keep longer-term memory — notes, preferences, past outcomes — stored in a database and put back into the context when relevant.
Memory isn't the model learning. It is data the system saves and chooses to show the model again, and it needs the same care as any stored data: what is kept, for how long, and who can see it.
Guardrails are the controls that don't depend on the model behaving: which tools it has, what permissions those tools hold in other systems, validation of its outputs, limits on steps and cost, and approval requirements for certain actions. They turn "the model probably won't do this" into "the system can't do this".
Evaluation checks — before launch and after every significant change — whether the agent reaches the right outcomes on representative cases. Monitoring records what it actually does in use: each step, each tool call, each result, so problems can be found and traced. Without both, there is no way to know whether an agent is working, or when it stopped working.
A practical way to picture one run of an agent — a common way of describing it, not a formal standard:
- Observe. Read the goal, the instructions and the current context.
- Decide. The model proposes the next step: call a tool with certain arguments, ask a question, or finish.
- Act. The software around the model checks the request against its rules and, if it is allowed, executes it.
- Inspect. The result is added to the context. Anthropic stresses that agents need to gain "ground truth" from the environment at each step, "such as tool call results or code execution", to assess their progress.
- Continue or stop. The loop repeats until the model judges the goal met, a person needs to decide, or a limit is reached.
The "decide" step is what is usually called reasoning or planning. Functionally, it is the model generating a plausible next step from its context. That can be very useful, and it can be wrong: the wrong tool, a misread result, a task declared finished too early.
The loop doesn't make the model more reliable. It gives it more opportunities to act, which is why the stopping rules matter. Anthropic notes that it is "common to include stopping conditions (such as a maximum number of iterations) to maintain control."
A hypothetical example: an agent that answers account managers' questions about customer orders. It has five tools:
- Read —
get_customer(id)returns a customer record. It changes nothing. - Query —
search_orders(customer_id, status)lists matching orders. It can only read, and only the order data the task needs. - Update —
add_order_note(order_id, text)adds a note to an order. Because it changes a record, its input is validated: the order must exist and the note has a length limit. - Prepare —
draft_email(to, subject, body)creates a draft in the account manager's mailbox. It doesn't send anything; a person reviews it. - Request approval —
request_approval(action, reason)sends a proposed consequential action, such as a goodwill credit, to a person, and the agent waits for the answer.
Asked why an order was delayed and whether the customer can be offered something, the agent might read the customer record, query the order and its shipment, add a note summarising what it found, draft an email to the customer and request approval for a credit. Each of those steps is a tool call that the system can log, check and, where necessary, refuse.
Notice what the agent doesn't have: a tool that issues credits, or one that sends email. Those capabilities exist elsewhere in the business; they were deliberately left out of its reach. That design choice does more for safety than any sentence in its instructions.
"Data" means several different things in an agent, and they behave differently:
- Context supplied at run time. The task, the request, the files attached to it. The agent sees them for that run only.
- Connected business data. Systems the agent can query through tools — a CRM, an order database, an inventory. The data stays where it is; the agent asks for what it needs, within the tool's permissions.
- Persistent state or memory. What the system chose to save from earlier runs and show the agent again. It is only as accurate as what was saved.
- Retrieval. When the knowledge an agent needs sits in documents — policies, manuals, past tickets — a retrieval system finds the passages that match the task and adds them to the context. This is retrieval-augmented generation (RAG). It improves answers only when the right passages are found, and building and evaluating LLM and retrieval integrations is a discipline of its own.
In every case, the model only "knows" what reaches its context. Missing or outdated data produces confident decisions on the wrong basis.
- Incorrect decisions. The model chooses a plausible but wrong step, misreads a tool result, or concludes too early that the task is done.
- Bad or incomplete context. Missing records, outdated documents or retrieval that returns the wrong passage produce well-formed answers built on the wrong facts.
- Tool failures. An API times out, returns an error or returns something unexpected. Unless the system handles errors explicitly, the agent may retry endlessly, work around the failure in a way nobody intended, or carry on as if the step had succeeded.
- Prompt injection and untrusted inputs. Text the agent reads — an email, a web page, a document — can contain instructions that change its behaviour. OWASP puts prompt injection first in its Top 10 for LLM applications and describes the indirect form, which occurs "when an LLM accepts input from external sources, such as websites or files". Anything an agent reads from outside should be handled as data, not as instructions.
- Excessive permissions. OWASP's Excessive Agency traces harm caused through agents to excessive functionality, excessive permissions and excessive autonomy: more tools, more access and more freedom to act than the task requires.
- Compounding errors. Each step builds on the one before, so an early mistake carries forward. Anthropic notes that agents' autonomy brings "higher costs, and the potential for compounding errors".
- Unclear stopping conditions. Without a step limit, a budget and a clear definition of "done", an agent can loop, repeat work or keep spending. Unbounded Consumption is another entry in the same OWASP Top 10.
None of these is exotic. They are the ordinary ways a system that decides from text can fail, and each has a design response: validation, narrowly scoped tools, explicit error handling, limits, logging and review.
An approval boundary is a rule, set in the system rather than in the instructions, that sends certain actions to a person before they happen. OWASP's guidance is direct: "Utilise human-in-the-loop control to require a human to approve high-impact actions before they are taken." Anthropic describes agents that "pause for human feedback at checkpoints or when encountering blockers".
Actions that usually warrant a boundary:
- Irreversible actions — deleting data, messaging a customer, publishing.
- Financial actions — refunds, credits, payments, orders.
- Actions outside the usual pattern — an amount above a threshold, an account flagged for review, a request unlike the ones the agent was tested on.
- Uncertain situations — conflicting data, a failed tool call, an ambiguous request.
An agent should also stop when it is blocked: information it can't obtain, a tool that keeps failing, a step limit reached. Stopping and explaining why is a correct outcome, not a failure.
A conceptual model — a way to reason about an agent, not a universal architecture:
| Stage | What happens | What controls it |
|---|---|---|
| Input | A task or an event starts a run | Who or what may start the agent |
| Context | Instructions, task, data and earlier results are assembled | What data the agent may see; the quality of retrieval |
| Decision | The model proposes the next step | The instructions; the tools on offer |
| Action | The system executes an allowed tool call | Tool permissions; validation; approval boundaries |
| Result | The tool returns data or an error | Error handling; treating returned text as untrusted |
| Evaluation | The step and its outcome are logged and checked | Logs; review; evaluation on test cases |
| Next step or stop | Repeat, hand over to a person, or finish | Step and cost limits; the definition of done |
Read down the last column: nearly every control sits outside the model. That is where an agent's reliability is designed.
You need an agent when a task requires choosing the next step from what is found along the way, and that choice can't be written as rules in advance. If the steps are fixed, a workflow is simpler and more predictable; if only reading the input is hard, a single model step inside a workflow may be enough. AI agents vs automation sets out that decision in detail.
When a task does call for an agent, most of the work is in the parts around the model: scoped tools, the right data, clear limits, approval boundaries and monitoring. That is the substance of developing AI agents that work across business systems.
An AI agent isn't a smarter model. It is a model placed in a loop with tools, data, state and limits. The model proposes each step; everything around it decides what that step can reach, what it can change and when the run has to stop. Knowing those parts is what makes it possible to judge an agent realistically: what it can do, how it can fail, and what has to surround it before it is trusted with real work.
Checked on 2 October 2026.



