AI
What it takes to put an AI agent into production
An agent that works in a demo is a long way from one you can trust with real customers and real systems. Here is what closes the gap.
It has never been easier to build an impressive AI agent demo. Give a capable language model a few tools, a clear instruction and a friendly test case, and within a day it will look up an order, draft a reply and update a record. The demo works, everyone is excited, and the question becomes how quickly it can go live.
The honest answer is that the demo is perhaps the first tenth of the work. The rest is about what happens when inputs are messy, when a tool fails, when a user tries to trick the system, and when the model is confidently wrong. None of that is exotic. It is ordinary production engineering applied to a component that behaves probabilistically.
Start with a narrow, valuable task
The agents that succeed in production tend to do one well-defined job. Triage incoming support emails and draft a reply for a person to approve. Extract fields from supplier invoices and flag anything unusual. Answer questions about internal policy documents with citations. Each has a clear input, a clear output, and a way to tell whether the result was right.
Open-ended agents that are asked to handle anything a customer might say are much harder to make reliable, because there is no bounded definition of success. If you are early in AI development, choose a task where a mistake is cheap to catch and the volume is high enough for the saving to matter.
Design tools and permissions deliberately
An agent is only as safe as the tools it can call. Treat each tool as an API you are exposing to a user who is capable but occasionally unpredictable. In practice that means:
- Least privilege. Give the agent read access by default and grant write actions one at a time, only where needed.
- Narrow tools over general ones. A tool that issues a refund up to a set limit on a specific order is far safer than one that runs arbitrary database queries.
- Validation outside the model. Check every tool call's arguments in code, against business rules, before anything executes.
- Idempotency and reversibility. Prefer actions that can be safely retried or undone, and log enough to reconstruct what happened.
Most of this is conventional API integration work. The difference is that the caller is a model, so the boundaries have to be enforced by the system rather than trusted to the prompt.
Put humans where the risk is
Human approval is not a sign that the agent failed. It is a design tool. Actions that move money, contact customers, change records of legal significance or cannot be undone should usually pass through a person, at least initially. The agent does the gathering, reasoning and drafting; a person confirms with one click. Over time, as you build evidence about accuracy on specific action types, you can relax approval for the low-risk ones.
The prompt is a suggestion. The permissions are the policy.
Evals are how you know it works
Without a way to measure quality, every change to a prompt, model or tool is a guess. An evaluation set is a collection of realistic inputs with known good outcomes, run automatically whenever something changes. It should include ordinary cases, awkward edge cases and examples that previously went wrong.
Some checks can be exact: did the agent call the right tool with the right order number? Others need judgement, such as whether a reply was accurate and appropriately toned, and can be scored with rubrics, sometimes using a separate model as a grader that you have checked against human judgement. The point is not a perfect score; it is to notice when a change makes things worse before your customers do. Keep adding real failures from production to the set, so it grows more representative over time.
Guardrails and prompt injection
Any text an agent reads, whether it is an email, a web page, a document or a support ticket, can contain instructions. Prompt injection is the risk that the model follows those instructions instead of yours. There is no complete fix at the model level, so the defence has to be structural:
- Treat all retrieved and user-supplied content as untrusted data, and keep it clearly separated from system instructions.
- Limit what the agent can do after reading untrusted content, especially sending data outside the organisation.
- Filter outputs for sensitive data before they leave the system.
- Require approval for high-impact actions regardless of how confident the model appears.
Input and output guardrails, such as topic restrictions, format validation and checks for personal data, add further layers. None is sufficient alone. Together, with tight permissions, they reduce the blast radius when something slips through.
Observability, cost and latency
You need to see what the agent did and why. That means tracing each run: the inputs, every model call, every tool call and its result, the final output, and the time and tokens each step used. When a user reports a strange answer, you should be able to replay the exact run. Good tracing also exposes patterns, such as a tool that fails often or a class of question the agent consistently fumbles.
Cost and latency deserve attention from the start. Multi-step agents can make many model calls per request, and both spend and response time scale with that. Useful levers include routing simple steps to smaller, faster models, caching repeated context, capping the number of steps per run, and setting per-user and per-day budgets with alerts. These are the same operational disciplines we apply to any system through cloud and DevOps practice.
Where to begin
Pick one workflow, build the eval set before you polish the prompt, launch with a person approving every action, and widen autonomy only as the evidence supports it. That approach is slower than the demo suggests, but it produces something the business can actually rely on, which is the only kind of agent worth running.