AI · 1 min read
What It Actually Takes to Run AI Agents in Production
Demos are easy. Reliability is not. Notes on evaluation, guardrails, and the unglamorous engineering behind useful agents.
By Saifulah Katpar · September 22, 2026
Every week brings a new agent demo that books flights, writes code, or runs a company. Very few of them survive contact with real users. The gap between a demo and a product is almost entirely engineering, not model quality.
Start with evaluation, not prompts
Before tuning a single prompt, write down what "correct" means. A small, honest evaluation set — fifty real tasks with expected outcomes — beats any amount of intuition.
- Capture real inputs from users, not invented ones
- Score outcomes, not intermediate steps
- Re-run the suite on every change
Constrain the action space
Agents fail most often when they have too many options. Give them fewer, sharper tools.
const tools = {
searchDocs: (q: string) => index.search(q, { limit: 5 }),
createTicket: (title: string, body: string) => tracker.create({ title, body }),
};
The best agent is the one that knows exactly what it is not allowed to do.
Keep a human in the loop where it matters
Automate the reversible, review the irreversible. That single rule removes most of the risk while keeping most of the speed.
This is a sample article — replace it with your own writing.
Get new articles by email
Occasional essays on technology, AI, and software engineering. No spam.