AI · 1 min read

What It Actually Takes to Run AI Agents in Production

Demos are easy. Reliability is not. Notes on evaluation, guardrails, and the unglamorous engineering behind useful agents.

By Saifulah Katpar · September 22, 2026

Every week brings a new agent demo that books flights, writes code, or runs a company. Very few of them survive contact with real users. The gap between a demo and a product is almost entirely engineering, not model quality.

Start with evaluation, not prompts

Before tuning a single prompt, write down what "correct" means. A small, honest evaluation set — fifty real tasks with expected outcomes — beats any amount of intuition.

  • Capture real inputs from users, not invented ones
  • Score outcomes, not intermediate steps
  • Re-run the suite on every change

Constrain the action space

Agents fail most often when they have too many options. Give them fewer, sharper tools.

const tools = {
  searchDocs: (q: string) => index.search(q, { limit: 5 }),
  createTicket: (title: string, body: string) => tracker.create({ title, body }),
};

The best agent is the one that knows exactly what it is not allowed to do.

Keep a human in the loop where it matters

Automate the reversible, review the irreversible. That single rule removes most of the risk while keeping most of the speed.

This is a sample article — replace it with your own writing.

Get new articles by email

Occasional essays on technology, AI, and software engineering. No spam.