AI · LLM Evals · Agents · Engineering

Why "It Feels Smart" Isn't a Metric: How to Actually Evaluate an AI Agent

1 October 2026 · Devectra

The demo is the worst test you have

Every AI agent demo looks the same: someone types a reasonable question, the agent calls the right tool, the answer comes back clean, the room nods. That demo proves exactly one thing — the agent can handle the path someone already knows works. It proves nothing about the other 10,000 paths a real user will take, and it proves nothing about what happens when a tool call fails, a document is ambiguous, or the user asks something slightly off-script.

We've written before about why RAG breaks in production and when a business actually needs an agent. This post is about the part that comes after you've decided to build one: how you find out, before a customer does, that it doesn't work.

Most teams don't have an answer beyond "we tried it a bunch and it seemed fine." That's not an evaluation. It's a vibe, and vibes don't scale past the five questions your team happened to think of.

What a demo can't show you

Agent failures mostly don't look like the model being dumb. They look like specific, repeatable breakdowns in a multi-step process:

  • Tool selection errors — the agent calls the wrong tool, or calls a tool when it didn't need to.
  • Parameter extraction mistakes — it picks the right tool but passes it the wrong arguments (the wrong date range, the wrong customer ID, a malformed filter).
  • Reasoning drift — on a long task, it loses track of the original goal somewhere around step 12 of 20.
  • Error non-recovery — a tool call fails or returns an empty result, and the agent doesn't adapt; it either hallucinates an answer anyway or gets stuck.
  • Policy violations — the task technically succeeds, but the agent did something it shouldn't have along the way (quoted a price it's not authorized to quote, accessed a record it shouldn't have touched).

None of these show up in a five-minute demo of the happy path. A single user task can trigger 20 to 60 model calls once an agent is chaining tool use, retrieval, and multi-step planning — and a failure in any one of those steps can produce an answer that still sounds confident. That's the trap: a bad answer delivered fluently reads as a good answer to anyone who isn't checking the work.

The three layers teams actually need

The useful way to think about agent evaluation is in layers, not one big score:

1. Model-level evals. Does the underlying LLM call produce a correct, well-formed response for a given input, in isolation? This is the simplest layer and the one most teams already do.

2. Trajectory evals. Did the agent take an acceptable path to the answer — right tools, right order, reasonable number of steps — not just land on a correct final output? An agent that gets the right answer by calling the wrong API twice and getting lucky is not a reliable agent.

3. Outcome evals. Did the task actually get done in a way that matters to the business — correctly, safely, within cost and latency bounds, without violating a policy?

Most teams that skip straight to "does the final answer look right" are only checking layer three, and only on the examples they thought to write down. The failures that actually cost you — the ones a support team discovers three weeks after launch — usually live in layer two.

LLM-as-judge works, but only if you calibrate it

Hand-grading every agent run doesn't scale, which is why using a second, more capable LLM as a judge has become the default approach for scoring agent outputs against a rubric. It's fast and it correlates reasonably well with human judgment — when it's set up correctly.

The step teams skip is calibration. Before you trust an LLM judge for anything that gates a release, hand-label a sample of runs yourself — a few hundred is enough to start — then run the judge on the same sample and measure agreement (Cohen's kappa is the standard statistic here). Below roughly 0.6 agreement, the judge isn't reliable enough to trust. Above 0.85, it's solid enough to sit in a CI gate. Skipping this step means you've swapped "we think it seems fine" for "a model we haven't checked thinks it seems fine," which is not the upgrade it looks like.

A workable split in practice: LLM-as-judge for the bulk of routine scoring, deterministic automated checks (schema validation, exact-match on structured fields, policy assertions) for anything that gates a deploy, and human review reserved for calibration and for the compliance-sensitive cases where you can't afford to be wrong.

Testing once isn't testing

A test suite you run before launch tells you the agent worked on launch day, against the inputs you anticipated. Production traffic will not stay inside that boundary — users phrase things you didn't expect, your data changes, the underlying model provider ships a quiet update. The current practice among teams running agents at real scale is continuous, sampled evaluation: score a slice of live traffic (often 5–20%) asynchronously and alert on changes in the pass rate, not just a static threshold. A drop from 94% to 87% over a week is a signal worth investigating even if 87% still sounds fine in isolation.

Uber's agent platform team spent over a year on this problem and concluded the hard part was never the tooling — it was building the habits: making evaluation the default on every project instead of an afterthought, and routing production failures back into new test cases so the eval suite actually grows instead of going stale. That's an organizational problem as much as a technical one, and it's usually the reason an eval setup that looked solid at launch quietly stops meaning anything six months later.

What this means if you're the one buying

If a vendor is proposing to build you an agent — for support, for internal ops, for anything customer-facing — ask what layer two and three look like, specifically: what happens when a tool call fails, how do they know if the agent takes a bad path even when the final answer looks right, and how is quality monitored after launch rather than just before it. "We tested it extensively" is not an answer. A described eval set, a judge they've calibrated against real examples, and a plan for sampling production traffic after go-live is.

Takeaway: an agent that looks smart in a demo and an agent that's reliable in production are different claims, and the gap between them is exactly the testing most teams skip. If nobody can tell you how they'd catch a silent failure three weeks after launch, that's the question to ask before you sign off, not after.

Have a system like this in mind?

Get a scoped plan ↗
Why "It Feels Smart" Isn't a Metric: How to Actually Evaluate an AI Agent — Devectra Blog — Devectra