AI Evaluation

Evaluating AI Agents in Production: Accuracy, Cost and Reliability

An agent that's 95% reliable per step succeeds end-to-end only about a third of the time across a 20-step chain — here's how to actually measure whether an agent is ready for production.

By Naeem Akhtar · 8 min read

01

Why per-step accuracy is the wrong headline metric

It's intuitive to ask "how accurate is the model?" and treat a high per-step accuracy number as evidence the agent is production-ready. It isn't, because failures compound across a multi-step task. An agent that's correct 95% of the time on each individual step completes an end-to-end task correctly only about 36% of the time across a 20-step chain — reliability multiplies against you with every additional step, tool call, or decision point.

This is why task completion — measured on the actual end state, not on whether each intermediate step looked reasonable — is the metric that matters, and why it needs to be measured on realistic multi-step tasks rather than isolated single-turn prompts.

02

The five dimensions that cover the real failure surface

DimensionWhat it catches
Task completionWhether the end state is actually correct, verified against ground truth — not whether the transcript looks plausible.
Tool-call accuracyWhether the agent picked the right tool, with the right arguments, at the right step.
Trajectory / step countWhether it reached the right answer via a sane path — an agent that loops five extra times before succeeding is a reliability risk waiting to happen under load.
Cost and latencyWhether the agent is economically and practically viable to run at the volume the business actually needs.
GroundednessWhether the agent's claims are actually supported by retrieved context or tool output, rather than a confident hallucination stitched in.

Shipping against only one or two of these — usually just task completion — is how teams end up with an agent that scores well in testing and then quietly burns through budget or produces ungrounded answers in production.

03

pass@k: measuring reliability, not just correctness

A single successful run doesn't tell you much about whether an agent is reliable — LLM outputs are non-deterministic, and a task that succeeds once can fail the next time on the exact same input. pass@k measures whether the agent completes the same task class successfully across k independent attempts, typically k=3 to k=10.

Why this matters

An agent with a high single-run success rate but a low pass@k score is telling you it's inconsistent — which is often worse for user trust than an agent that's consistently mediocre, because inconsistency is unpredictable and hard for a team to build workarounds around.
04

Offline and online evaluation are both required, not either/or

Offline evaluation runs a fixed dataset of known tasks with known-good outcomes, in CI, on every meaningful change — this catches regressions reproducibly before anything reaches a real user. Online evaluation scores real production traffic as it happens, catching drift, novel failure modes offline testing never anticipated, frustrated users, and adversarial or edge-case inputs that a curated test set wouldn't include.

  • Offline: fast, reproducible, catches known failure modes before release, but blind to anything not represented in the test set.
  • Online: catches real drift and novel failures, but by definition happens after the agent is already live in front of users.
  • Teams that skip offline testing ship regressions repeatedly. Teams that skip online monitoring find out about drift from a support ticket instead of a dashboard.
05

Eval-gated promotion: making this a process, not a one-time check

The teams that consistently ship reliable agents treat evaluation as a continuous discipline: automated metrics for coverage, model-based screening for efficiency at scale, and human expert judgment for the correctness only domain knowledge can verify — combined, not chosen from. No agent version should move from development to staging, or staging to production, without meeting defined thresholds on all five dimensions simultaneously. This is exactly what a decisioning-layer rebuild caught on a vendor compliance pipeline we worked on: the original set of 24+ disconnected automations had no shared evaluation or visibility, so failures were discovered by a partner complaining, not by a dashboard. A single pipeline with a standing evaluation set turns that into something the team can see coming.

06

Frequently asked questions

Why does an agent with 95% per-step accuracy fail so often end-to-end?

Because reliability compounds multiplicatively across steps: 0.95 raised to the 20th power is roughly 0.36. A long multi-step task with many decision points or tool calls needs much higher per-step reliability than an intuitive read of "95% accurate" suggests, which is why task-completion rate on realistic multi-step tasks is a more honest metric than per-step accuracy alone.

What is pass@k in AI agent evaluation?

pass@k measures whether an agent can successfully complete the same task class across k independent attempts (commonly k=3 to k=10), rather than relying on a single run. It's a reliability metric — an agent that succeeds inconsistently across repeated attempts on the same task is a production risk even if any individual run looks fine.

Do I need both offline and online evaluation for an AI agent?

Yes. Offline evaluation (a fixed set of known tasks run in CI on every change) catches regressions before release but is limited to what's in the test set. Online evaluation (scoring real production traffic) catches drift and novel failure modes offline testing can't anticipate, but only after the agent is already live. Relying on only one leaves a real gap.

Should cost and latency be part of an AI agent's evaluation criteria?

Yes — an agent that's accurate but slow or expensive to run isn't shippable at real business volume. Cost and latency should be tracked as first-class evaluation dimensions alongside accuracy, tool-call correctness, and groundedness, not measured separately or added as an afterthought once accuracy looks acceptable.