Close the Eval Gap: Why an Agent That Passes Its Evals Is Where the Real Work Begins
AI agent evals can pass while production fails. Learn why static evals miss real-world failures and how continuous production intelligence closes the gap.
Hamel Husain's essay "Your AI Product Needs Evals" is, correctly, one of the most cited pieces of practical wisdom in applied AI right now. His central claim is hard to argue with:
"Unsuccessful products almost always share a common root cause: a failure to create robust evaluation systems."
He's right. Teams that can't evaluate quality, can't debug failures, and can't change behavior quickly are teams that stay stuck in demo purgatory forever. His three-level framework - unit tests, human and model evaluation, then A/B testing - is the correct way to build an AI product. We'd tell any team building an agent to read it before they read anything we publish.
But there's a sentence buried in that same essay that we think deserves more attention than it gets:
"You can never stop looking at data - no free lunch exists."
Most teams read that as a footnote about diligence. We read it as the whole argument. Because the honest question isn't whether you need evals. It's: evals of what data? And the uncomfortable answer is that no eval suite, however well built, evaluates the data you haven't seen yet - which, the day you ship, is all of it.
The Two Loops
There are really two different feedback loops running on any AI agent, and they get treated as one.
Loop 1: The Eval Loop. This is Husain's territory - the one you control. You write the test cases. You curate the golden answers. You decide what counts as a pass. It runs before launch, and it runs fast, because you're the one setting the questions.
Loop 2: The Production Loop. This is the one nobody controls. Real users, with real intent, phrasing things you didn't anticipate, in sequences you didn't design for, hitting edge cases your eval set never contained because you didn't know they existed. It runs the moment the agent goes live, and it never stops.
Loop 1 tells you whether your agent does what you built it to do. Loop 2 tells you whether that was ever the right question. An agent can pass every test in Loop 1 and still fail constantly in Loop 2 - not because the eval suite was badly built, but because an eval suite, by construction, can only test what someone thought to test.
That's not a criticism of evals. It's a description of their boundary.

The eval doesn't hold up on its own terms
Here's where it gets uncomfortable, because the evidence for that boundary isn't hypothetical - it's showing up in the numbers teams are now publishing about their own agents.
Research from Galileo and Google Cloud,( reported by Prefactor) , found enterprise agents scoring roughly 60% success on a single-run evaluation - and then dropping to around 25% success across eight consecutive runs in production. That's not a rounding error. That's a 35-point collapse between "passed the eval" and "works when a real person uses it more than once."
The math behind it is unglamorous but it's exactly why this happens. If an agent is 95% accurate at each individual step - a perfectly respectable eval score - and a real task takes ten steps end-to-end, compounding accuracy across the chain puts you at roughly 60% correct by the time the task actually finishes. Every step-level eval can look clean. The chain still breaks.
The same reporting cites a Venture Beat survey from mid-2026: only 5% of organizations say they fully trust their automated evaluations, and 29% report a meaningful mismatch between what their evals say and what actually happens in production. That's not a tooling problem you fix by writing more test cases. It's a structural gap between a loop that's finite by design and a production surface that isn't.
A separate argument from( AgentForge Hub )makes the same point from a different angle, and states it more bluntly than we would:
"A model can climb a public leader board and still fail your onboarding workflow." "Production agents are systems, not just models."
That's the crux of it. An eval tests a model, or a well-scoped feature of a product, against a fixed idea of correct behavior. A production agent is a system - model, retrieval, tools, APIs, guardrails, and a real user with their own goals - operating in an environment nobody fully specified in advance.
Klarna's customer service agent is the case everyone cites for a reason: strong outcome metrics at scale (millions of conversations a month, resolution time down from roughly eleven minutes to two), and still a quality gap on complex cases significant enough that human agents were reintroduced into the loop.
That's not a failure of Klarna's evals. It's evidence that outcome-only evaluation, however well it scores, can miss exactly the failures that only show up in the weirder 5–10% of real conversations.
What production actually changes about an agent
The thing that changes the moment an agent goes live isn't the model. It's the shape of the input.
An eval set, however large, is a distribution its authors chose. Production traffic is a distribution nobody chose - it's whatever your actual users decide to say, in whatever order, with whatever context, on whatever day.
Morgan Stanley's DevGen.AI work migrating COBOL is a useful illustration of why this matters even when per-step error rates look small: at the scale of nine million lines of code, a 1% step-level misclassification - something well within an acceptable eval tolerance - still propagates into real downstream failures, because production doesn't sample from your test set. It samples from reality.
This is also where the qualitative failures live - the ones an outcome-only eval genuinely cannot see, because the agent's output looked fine. An agent can answer a question using the wrong source document and produce a response that reads as completely correct, right up until the same question is asked with slightly different context and the wrong source stops coincidentally giving the right answer. Silent failures like this don't fail uniformly. They fail when context shifts - which is precisely the condition an eval set, built before you knew what the shifts would be, cannot contain.
None of this is an argument that evals don't work. It's an argument that they're solving a different problem than the one that starts the moment your agent is live.
Why the fix has to be as continuous as the failure
If Loop 2 never stops, then the response to it can't be a one-time exercise either. A quarterly audit of production transcripts, or an engineer reading a sample of conversations when something looks off, is applying Loop 1 thinking - bounded, scheduled, sampled - to a problem that is by definition unbounded, continuous, and unsampled in how it actually occurs.
The teams we talk to who are furthest along on this have stopped asking "did we pass the eval" as the final question and started asking three different ones, continuously:
What's actually failing in production right now, and how often?
What are users asking for that the agent can't do - not a bug, just a gap nobody built?
When we ship a fix for either of those, did the failure rate actually go down, or did we just believe it would?
Those three questions don't have good answers if the only instrument you have is a fixed eval suite and a team reading a handful of transcripts a week. A team can realistically read maybe 1–2% of what a moderately busy agent handles. The failures that cost you the most customers are disproportionately likely to be sitting in the 98% nobody opened.
Close the Gap
This is the gap VoxMith is built to close: what happens after your AI agent goes live.
Static evals prepare an agent for production. VoxMith learns from what actually happens in production - continuously analyzing real customer interactions after deployment.
1. Observe
VoxMith continuously captures and analyzes 100% of production interactions - including prompts, models, RAG, workflows, tools, APIs, guardrails, and responses.
You see how your agent actually behaves with real users, not just how it performs in a test environment.
2. Diagnose
VoxMith evaluates production conversations, detects failures, and finds their root causes.
It surfaces hallucinations, workflow failures, policy violations, tool and API failures, drop-offs, and escalations - then traces them back to the prompt, knowledge, workflow, tool, model, or integration responsible.
3. Learn → Improve
VoxMith turns those production signals into action.
It identifies recurring patterns, capability gaps, and user friction, then generates engineering recommendations or, with approval, creates PRs and engineering tickets with the production evidence attached.
After the fix ships, VoxMith measures whether production performance actually improved and detects regressions. That's the proof loop: every release is measured against what happened in the real world.
Evals tell you if your agent is ready for production. VoxMith tells you what production teaches you after it gets there.
Curbing the gap isn't the same as distrusting your evals

None of this is an argument for skipping Husain's framework, or for treating evals as theater.
The Level 1 and Level 2 work - fast assertions, human and model review, the discipline of actually looking at your data before you ship - is what gets an agent good enough to be worth deploying in the first place. Skipping that step doesn't make a team more production-focused; it just means they ship something worse and find out later, more expensively, in front of real customers.
The choice isn't between rigorous pre-launch evaluation and continuous production intelligence. It's a false one if you frame it that way. The teams getting this right run both: a tight eval loop to earn the right to ship, and a continuous production loop to find out what shipping actually revealed. An agent that only has the first will pass its tests and still surprise you in the field. An agent that only has the second was never good enough to deploy responsibly in the first place.
A team that stops at "it passed the evals" hasn't finished the job. It's just reached the point where the real feedback loop - the one it doesn't control, the one that never stops - begins.
VoxMith is a production intelligence platform for conversational AI. We read 100% of your agent's production conversations, cluster what's failing, surface what users are asking for that the agent can't do yet, test candidate fixes against your own real traffic before they ship, and verify whether the fix actually worked once it's live. Book a call to see it against your own agent's traffic.
Sources:
FAQs
What are AI agent evaluations (evals)?
AI agent evaluations are tests used to measure how reliably an agent performs specific tasks against predefined test cases, expected outcomes, or quality criteria.
Why do AI agents pass evals but fail in production?
Evals test a fixed set of scenarios, while production exposes agents to unpredictable user inputs, edge cases, context shifts, tool failures, and workflows that weren't included in the evaluation set.
Are AI evals enough to monitor an agent after deployment?
No. Evals are essential before deployment, but they don't replace continuous analysis of real production interactions. Production monitoring reveals failures and capability gaps that a fixed eval set may never contain. Pasted text
How can teams improve AI agents after they go live?
Teams should continuously observe production conversations, diagnose failures and their root causes, identify recurring patterns, implement fixes, and verify whether those fixes actually improve production performance.


