Agent chats can score perfect and still fail: LangChain, Conviva, CoreWeave warn executives
Evaluation needs cohorts, baselines, and monitoring, not just “looks good” trace scoring and cost-heavy judge models.

At VB Transform 2026, Harrison Chase of LangChain, Hui Zhang of Conviva, and Emmanuel Turlay of CoreWeave argued that agent evaluations can look flawless while products remain broken. Their answer: shift from single-trace scoring to contrastive, baseline-driven cohort analysis and build monitoring plus targeted offline evals.
A single AI agent conversation can look flawless on its own, get high scores, and still point to a broken product. That gap, leaders from LangChain, Conviva, and CoreWeave argued at VB Transform 2026, is forcing enterprises to rethink how they evaluate agents in the first place.
Harrison Chase, CEO of LangChain, framed it as a mismatch between evaluation theater and shipping reality. Teams often build an eval set, freeze development because they cannot “prove” performance, then stall. Chase’s blunt alternative: “The best teams launch and then iterate,” and he treats evals as a living specification, like a product requirements document. In his framing, “Evals are like the new PRD,” defining what your agent should and shouldn’t do.
That “still broken” problem has a deeper root than impatience. Hui Zhang, CTO and co-founder of Conviva, criticized a common enterprise habit: scoring traces one at a time. Even if you score 50 examples or an entire population of traces, you can miss the signal that only appears when you compare cohorts of users against a baseline. Zhang called this contrastive analysis, and he illustrated it with a retail scenario: a shopper asks an agent for a running shoe before a half marathon; the agent asks qualifying questions; the shopper buys the shoe. If you score that interaction in isolation, it may look fine. But when Zhang compares the full user population against a baseline category, two cohort-level metrics move in ways a single trace never would: the clarification ratio, how many follow-up questions the agent asks before completing a task, was three times higher than baseline for that shoe category; and the completion rate outside the conversation, how often shoppers finished the purchase outside the conversation, was five times higher than baseline for the same category.
The implication is uncomfortable for anyone building an “AI agent scorecard.” An agent can pass your trace grading and still be failing the job your business actually needs done. Zhang added that teams also often lack a second data source: what happens before, between, and after the conversation, not just the trace itself. In other words, the conversation log is the visible surface, but the product truth lives in user behavior around it. If you only grade the conversation, you are likely optimizing for the part that is easiest to measure.
Chase and Zhang also discussed why the industry’s judge models do not automatically solve this. Agent-as-judge, using one AI agent to judge another agent’s output, hasn’t replaced LLM-as-judge, which Chase said remains the default. But there is still a scalability trap: you can grade outcomes, and you can still struggle to ground the judgment to something reliable. Zhang tied it to a core tension: automated judging gives scalability but can be ungrounded, whether the judge is an agent or an LLM; grounding is hard; then you use humans; and “that’s just not scalable.” So the evaluation question becomes: how do you get the signal without paying the human bill for every edge case?
Emmanuel Turlay, director of engineering at CoreWeave, offered a cost-minded approach to judging. He said teams should start with the most capable model available to prove a task is solvable, then work down. If it cannot be done with a top-tier model, he argued it will not work with a smaller one. Once a pattern is viable, teams can sample only a fraction of traffic instead of judging every interaction. And for simpler tasks like binary classification, the path can be to move to smaller open source models. The practical goal is not to eliminate models from evaluation, but to match the model to the job.
Turlay also pushed back on the idea that an exhaustive pre-launch test suite is the endgame. He said he tried to reach 100% coverage for tests and still had bugs in production. The fix, he argued, is broad, always-on monitoring. Set wide online checks first. Use those to identify failure classes as they occur. Then build a targeted offline evaluation set around the problems that surface. Chase’s eval-paralysis warning fits neatly here: the goal is not to build a perfect gate that blocks every release, it is to build an evaluation system that keeps improving as reality changes.
When it comes to “human in the loop,” none of these leaders suggested humans are optional forever. Turlay pointed to accountability, borrowing from his prior work on a self-driving car company. Even with aggressive iteration cycles, he said someone still had to sign off on deployment, because companies need someone legally responsible. The same logic extends to legal, finance, and healthcare: before a human can be removed entirely, it will “be a while” before agents can do that on their own. Zhang agreed that humans have to remain a guardian on corner cases, even as automation can beat individual humans at pattern-level scale. Chase went further, saying human-in-the-loop is also important for building trust in how agentic systems work, and for memory and learning from systems. “There has to be interactions in order for the system to learn.”
For decision-makers, the strategic stakes are simple. If you evaluate only the conversation trace, you can ship a product that looks good in dashboards but drives worse outcomes in the real world, like the cohort shifts Zhang described. And if you rely on an eval set as a one-time gate, you risk eval paralysis instead of iteration. The leaders at VB Transform 2026 are pushing a more production-minded philosophy: treat evals as a living spec, use contrastive cohort analysis against baselines, watch continuously in production, and size judge models to the economic reality of what you are trying to catch.
This story's Key Insights and Take-aways are locked.
Create a free account to unlock Executive Actions for one credit.
Register to UnlockAlways free for Executives Club members. Join the Club
More in Technology

Nvidia folds CPUs and GPUs into Vera Rubin to control more of AI data centers
The Vera Rubin platform merges CPU and GPU compute into one system, signaling Nvidia’s push to own the whole stack.

Google launches Gemini 3.5 Flash Cyber to patch vulnerabilities fast, cheaply
An AI security model built on Gemini 3.5 Flash aims to let agents scan more code paths at low cost.

Suno breach exposed 55M users with names, phone numbers, and addresses, report says
Have I Been Pwned says an attacker took identifiable customer data, turning AI creativity into an urgent security problem.

