Introduction
Observability is ubiquitous in AI agent monitoring - latency, token counts, error rates flood dashboards. But high-level metrics dont guarantee the agent did the right thing. Two runs can look identical on every chart while producing opposite business outcomes. This piece explores why tracing alone isnt enough and what it takes to close the gap.
What Happened
A support agent processes refund requests. Two runs side by side. Every observability field - model, duration, token usage, finish reason, span status - is nearly indistinguishable. Both complete in about 3.2 seconds, consume roughly 2850 input tokens, and return a clean stop signal. On latency, spend, and error-rate charts they plot as the same point, twice. Yet in one run the agent refunded a 40 order; in the other, a 2300 order. The traces show the API call succeeded, the span is green, and from the runtime perspective nothing failed. The difference never appears in the telemetry because it lives outside the request-response cycle, in the business logic that follows.
Why This Matters
Observability answers what happened evaluation answers should it have A survey of 1340 engineers found that 57 percent of teams have agents in production 89 percent have observability implemented but only 52.4 percent run offline evaluations and 37.3 percent run online ones Quality accuracy relevance consistency tone is the top barrier to getting agents into production yet it is the least measured The core problem every trace attribute is a property of the call not the outcome Correctness requires reference to something the request does not contain annotated goal states database diffs business rules The article highlights the tau bench methodology which validates agents by comparing the final database state against an annotated goal not by reading the transcript It also introduces the pass k metric showing that an agent succeeding 90 percent of the time achieves full reliability across eight runs only about 43 percent of the time. Even state-of-the-art function calling agents succeed on fewer than half of tasks when evaluated this way. The piece also warns against uncritically trusting LLM as judge setups noting that high consistency can coexist with severe position bias and that judge validation often overstates discriminative ability without chance correction.
Key Takeaways
- Pull fifty real traces from production weighted toward unusual or escalated sessions and write down what should have happened as a program checkable assertion order IDs row counts exit codes
- Replay these assertions on every change prompt edits model version bumps tool-schema updates dependency upgrades gate releases on the results
- Report pass k not just the mean an agent that succeeds 90 percent of the time fails all eight attempts for nearly 57 percent of users
- Turn every production failure into a new labeled case the evaluation set strengthens exactly where the system is weakest
- When using LLM judges calibrate with chance corrected statistics like Cohens kappa and swap option positions to check for position bias
- Where possible verify the resulting state rather than the transcript assert on database rows test suite outcomes or document validity rather than re reading the agents output
Conclusion
Observability is solved buyable and nearly universal Evaluation is the hard part but its the one your customers are actually asking about Roughly half of teams running agents in production still have no systematic way to answer whether their agent is doing the right thing. If you have 89 percent of the stack covered by traces the remaining work isnt another dashboard its fifty labeled traces and an assertion that fails. That is where quality lives and that is where production readiness actually begins.




Discussion
Join the conversation
Thoughtful reactions, questions, and follow-up ideas help shape the next story.