The Blind Spot We Closed in Video Just Reopened in AI Agents

Around 2010, one of the first major broadcasters to launch a streaming product had a problem nobody could name. The apps loaded. The streams played. Every dashboard was green. And viewers were disappearing before a single frame appeared on screen.

The metric that explained it didn’t exist. We had to go find it and then name it: video start time. How long a person stares at a black rectangle before the content shows up. Once you can see that number, you can fix it. Until you can see it, you’re looking at green dashboards wondering where everybody went.

That’s how the quality of experience category actually got built. Not on a whiteboard — in the gap between what the systems reported and what the humans lived through. We’d been at it for a few years by then, and we’ve spent the twenty since turning that gap into a discipline for the world’s largest media companies.

That same gap just reopened, in a different room.

Consumer-facing AI agents are being deployed right now by brands with mature products, mature websites, and no way to answer one simple question: did the agent create a good experience for the person on the other end of the conversation?

Same story, bigger room

The market dynamics are almost identical to what we lived through in streaming.

Back then we were sitting across from billion-dollar broadcasters who had been told they were putting video on the internet whether they were ready or not. Today we’re sitting across from the same class of company — global e-commerce brands, airlines, multi-brand hospitality groups — who have been told they’re launching an agent whether they’re ready or not. Same compelling event, same compressed timeline, same executive on the hook for something they’ve never had to measure before.

The architecture parallel runs just as deep.

In streaming, the encoder and the decoder were never the differentiator. Everything wrapped around the video decided whether the experience held up — the CDN, the player, the origin server, the device. Two services could run the identical codec and deliver completely different experiences.

The model is not the differentiator in agentic AI either. It’s the harness: the tools, the context, the permissions, the memory, the data, the evaluation, the observability, and the production environment wrapped around that model. Same AI driver — the performance is all about the car. Practical agent performance is a function of everything surrounding the model, and almost none of it is the model.

That isn’t a new insight for us. It’s the one we’ve spent two decades proving, pointed at a market that didn’t exist when we learned it.

Four things have to be true

We didn’t get to build the quality of experience category by accident. It took four things being true simultaneously, and not one of them was optional. The same four are non-negotiable for agent experience.

Objective experience metrics. Not a vibe. Not a single sentiment score. Something you can point at, track over time, and improve.

Full population, not a sample. You cannot sample your way to the truth about how an agent is performing. The rare-but-critical failure patterns vanish the moment you sample, and those are usually the ones costing you the most money.

Continuous measurement in production. Here’s the part worth being precise about: pre-launch evals are necessary. They catch regressions, they block bad releases, and running an agent without them is reckless. They are also, by definition, pre-production. No curated test set tells you how an agent behaves when real people with real money start asking things you didn’t anticipate, in an order you didn’t design for. The wild is the only environment that matters, and it has to be watched live.

Contrastive, multi-dimensional analysis. A single macro number tells you almost nothing. You have to be able to cut into segments, cohorts, and situations to find where an agent is struggling and under what conditions — otherwise you know something is wrong and nothing about what to fix.

Experience, Population, In Production, Contrastive. We call it EPIC. It’s the same discipline that built our video business, given a name for a market that now needs it.

The metrics most agent teams aren’t measuring yet

Video start time didn’t exist until somebody went looking for it. Then it became one of the most important numbers in the industry.

The same thing is happening with agents right now, and these are the ones we’d start with.

Turns to intent. How many exchanges it takes before the agent understands what the person actually wants. Every unnecessary turn is friction for the customer and tokens off your bill.

Clarification ratio. How often a person has to restate the request they already made.

Manual search between turns. Whether someone quietly opens a second tab mid-conversation because the agent isn’t getting them there. This is the clearest signal that a conversation scoring well on paper is failing in practice — the customer doing the agent’s job for it, in real time.

Scope-narrowing rate. When a person simplifies their own request because the agent couldn’t handle the real one. That is not the agent succeeding. That’s your customer dumbing down what they came for until it fits what the agent can do.

Post-conversation correction rate. How often someone goes back and fixes the record after a chat the agent reported as resolved. Its close cousins matter just as much: verification-seeking, when the customer logs in to check the agent’s work, and post-conversation escalation, when they open a ticket about the thing that was just marked done.

None of these numbers mean anything at the level of a single conversation. They only mean something across the whole population — which is precisely why nobody reading transcripts one at a time has ever found them.

And notice what they share. Every one of them is a behavior that costs the customer patience and costs you tokens at the same time. Fewer wasted turns is a better experience and a smaller bill. Those aren’t two projects. They’re one problem viewed from two angles.

Your agent passed every eval. Do you know if it worked for the user?

That’s the question the industry hasn’t been asking, and it’s the one we’ve built our second act around.

Give the two incumbent answers their due. Observability tells you the system ran — spans, latency, errors, cost, and it does that well. Evals tell you a response scored well against a rubric somebody wrote in advance, and you want that rubric. Both are real engineering and both belong in your stack.

Neither one asked the user anything.

Neither can tell you whether the person on the other end of that conversation got what they came for, or gave up somewhere the trace never reached. A transcript can be helpful, relevant, correctly clarified, and score a fraction under perfect while the shopper it belonged to is already on a competitor’s site. Nothing in the trace tells you that. The behavior around the trace does.

We didn’t learn this in a lab

We learned it supporting live digital experiences across 10 Olympic Games, 20 Super Bowls, and 5 FIFA World Cup tournaments — in production, at population scale, in real time, on nights when a few seconds of frozen screen meant a viewer was gone for good.

That’s the only place this discipline can be learned. There’s no simulator for a penalty shootout, and there isn’t one for a customer who’s had enough.

Your evals grade the agent. We think the user should grade it. Fifteen years ago that argument took the video industry the better part of a decade to accept. Agent teams don’t have a decade — but they do have every customer they’ve ever served, already telling them the answer.