I gave two talks at Dreamforce, a show where the stakeholder who matters most never came up

I spoke twice at Dreamforce this year. The first talk was about a category gap: agents are the new front door, and nothing in the stack was built to tell you whether they deliver. The second was a case study: an agent at the top of a self-serve funnel, measured first message to purchase.

Then I walked the floor for three days. Build, orchestrate, deploy, govern, trace, guardrail, evaluate. The tooling is extraordinary; half of it did not exist eighteen months ago. And I did not see one product, one session title, or one conversation about the stakeholder every one of those agents is pointed at.

The user.

Not the model. Not the developer. The person who decides whether any of this was worth their time.

Two talks, one argument

I started the first one by level-setting on something I think the market forgets. Agents are foundationally about automation, and automation is a promise about the machine. So the machine is what got built, and the machine is what filled that floor: functions, tools, orchestration, guardrails, evals. Every one of them answers the same question. Can the agent do the thing?

No one was talking about the key stakeholder in all of this. Automation is the technical half, and judging by that show floor, any AI team at any company can now automate a workflow. However, the promise underneath it was never made to the workflow. It was made to the person on the other end, and it is a harder promise to keep. The promise of fewer steps than the thing it replaced. Of less time and less repeating yourself. Of something they choose again rather than endure. Easy, efficient, pleasant. That is the user’s productivity, and it goes beyond the technical functionality of your agent. They are grading it against whatever they did before; searching the site, reading the docs, calling somebody, every session, and no one is even talking about how to instrument for that comparison. An agent that automates more while making the person work harder has just moved the work around and sent you a bill for it. And worse: all that engineering is worth nothing if the user would rather not deal with the agent at all. That is an existential blind spot.

Then I put one conversation on the screen. A shopper looking for wide-fit shoes.

On the left, her conversation as a modern stack saw it: 0.96 on helpfulness, 0.95 on relevancy, 0.94 on clarifying-question quality, ticket marked resolved. By every standard this industry uses, that is a success, and nobody would ever look at it again.

On the right, what she actually did: restated the ask four times, opened the catalogue in another tab because she had given up on the agent finding it, then quietly took less than she came in asking for. She said nothing. She left.

Nothing in the right column appears anywhere in the left. That is the whole show in one slide.

The second talk landed that argument in a real-world PLG project, where revenue is on the line and nobody, no rep and no CSM, rescues a bad thirty seconds. So the measurement has to start outside the agent: conversations, traces, site events and product events, joined on one identity and one timeline, the whole funnel from first message to upgrade. Engagement, experience, outcome. Everyone talks about the first column and the third. No one owns the middle, and it is what moves the third.

From there it is systematic. Start with the outcome you do not want, failed conversations, then find the signal under it carrying most of that failure, mid-intent abandonment. Slice until the worst cohort separates out: the people who repeated the same question three or more times. The cause names itself. The agent guessed at intent instead of asking, so ask once and never guess a default. Ship it, measure again, and close the loop where agents get built, which is why this Conviva customer wires their agents to us through MCP.

And the financial impact is massive, because that middle column is also the budget. Cost grows with roughly the square of the turn count, so nine turns instead of three is closer to an order of magnitude. Every wasted turn passed its eval and threw no error, so it never shows up as a failure. It shows up as usage. Your inference bill is an itemized invoice for your agent’s misunderstandings.

What I kept recognizing on that floor

My reaction walking the show was not that anyone had got it wrong. It was recognition. We have spent twenty years learning how to tell whether something was actually good for the person on the other end, and I was watching a market solve the first half of that problem with real skill while the second half had not surfaced yet.

Nobody chose this stack. Observability came out of the server room and answers whether the system is running. Evaluation came out of the release gate and answers whether an output is acceptable. Both are excellent at their jobs. They are simply not the right instruments for the question the user is keeping score on: was this worth what it cost me?

I understand the instinct to close that gap with a judge model. When the thing you are measuring is made of language, measuring it with a language model feels natural. It is not optimal, for a stack of reasons I wrote a paper about, linked below. The short version is that there is already a better judge sitting in every conversation, and it is not a model. It is the person: repeating themselves, rephrasing, opening a search tab mid-conversation, finishing the purchase somewhere else, not coming back. That is not an opinion. It is what happened.

Which brought me to the part that explains the whole floor. Every tool there starts at the agent, because the agent is the obvious place to start and the easy one. The code sits there and the trace is already written. But it is what you do with those traces, combined with clickstream data, that matters most. Quality-of-experience metrics, deterministically computed over the full population, with high-contrastive dimensionality, in real time. That is where most will break, and I say that as someone whose company had to build all of it. Sample five percent. Review these traces. The limits arrive dressed as methodology. Those are not chosen methods. They are the output of a stack that is not there yet.

Which is probably why we were standing there alone. Twenty years of quality-of-experience engineering, walking into a new market surrounded by young technology. We did not start at the agent. We started where the hard part is, because we had already built it once.

Building on our existing expertise and market dominance

Video started on delivery metrics. Clean dashboards, churning subscribers, until the subject of measurement moved from the system to the viewer, on every session, in real time, dimensional enough to name a cause rather than raise an alarm. We pioneered that shift, and we still lead the innovation in it. The failure mode here is identical: a functioning system, a dissatisfied customer, nothing between them.

Dreamforce proved this industry can build agents fast. The instruments to prove they worked for anybody were not in the building. An eval is an opinion. Behavior is evidence.

 

The full argument, including why judged and sampled signal cannot carry a production operation: Why Agent Operations Needs an Experience Layer, and How We Define What It Is.