Why does outcome-based pricing turn the evaluation suite into a revenue system? Support automation vendors now price per resolved conversation, at roughly 1 to 2 dollars per resolution, and some charge only when the agent finishes with no human handoff at all. The moment a vendor bills that way, the definition of a successful outcome becomes the definition of a billable event.

So the rubric that decides whether a ticket counts as resolved is no longer a quality artifact. It is the billing engine. An eval failure stops being a bug in a backlog and becomes an invoice that never goes out, or one the customer disputes. Whoever writes the rubric writes the revenue, which is a strange amount of power to leave with whoever happened to set up the test harness.

Pricing conversations used to end at the pricing page. Under outcome pricing they end inside the evaluation code, and most companies have not noticed the handoff.

What actually changes when you charge per outcome

Three things move at once. Revenue becomes a function of model performance, so a regression is a revenue event. Gross margin absorbs the variance, because the inference cost of a failed attempt is real and the revenue for it is zero. And the customer relationship inverts: the vendor now carries the risk that the technology does not work, which is exactly why buyers like it.

That risk transfer is the whole appeal. No outcome, no cost. It also means the vendor is running a business whose margin depends on a distribution it does not fully control, which is a genuinely different company than the one selling seats. Seats are sold once and paid monthly. Outcomes are earned continuously.

Who owns the definition of resolved?

This is the question nobody has a clean answer to. In a per-seat company, revenue recognition is a finance policy written in a document. In a per-outcome company, revenue recognition is a boolean returned by a system, and that boolean was specified by whoever built the eval set.

A ticket that ended with a polite deflection and no resolution. A task the agent completed on a wrong assumption. A conversation the user abandoned in frustration, which technically closed. Each of those is a judgment call, and each judgment call is worth money. I argued in Evals Are the New QA that the scenario library is the PRD of an AI product. Under outcome pricing it is also the rate card.

Which means the definition needs the same scrutiny a pricing decision gets. Product, finance, and the people who own the model all have standing in the room, and the version that survives is the one a customer would accept if they audited it line by line, because eventually one of them will.

The incentive problem is not hypothetical

Any metric that becomes a payment becomes a target. An agent optimized against a resolution definition will find the cheapest path to satisfying that definition, and so will the team tuning it, without anyone deciding to cheat. This is the oldest failure in growth measurement wearing new clothes, and I wrote the general version in Metrics That Move Teams.

The defense is a second metric the first one cannot fake. Resolution paired with reopen rate. Completion paired with whether the human redid the work an hour later. Deflection paired with satisfaction. A single billable metric with no counterweight will drift toward whatever is easiest to claim, and the drift will look like growth right up until renewal.

Why hybrid is winning

Pure outcome pricing gets the headlines and hybrid gets the adoption. A base platform fee plus variable consumption or outcome charges is the most common structure in 2026, and it is the sane transition. It keeps enough revenue predictable to run a company while the definition of the outcome is still being learned, and it protects the customer from a bill that scales with something they cannot forecast either.

The failure mode to avoid is picking the pricing model before you can measure the outcome. Plenty of companies announced outcome pricing and then discovered their instrumentation could not distinguish a resolved case from a closed one. That is the same production gap I described in Production Is an Org Problem, except now it is attached to an invoice.

What this means for a growth team

The pricing page and the evaluation suite are now the same document written in 2 languages, and somebody has to own both readings. That is a growth job, because it sits exactly where product decisions meet revenue, which is the argument I made in Pricing After Seats before the meter got this specific.

The practical starting point is smaller than a pricing migration. Write down what your product would charge for if it charged per outcome, then check whether your instrumentation could defend that number to a customer who disagreed. Most teams find out they are 1 definition and 2 events short. Finding that out before you announce the pricing is worth considerably more than finding out after.