
Scott Woody published a pricing essay on the Stripe blog on October 1, "Why I tried to kill token billing (and why we kept it)." Half of it is about token billing. The half I want to argue with is about output-based pricing, because it names a unit I have run through a billing system, and the unit did not behave the way the essay says it does.
The short version
Scott Woody argues on the Stripe blog that token billing is useful infrastructure and a bad customer-facing price, that outcome-based pricing is a myth for almost everyone because of a telemetry gap, and that the answer is output-based pricing: an objective, verifiable metric, like Fin's resolved support conversation, which he says you can objectively count. I priced per resolved ticket in a per-seat sunset. The definition (no escalation within 48 hours) held for 98% of cases, and the other 2% produced more than 200 contested cases a week, 600 disputes a month against 50 projected, and a dispute team that grew from 4 analysts to 14. Rewriting the definition to "no follow-up in 7 days" cost revenue and cut disputes by 60%. An output is countable. That does not make it agreed. The rewrite: price the unit whose disputes you can afford, and put the definition, the dispute window, and the arbitration path in the contract before the meter.
What Woody argues
The essay is short and worth reading in full. Here is what the argument needs.
Token billing, passthrough pricing with a markup, lets your costs determine your price. An invoice that shows which models ran, how many tokens they used, and what markup was applied defines your value as the spread over somebody else's model cost. As models get cheaper and more interchangeable, he writes, customers challenge the markup or route around you.
His better model is unified credits. Token-level metering stays on the back end for margin tuning, and the customer draws down a credit balance against operations the product performs. The invoice shows enrichments or generated images, not ten models and their markups.
Then the claim. He sorts pricing metrics into three. An input is a cost, like a token. An outcome is the customer making more money, building pipeline, or reducing churn. An output is an objective, verifiable metric, like generating an image or sending an email.
Outcome-based pricing, he says, is for all intents and purposes a myth for everyone without monopoly-like control of the telemetry or a contract large enough to pay for instrumenting the workflow. Sell a sales agent on revenue and you have invited a dispute: was it the agent, the rep, or marketing?
Outputs avoid that. Fin's price per resolved support conversation is the poster child for outcome pricing, and he argues it is really an output: the outcome of support is happy customers, and a resolved ticket is something you can objectively count. His advice to everyone else is to define an output metric, call it an outcome, and convince the market.
The unit he picked is the unit I priced
I am going to take the one example he chose.
In Field Report: What Broke When We Killed Our Per-Seat Tier, the priced unit was a resolved ticket. We defined resolved as "customer didn't escalate within 48 hours." It worked for 98% of cases, and it had held through the pilot.
The 2% that broke it were customers who reopened a ticket after 60 to 72 hours with a follow-up that was technically a new issue but related to the original. The system counted the original as resolved and billed it. Some customers contested. Our position was technically correct. Their position was that it was the same issue.
At low volume we handled those one by one. At sunset volume the contested cases passed 200 a week. Dispute volume in month 20 was 600 a month. We had projected 50. The dispute team went from 4 analysts to 9 to 14 in three months, and NPS dipped 8 points while they ramped.
Nothing in that story is a telemetry gap. Every event was logged. The count was exact. Both sides were reading the same number.
They disagreed about what the number meant.
Countable is not agreed
That is the property the output argument skips. An output is countable by the vendor. An invoice gets paid when the unit is agreed by the customer. Those are separate tests, and a unit can pass the first for months before it fails the second.
Moving from outcome to output does change the argument. It stops being "did your product cause this revenue," which is a few very large disputes a year. It becomes "did this one count," which is a small dispute on a small fraction of a very large volume. In my case the small fraction was 2%, and the volume turned it into more than 200 cases a week.
The other two redesigns in Per-Outcome Pricing: What Gets Clearer and What Gets Terrifying say the same thing from different sides. A qualified meeting at $40 is an output by his definition, and the word doing all the work is "qualified." An accepted code suggestion at $1.20 is as objective as a unit gets, and half the customers had a CFO meltdown anyway, because finance could not forecast against something that granular.
So I would not call outcome pricing a myth and output pricing the cure. They sit on one line. The further you move toward the customer's business result, the fewer and larger the arguments. The further you move toward your own activity, the smaller and more frequent they are. You are choosing a dispute profile either way.
The rewrite
"Define your output metric, call it an outcome, and convince the market" becomes this: price the unit whose disputes you can afford, and prove you can afford them before the meter is live.
Three things, in this order.
Write the definition from the customer's reading, not the system's. "No escalation in 48 hours" was our system's reading. "No follow-up in 7 days" was closer to theirs. We made that change at month 20, it cost some revenue, and it cut disputes by 60%. It would have been cheaper in month zero.
Put the dispute window and the arbitration path in the contract next to the definition. The five terms are the unit definition, the dispute window, the arbitration mechanism, a committed minimum, and a price ceiling. Pricing for AI Products has the full sequence.
Stress the definition before launch. Run it against 10x the dispute volume the pilot suggests, and staff dispute analysts to 5x. The edge case that is rare in a pilot is a weekly queue at scale.
On the token-billing half of the essay I have nothing to subtract. An invoice should not be a breakdown of the vendor's costs, and unified credits are a reasonable way to keep the metering behind the wall. That half is the argument running through SaaS to AI Business Models: the price has to move off the cost and onto what the customer got. My addition is only that "what the customer got" is a sentence two parties have to sign, and renaming it from outcome to output does not shorten the negotiation.
Related answer: Is outcome-based pricing a myth?
Sources: Why I tried to kill token billing (and why we kept it), Scott Woody, Stripe blog, October 1, 2026. Field Report: What Broke When We Killed Our Per-Seat Tier, falkster.com, May 11, 2026. Per-Outcome Pricing: What Gets Clearer and What Gets Terrifying, falkster.com, May 18, 2026.
Frequently asked
What is output-based pricing?+
In Scott Woody's 2026-10-01 essay on the Stripe blog, an output is an objective, verifiable metric like a generated image or a sent email, as opposed to an input (a cost, like a token) or an outcome (the customer making more money or reducing churn). He argues outputs are the best way to align AI consumption pricing with value today, and that Fin's price per resolved support conversation is an output that gets marketed as an outcome.
Why does Scott Woody call outcome-based pricing a myth?+
Because of what he calls a telemetry gap: what counts as the outcome, whether it occurred, and whether the product caused it. He says outcome pricing works only with monopoly-like control of the transaction and its telemetry, or with a ticket price high enough to pay for instrumenting the workflow end to end. His example is a sales agent priced on revenue, where the vendor invites a dispute over whether the agent, the rep, or marketing made the sale.
Is a resolved support ticket an objective unit to bill on?+
It is countable. It is not automatically agreed. In the per-seat sunset I ran, resolved meant no escalation within 48 hours, which held for 98% of cases. The 2% that reopened after 60 to 72 hours with a related follow-up produced more than 200 contested cases a week at scale, and disputes reached 600 a month against a projection of 50.
How do you fix a unit definition that customers keep disputing?+
Tighten it toward the customer's reading, and accept the revenue hit. We rewrote resolved ticket at month 20 to require no follow-up in 7 days. It cost some revenue and cut disputes by 60%. The better fix is before launch: test the definition at 10x the dispute volume you project, and staff dispute analysts to 5x.
What should be in the contract before an output or outcome meter goes live?+
Five things in writing: the unit definition, the dispute window, the arbitration mechanism, a committed minimum, and a price ceiling. The first three are the ones that decide whether a countable unit is also an agreed one.
Is token billing a good customer-facing pricing model?+
Woody's answer is no for almost every company that is not a model provider, and I would keep that part of his essay whole. An invoice that lists models, tokens, and markups defines the vendor's value as the spread over someone else's cost. He points to unified credits as the presentation layer: token-level metering stays on the back end, and the invoice shows the work the product performed.

Comments (0)
Sign in with LinkedIn to leave a comment.
Sign in with LinkedIn