FoundationNew·Falk Gottlob··6 min read

Three Token Bills, One Unit

The AI token bill got measured per task (SoL-Pi), per engineer (Larridin), and per company (GitLab) in one week. Product's unit is per outcome. Here is why.

token costgross marginSoL-PiNVIDIALarridinJason LemkinBasia KubickaSaaStrGitLabcost per outcomemodel routingfield notes
Helpful?

Foundation Falkster cover: three paper receipts of different lengths hanging side by side from one rail, with a single ruler laid straight across all three at the same height.

Three token bills landed this week, in three different units, from three people on my reading list. Each one is right for the person who asked. None of them is the product team's.

The short version

In one week the AI token bill was measured per task (NVIDIA's SoL-Pi, via Basia Kubicka: 49% fewer tokens on 51 benchmark tasks for 6% of score), per engineer (Larridin, via Jason Lemkin: median $213 a week, and a low-AI cohort flat at 1.9x output across a 20x spend range), and per company (GitLab: 400 basis points of gross margin). Product's unit is per outcome: one line per AI workflow with outcomes, tokens per outcome, cost per outcome, realized price per outcome, and gross margin. It is the only unit of the four that tells you whether to cut the tokens or leave them alone, because a per-task saving only counts on the tasks your workflow actually runs, a per-engineer bill says nothing about what shipped, and a per-company margin arrives a quarter late.

Per task

Basia Kubicka wrote up NVIDIA's SoL-Pi this morning. NVIDIA tested 152 ideas for making coding agents cheaper. Each one had to hold task quality inside a limit while cutting cost or tokens. Four passed, and they ship as an extension for the Pi coding agent: edit and test in the same call, park long tool output in a file and fetch it on demand, have a cheaper model read long logs and report with quotes that get checked against the real log, and compact history only when the math says the rewrite pays for itself.

On 51 EdgeBench tasks against plain Pi: 44.7% to 49.0% fewer tokens, about one third less spent on model calls, roughly 94% of Pi's average score. Everything off by default.

Her closing question is the right one to ask, and I want to answer it properly further down: would you trade 6% of benchmark score for one-third lower agent cost?

Per engineer

Jason Lemkin's Larridin piece landed yesterday. Larridin's first benchmark, from production billing and engineering telemetry over the four weeks ending August 2, says the median engineer bills $213 a week in AI coding tokens, about $920 a month. The 90th percentile bills $911 a week. A spread of more than 10x, and both are floors, because flat-fee plans hide what gets consumed. Annualize the top and you are near $47,000 a year per engineer in tokens alone.

The finding I keep coming back to is the cohort split. Same company, same tools, same prices, same starting spend of about $170 a week. Deeply AI-native engineers reached 11.8x output at around $1,300 a week and had not hit a ceiling. Partial adopters saw marginal payoff halve past about $600 a week. And the low-AI cohort topped out around 1.9x and stayed flat across a 20x spend range, from $21 to $421 a week.

Extra dollars bought activity. No additional output.

Larridin is careful about this. Output is not lines of code, it is merged PRs scored for complexity, discounted for missing tests, scaled by churn, and the relationships are associational. Lemkin's read: there is no universal ideal budget, track your own curve. Fair. But notice what the unit is. Dollars per engineer per week. It cannot see what any of those dollars shipped.

Per company

GitLab, last week, in Lemkin's teardown of the quarter. Subscription cost of revenue up 76% on revenue up 21%, non-GAAP gross margin down to 86% from 90% a year earlier, the biggest single drop in the quarter GitLab started selling agent consumption. Lemkin calls the AI cost line an architecture decision that becomes an accounting one. I would put it differently, and the point here is small: a gross margin number arrives a quarter after the decisions that made it, and it arrives at the company level, where nobody can see which workflow did it.

Product's unit

Per outcome.

One line per AI-driven workflow. Outcomes delivered this month. Average tokens per outcome. Average cost per outcome. Realized price per outcome. Gross margin. Red dot on any line where margin is negative or declining. That is the line in The CPO Mandate 2026: What Boards Expect From Product, and it has been there since April.

It survives the other three because each of them goes blind somewhere.

Per task goes blind on mix. When Spotify published a 90% token saving, I wrote The Enforceable Half about keeping two claims apart: 90% on bulk reads is not 90% off the bill, because your bill depends on how much of your workload is bulk reads. SoL-Pi's 49% is the same shape. It is 49% on 51 EdgeBench tasks. Yours is whatever share of your outcomes look like those tasks, and you only know that share if you count outcomes.

Which is the honest answer to Basia's question. Six percent of score for a third of cost is not one trade. It is a trade per workflow. Where the outcome tolerates a weaker score, a classification step, a draft that a human edits anyway, take it. Where it does not, refuse it. That is the whole job of the routing table per surface in Gross Margin Is Your Job Now, and the per-outcome line is how you know which surfaces are which.

Per engineer goes blind on output. On Larridin's flat cohort, spend moved 20x and output did not move. Cut that cohort's tokens by a third and you have saved a third of money that was producing nothing. The lever there is not the harness. It is fluency, and no token dashboard will tell you that. An outcome count will, because it is the outcomes that stayed flat.

Per company goes blind on time and on which workflow. A quarter late, and aggregated past the point anyone can act.

What to do this week

Pick your most-used AI workflow. Write its five numbers on one line. Outcomes, tokens per outcome, cost per outcome, price per outcome, margin.

If you cannot fill in the first one, stop. The other four are not a bill. They are a guess, and every cost cut you make against a guess is Larridin's flat cohort: a third off money that may have been producing nothing, or a third off the one workflow that was carrying the margin.

This one belongs to the running argument on SaaS to AI Business Models: the pricing model moves to the outcome, and so does the cost model, because the two have to share a denominator or neither one means anything.

Related answer: What unit should product use to measure AI token cost?

Sources: Basia Kubicka, "NVIDIA's agent harness SoL-Pi uses 49% fewer tokens," LinkedIn, September 28, 2026. SaaStr AI App of the Week: Larridin. The Median Engineer Now Bills $213 a Week in AI Coding Tokens, Jason Lemkin, SaaStr, September 27, 2026. 5 Interesting Learnings from GitLab at $1.13 Billion in Revenue, Jason Lemkin, SaaStr, September 25, 2026.

Share this post

Frequently asked

What did NVIDIA's SoL-Pi harness measure?+

As Basia Kubicka reported it, NVIDIA tested 152 ideas for making coding agents cheaper and shipped the 4 that passed as SoL-Pi, an extension for the Pi coding agent. On 51 EdgeBench tasks against plain Pi it sent and received 44.7% to 49.0% fewer tokens, spent about one third less on model calls, and kept roughly 94% of Pi's average score. Every mechanism is off by default. That is a per-task number, on those 51 tasks.

What does the Larridin benchmark say the median engineer spends on AI coding tokens?+

Per Jason Lemkin's SaaStr write-up of Larridin's August benchmark, the median engineer bills $213 a week, roughly $920 a month, and the 90th percentile bills $911 a week, a spread of more than 10x. Both are floors because flat-fee plans hide consumption. The sharper finding is that a low-AI cohort topped out around 1.9x output and stayed flat across a 20x spend range, $21 to $421 a week.

What unit should product use for AI cost?+

Per outcome. One line per AI-driven workflow with five numbers: outcomes delivered this month, average tokens per outcome, average cost per outcome, realized price per outcome, and gross margin, with a red dot on any line where margin is negative or declining. It is the line in my 2026 CPO board deck and the only unit that says whether to cut tokens or leave them alone.

Should you trade 6% of benchmark score for a third off agent cost?+

Not as one trade. It is a trade per workflow. Where the outcome tolerates a weaker score, take it. Where it does not, refuse it. That is what a routing table per surface is for, and it is the same reason Spotify's 90% token saving on bulk reads was never 90% off anyone's bill.

Why doesn't a per-engineer token bill tell you what to cut?+

Because it says nothing about what shipped. On Larridin's low-AI cohort, spend moved 20x and output did not move, so cutting that cohort's tokens by a third saves a third of money that was producing nothing. The lever there is fluency, not the harness.

THE SHORT ANSWER

PART OF

SaaS to AI Business Models

About the author

Falk Gottlob

Falk Gottlob

Product Executive · Founder, Falkster.AI

Thirty years shipping product, from Microsoft Research and Adobe to Salesforce, where he grew Quip into what became Slack Canvas. Four startups, five exits, including a $6.5B healthcare platform and a company Microsoft bought. Four-time Chief Product Officer. Now founder of Falkster.AI, an agentic AI company run by its own agents. This notebook is written from inside the build, not above it.

Comments (0)

Sign in with LinkedIn to leave a comment.

Sign in with LinkedIn
  • Be the first to comment.

Keep Reading

Posts you might find interesting based on what you just read.