What unit should product use to measure AI token cost?

THE SHORT ANSWER

Per outcome. One line per AI-driven workflow with five numbers: outcomes delivered this month, tokens per outcome, cost per outcome, realized price per outcome, and gross margin. A per-task saving like SoL-Pi's 49% only counts on the tasks your workflow runs, a per-engineer bill like Larridin's $213 a week median says nothing about what shipped, and a per-company margin like GitLab's arrives a quarter late. Only the per-outcome line says whether to cut the tokens or leave them alone.

Per outcome. Everything else goes blind somewhere.

The line

One line per AI-driven workflow, five numbers: outcomes delivered this month, average tokens per outcome, average cost per outcome, realized price per outcome, gross margin. Red dot on any line where margin is negative or declining. It has been in The CPO Mandate 2026 since April.

Where the other units go blind

Per task goes blind on mix. NVIDIA's SoL-Pi, as Basia Kubicka wrote it up, cut tokens 44.7% to 49% on 51 EdgeBench tasks for roughly 6% of score. Spotify's shunt cut 90% on bulk reads. Neither is a number off your bill; your bill moves by whatever share of your outcomes look like those tasks. The Enforceable Half has the Spotify version of this caution.

Per engineer goes blind on output. Larridin's benchmark, per Jason Lemkin, puts the median engineer at $213 a week and the 90th percentile at $911. The finding that matters is the low-AI cohort: flat at about 1.9x output while spend moved 20x. Cut that cohort's tokens by a third and you saved a third of money that was producing nothing. The lever is fluency, not the harness.

Per company goes blind on time and on which workflow. GitLab's 400 basis points of gross margin arrived a quarter after the decisions, aggregated past anyone's ability to act.

The trade

Six percent of score for a third of cost is not one decision. It is one per workflow, taken where the outcome tolerates it and refused where it does not, and the per-outcome line is how you tell the two apart.

The full argument is in Three Token Bills, One Unit, part of SaaS to AI Business Models.

SOURCES

THE LONG VERSION

RELATED ANSWERS

Last reviewed 2026-09-28 · 2 min read