
Jason Lemkin put real numbers on support agents on October 6. He wrote up a benchmark Gorgias runs on 13 support AI vendors: the best agents fully resolve about 70% of conversations, and the median vendor resolves 48%.
The more useful part is a paragraph further down, about what "resolved" means. I've priced that word. It has at least three definitions, and each one gives your agent a different rate.
The short version
Per Jason Lemkin's SaaStr write-up, Gorgias tested 13 vendors on 212 live ecommerce stores and found the top five resolve 64% to 75% of conversations, with a median of 48%. The benchmark counts a resolution only with zero human touch and no push out of the channel. Gorgias's own customer reporting, Lemkin notes, counts one once 72 hours pass without a human agent helping, so a customer who gave up counts. Resolved is also the billing unit, at $0.90 each. I priced per resolved ticket with a 48-hour definition that held for 98% of cases, and the other 2% became more than 200 contested cases a week. So before a renewal, take 30 conversations your agent closed last month and score each three ways: zero touch, the invoice definition, and no related follow-up in seven days. Report three rates. Re-run it monthly, because nine of ten vendors in the benchmark lost automation in four weeks.
What the benchmark found
I read Lemkin's post in full. I haven't opened the benchmark itself, so every figure here is as he reports it. He also flags the obvious caveat himself: Gorgias runs the benchmark, SaaStr Fund led Gorgias's seed, and Gorgias ranks fifth on its own board.
The setup is the good part. 212 live mid-market ecommerce stores carrying real inventory. Every agent gets the same customer messages. A judge scores each response blind to the vendor against 26 binary checks, and factual claims like price, policy, and SKU are checked against the live store.
The results, pooled over four weeks:
- The top five vendors resolve 64% to 75%. The median is 48%. Weighted by volume the field resolves 46%.
- Automation and quality come apart. Ada resolves 68% with a blind quality score of 39 out of 100. Sierra resolves 48% and ties for the top quality score at 72.
- Nine of the ten vendors with enough history lost automation over the last four weeks, by 4 to 23 points. The benchmark doesn't say why.
- Policy questions score best (returns policy at 77.4). Order tracking, usually the biggest ticket category for an ecommerce brand, scores 60.1.
Then the paragraph I'd tape to the wall. The benchmark counts a conversation as automated only when the AI handled it with zero human touch, no handover, and no deflection out of the channel. "Email us" counts against the agent. In Gorgias's own customer-facing reporting, Lemkin writes, an interaction counts as automated once the AI resolves it and 72 hours pass without a human agent helping. His next two sentences: "A customer who gave up and never came back counts as a resolution there. In the benchmark, that customer counts as a failure."
And resolved is the pricing unit. Gorgias AI Agent charges $0.90 per resolved conversation. Lemkin's advice is to get each vendor's definition of a billable resolution in writing before you sign.
Do that. Then go one step further, because having it in writing didn't save me.
I priced that unit
The full story is in Field Report: What Broke When We Killed Our Per-Seat Tier, and I used it on October 2 in The Output Gets Disputed Too to argue about pricing. Here it's the measurement that matters.
We priced per resolved ticket. Resolved meant the customer didn't escalate within 48 hours. That was the written definition, and it held for 98% of cases, pilot included.
The other 2% were customers who came back after 60 to 72 hours with a follow-up. To us it was technically a new issue. To them it was the same one. The system had already counted the original as resolved and billed it. At sunset volume those contested cases passed 200 a week, and disputes hit 600 a month against a projection of 50.
At month 20 we rewrote the definition to no follow-up in 7 days. It cost some revenue and cut disputes by 60%.
Look at where my 2% landed. 60 to 72 hours. A 48-hour quiet window missed all of them, and a 72-hour window sits right on the edge. I'm not saying anything about Gorgias's customers here. I'm saying a quiet window is a guess about how long your customers take to notice the answer was wrong, and mine took about three days.
So that's three definitions of the same word. The benchmark's. The invoice's. And the one I got pushed into.
The drop: one sheet, three columns
Lemkin recommends a cold test: take 30 real tickets and count full resolutions yourself. Run that on your own agent, with three columns.
- Pull 30 conversations your agent closed last month. Closed by the agent's own count, picked by you and not by the vendor. Weight them toward your real ticket mix. If half your volume is order tracking, half the sheet is order tracking.
- Column one, zero touch. No human joined. No handover. No "email us," no contact form, no "call us." This is the benchmark's rule, and it's strict on purpose.
- Column two, the invoice. Did this conversation count as a billable resolution under your contract? Use the vendor's definition exactly as written, quiet window and all.
- Column three, seven days. Did that customer come back within a week about the same thing, on any channel? Check email and phone too. If they did, it's a no.
- Report three rates, side by side. Not an average.
Then read the gaps. Between column one and column two is the quiet. Some of it is people who got their answer and left. Some of it is people who gave up, and from a transcript you often can't tell which. That's what column three is for. The gap between two and three is the conversations you paid for that didn't stay resolved, and those are the ones you'll end up disputing.
Thirty rows is an afternoon for one person. It won't give you a statistically tidy number. It'll show you whether the three rates are close together or far apart, which you need to know before the renewal call.
Put it on the calendar monthly. Nine of ten vendors lost ground in four weeks in that benchmark, and Lemkin says SaaStr re-checks every one of its own 20-plus agents on a schedule, including the ones that seem to be working. That's the argument of Enterprise AI Agents in one line: what matters is whether the agent still completes the work after ninety days, and what each successful outcome costs. You can't price a successful outcome until you've picked which of the three you mean.
One thing to try this week: find the sentence in your contract that defines a billable resolution, and write the number of hours in its quiet window at the top of the sheet. If nobody on your team can find the sentence, you've learned something already.
Related answer: How do you measure an AI support agent's resolution rate?
Sources: Jason Lemkin, "How Much of Your Customer Support Can AI Really Resolve? The Best Get About 70%. The Median Is 48%. Here's the Real Data From 13 Vendors," SaaStr, October 6, 2026, reporting the public benchmark run by Gorgias. Field Report: What Broke When We Killed Our Per-Seat Tier, falkster.com.
Also on Medium
Full archive →AI Agents and the Future of Work: A Pixar-Inspired Journey
What product managers can learn about AI agents from how Pixar runs a film team.
Many AI Agents Are Actually Workflows or Automations in Disguise
How to tell agents from workflows from cron jobs, and why it matters for what you ship.
Frequently asked
How much customer support can AI agents resolve today?+
In the public benchmark Gorgias runs, as written up by Jason Lemkin at SaaStr on October 6, 2026, the top five of 13 vendors resolve 64% to 75% of conversations with no human involved, the median vendor resolves 48%, and the field resolves 46% weighted by conversation volume. The test runs on 212 live mid-market ecommerce stores, with a judge blind to the vendor scoring each response against 26 binary checks.
Why do resolution rates differ between a benchmark and a vendor dashboard?+
The definition. The benchmark counts a conversation as resolved only with zero human touch, no handover, and no deflection out of the channel. Per Lemkin, Gorgias's own customer-facing reporting counts an interaction as automated once the AI resolves it and 72 hours pass without a human agent helping, so a customer who gave up and never came back counts as a resolution there and as a failure in the benchmark.
What are the three resolution rates for an AI support agent?+
Zero touch: no human, no handover, and no push to email or a form. The invoice rate: whatever your contract defines as a billable resolution, usually a quiet window measured in hours. And the seven-day rate: no related follow-up from that customer within a week on any channel. Score the same 30 closed conversations all three ways and report the rates side by side.
Why seven days?+
From a definition I had to rewrite. In a per-seat sunset I ran, we priced per resolved ticket and defined resolved as no escalation within 48 hours. That held for 98% of cases. The other 2% reopened after 60 to 72 hours with a related follow-up, and at volume that produced more than 200 contested cases a week. At month 20 we changed the definition to no follow-up in 7 days. It cost some revenue and cut disputes by 60%.
How often should you re-test a support agent's resolution rate?+
Monthly. In the benchmark Lemkin reports, nine of the ten vendors with enough history lost automation over the last four weeks, with declines from 4 to 23 points, and nine of ten also declined on quality. The benchmark does not explain the declines. The rate you saw in the pilot is not the rate you have now.
Does a high automation rate mean a good support agent?+
No. In the same benchmark Ada resolved 68% of conversations with a blind quality score of 39 out of 100, and Sierra resolved 48% with a quality score of 72. A conversation closed with the wrong return window creates a second contact. That is why the third column, no related follow-up in seven days, belongs next to the first two.

Comments (0)
Sign in with LinkedIn to leave a comment.
Sign in with LinkedIn