
Vercel's inbound sales agent used to be a 1,000-line prompt. It is now 14 rules in code, with a model for the parts that need judgment and a second agent that decides which rule breaks are fine. I would copy the first two parts this week. The third is the one I would change before I did.
The short version
Tomasz Tunguz wrote up how Vercel COO Jeanne DeWitt Grosser's team built the agent that runs Vercel's inbound sales: a 125-line prompt by the best SDR, a summer of human review, then a prompt that grew to 1,000 lines until the team moved the deterministic parts into 14 rules. A second agent watches for breaches of those rules and "fixes it or blesses the exception." Grosser samples about one in 100 sends once the agent is competent. The rewrite is small: sample the sends, and have a person read every blessed exception, counted by rule. When I sunset a per-seat tier at a $40M ARR company, the plan modeled 50 exceptions a week and we got 600, and nobody was counting them. An approved rule break is the only place a rule tells you it is wrong.
What Vercel built
Tunguz's post is short and worth reading whole. I read it in full, and what follows is his account.
Vercel started a go-to-market engineering team and handed inbound to one engineer at roughly 20% of his time. The first agent was a prompt of about 125 lines, written by the best SDR on the team. It kicked off in June. For the first phase the agent researched, qualified, and drafted, and did not send. By August the human was out of the loop.
Grosser's description of that phase is the best short statement of agent QA I have read this month. Manage it like a new rep. Read 100% of the first hundred emails. Once the rep is generally competent, sample about one in 100.
Then the business changed. Vercel added product surface and moved upmarket, and the prompt grew to 1,000 lines. The model stopped following some of it. So the team divided the prompt into rules and judgment. Engineers encoded the rules. Inbound now runs on 14 of them. A lead arrives, the system checks Salesforce for an open opportunity on the account, and if there is one it goes to the account executive. In her words, the model should not get to decide.
Jan Oberhauser, n8n's CEO, drew the same line in his conversation with Aakash Gupta the day before: when you cannot rely on something working 95% of the time, you put it on a path you can be sure of, and you decide on the canvas where a human approves. Gupta's write-up turns that into a rule for moving a workflow over: human approval on any tool that sends, creates, updates, or deletes. It is the split I argued for in Your 'AI Agent' Is Probably Just a Cron Job, and it is good to see it with a company's name on it.
The part I would change
Because the rules are explicit, Vercel runs a second agent that watches for breaches. When the system breaks one of the 14, Tunguz writes, the escalation agent determines whether the breach should have occurred and "fixes it or blesses the exception."
Fixing is fine. Blessing is a decision that a rule did not apply to this lead. That is the most informative event in the whole system, and in this design a model makes it.
The post does not say whether those blessings are logged, counted, or read by a person. They may well be. I am not claiming a gap at Vercel. I am saying the diagram, as most teams will copy it, has an approval step with no approver.
What 600 exceptions taught me
I ran a rules system once where exceptions were granted quietly. It was a pricing migration, and I wrote it up in Percent Migrated Is a Vanity Metric.
We sunset a per-seat tier at a $40M ARR company. The plan modeled 50 disputes a week. We got 600. There were 200 contested cases a week on the unit definition alone, and three reps quietly negotiated unauthorized extensions. Every one of those accounts counted as migrated. The dashboard was accurate and useless.
What I took from it: an exception is the one artifact that only gets created when the model does not fit. Nobody files a carve-out for a rule they think works. And it was countable for one reason. Someone had to approve it, so it passed through a chokepoint with a name on it.
Hand the approval to an agent and the chokepoint is still there. The name is not. The count exists only if someone decides to keep it.
The rewrite
Keep the 14 rules. Keep the model for judgment. Keep one in 100 for sends. Add three lines.
- Every blessing is logged with the rule, the lead, and the reason. One row each. The reason is the escalation agent's own sentence, which is a claim and not a finding, the same way a completion report is in The Receipt Gate.
- A person reads all of them, every week. Not a sample. Sends are many and mostly fine, so sampling works. Blessings are few and each one is a disagreement between a rule and a real lead.
- A rule that keeps getting blessed is rewritten, or becomes rule 15. Count by rule. A blended number says something is off without saying where.
Two readings of the count, both from the migration. If it rises right after a rule change, that is good news. The rules are meeting real leads while a fix is one line. If it stays flat and high after the rule has been rewritten twice, the rule is not the problem. In my case it was the comp plan. In an inbound system I would look at what upstream of the rules changed: the form, the routing data, the definition of an open opportunity.
This is the same count I would want on any agent in Enterprise AI Agents, where the claim is that what matters is whether a deployed agent is still doing production work after ninety days. Fourteen rules written in August will not fit next August's leads. The blessings are where you find out which ones.
Pick one thing this week. If you run an agent with an escalation path, ask how many exceptions it approved last month, by rule. If nobody can say, that is the finding.
Related answer: Who should approve exceptions to an AI agent's rules?
Sources: Tomasz Tunguz, "How to Automate Inbound," tomtunguz.com, October 6, 2026, a write-up of his conversation with Jeanne DeWitt Grosser, COO of Vercel. Aakash Gupta with Jan Oberhauser, "n8n CEO: Why we aren't dead, how to build great AI products, and what I look for in an AI PM," Product Growth, October 5, 2026 (newsletter version). Percent Migrated Is a Vanity Metric, falkster.com, September 20, 2026, for the 50, 600, and 200 figures.
Also on Medium
Full archive →AI Agents and the Future of Work: A Pixar-Inspired Journey
What product managers can learn about AI agents from how Pixar runs a film team.
Many AI Agents Are Actually Workflows or Automations in Disguise
How to tell agents from workflows from cron jobs, and why it matters for what you ship.
Frequently asked
How did Vercel build its inbound sales agent?+
Per Tomasz Tunguz's write-up of his conversation with Vercel COO Jeanne DeWitt Grosser (October 6, 2026): one go-to-market engineer spending roughly 20% of his time, a first prompt of about 125 lines written by the best SDR on the team, and a human in the loop from June until August while the agent researched, qualified, and drafted without sending. Over the following year the prompt grew to about 1,000 lines, and the team split it into 14 deterministic rules in code, with the model kept for judgment.
Why move rules out of the prompt and into code?+
Grosser's reason, as quoted by Tunguz, is that the model will not always follow rules embedded in a long prompt, and a lot of qualification is a rule, not a judgment. Her example: when a lead arrives, a Salesforce lookup checks for an open opportunity on the account and routes it to the account executive. The model should not get to decide that.
What is a blessed exception in an agent system?+
In Vercel's design, a second agent watches for breaches of the 14 rules. When the system breaks one, that agent decides whether the breach should have occurred, and either fixes it or blesses the exception. A blessed exception is an approved rule break, and it is the only record that a rule did not fit a real case.
How often should a human review an agent's work?+
Two different rates. For routine output, Grosser's rule is to read 100% of the first hundred emails and then sample about one in 100. For blessed exceptions, I would read all of them every week. There are few, and each one marks a place where a rule and reality disagreed.
What does a rising exception count mean?+
Right after a rule change, a rising count is good: the rules are meeting real cases while changes are cheap. A count that stays high after the rule has been rewritten twice points at something outside the rule. In a pricing migration I ran, that something was the comp plan.
Where do the 50 and 600 figures come from?+
From sunsetting a per-seat pricing tier at a $40M ARR company. The plan modeled 50 disputes a week and the actual number was 600, alongside 200 contested cases a week on the unit definition and three reps negotiating unauthorized extensions. Every one of those accounts counted as migrated. The write-up is Percent Migrated Is a Vanity Metric on falkster.com.

Comments (0)
Sign in with LinkedIn to leave a comment.
Sign in with LinkedIn