
Sierra is the fastest-growing enterprise software company of this cycle, and almost everyone explains it wrong. The usual story is founder gravity: Bret Taylor co-created Google Maps, ran Salesforce as co-CEO, chairs the OpenAI board, so of course the Fortune 50 buys from him. The gravity is real and it explains the first meeting. It does not explain $100M in ARR in seven quarters, $150M in eight, roughly $200M by May 2026, and a $15.8 billion Series E, in a market crowded with well-funded competitors selling what sounds like the same thing.
I spent time going through Sierra's numbers, Taylor's long-form interviews, and their engineering writing, because I wanted the mechanism, not the mythology. What I found is not one trick. It is six deliberate operating choices that reinforce each other, and nearly every one of them is stealable by companies that will never raise a dollar at a $10 billion valuation. This is the deepest case study I know of for the thesis I have been writing all year: the constraint moved from building to landing, and the companies winning right now are the ones that restructured their entire business around that fact.
The short version
Sierra sells AI agents for customer experience, launched February 2024, and hit $100M ARR in seven quarters, faster than almost any software company on record, reaching roughly $200M by May 2026 with 40% of the Fortune 50 as customers. The mechanism is six choices, not one. They price the outcome: a pre-negotiated rate per resolved case, escalations to humans free, which makes Sierra's revenue depend on the product actually working. They own the last mile with forward-deployed agent engineers, deliberately absorbing the implementation risk every SaaS vendor spent two decades pushing onto the customer. They invented an engineering discipline for non-deterministic software, where annotated real conversations become regression tests so an agent never makes the same mistake twice. They build capabilities ahead of the models and throw the code away without sentiment when models catch up. They land one high-volume channel at a giant brand, prove resolution and CSAT, then expand. And they market with credibility artifacts, public benchmarks and engineering essays, instead of category slogans. Each choice is stealable. The compounding of all six is the moat.
First, the numbers, so we agree this is real
Sierra was founded in 2023 by Taylor and Clay Bavor, a twenty-year Google veteran, and launched publicly on February 13, 2024. Sacra estimates ARR at $26M at the end of 2024, around $130M at the end of 2025, and $200M by May 2026. Taylor's own account in a March 2026 interview: "We reached $100 million in ARR in seven quarters, $150 in eight quarters." The funding curve tracked it: $110M from Sequoia and Benchmark at roughly $1B, $175M at $4.5B in October 2024, $350M at $10B in September 2025, and a $950M Series E at $15.8B in May 2026 led by GV and Tiger Global.
The customer list is the more telling number. Roughly a third of Sierra's clients have over $10 billion in revenue. Named accounts include Cigna, Nordstrom, ADT, Chime, SiriusXM, Sonos, WeightWatchers, Nubank, Ramp, Rivian, Rocket Mortgage, Sutter Health, Singtel, and Wayfair. The results they publish are specific: SoFi's net promoter score up 33 points, Ramp automating 90% of support cases, typical automation between 70 and 90%, CSAT above 4.5 out of 5. Voice, which Sierra launched later, overtook text as the primary channel within about a year.
Hold that list against the competitive field. Decagon, founded the same year, went after internet-native logos like Duolingo and Notion. Zendesk, Intercom, and Salesforce are bolting agents onto installed bases. Sierra went after the hardest customers in the market, the ones with procurement departments, compliance regimes, and hundred-million-call volumes, and won them first. That was not an accident of connections. It was a sequence of choices. Here they are.
Choice 1: Price the outcome, so the revenue depends on the landing
This is the foundation everything else stands on, and it is worth being precise about the mechanics because "outcome-based pricing" gets used loosely.
Sierra's model, in Taylor's words: "If the AI agent resolves the case, no human intervention, there's a pre-negotiated rate for that. If we do have to escalate to a person, that's free." Reported rates are around $1.50 per resolution in support contexts. Contracts are custom, typically starting around $150,000 a year, and blend volume-based pricing for routine interactions with pay-per-resolution for complex ones. There is no self-serve tier.
Two details in that design are load-bearing, and both are commonly missed.
First, escalations are free. That is not a discount, it is an incentive structure. Every failed resolution costs Sierra the revenue and still costs Sierra the compute. The company loses money on its own product's failures, in direct proportion to how often it fails. I have never seen a cleaner mechanical answer to the question I asked in who owns landing: who is accountable for whether the thing actually works? At Sierra, the P&L is.
Second, Taylor explicitly separates outcomes from usage, and his argument deserves to be quoted because most of the industry conflates the two: "If you have an AI agent that is making sales for Stripe to small businesses, and I told you I will sell one-tenth the number of new Stripe GMV but I'll use one one-hundredth of the tokens, you probably wouldn't care. You care about the value to your top line. There's not a strong correlation between token usage and value." Usage-based pricing is charging for storage. Outcome-based pricing is charging for the thing the customer actually wanted. The difference matters operationally: under outcomes, "reducing your token utilization for the same outcomes is your problem, not your customer's," which turns cost optimization into Sierra's own margin engine rather than a customer negotiation.
The deeper effect is cultural, and Taylor names it directly: the separation between software and implementation, where the vendor ships and the customer's failure to adopt is the customer's problem, is "the essence of so many of the problems in the software industry." Outcome pricing deletes that separation. "You become more accountable to help them be successful because until they do, they can't use it," and there is "a strong incentive for the software company to have skin in the game to navigate that last mile."
This is the Snowflake move, executed one level deeper. Snowflake denominated revenue in usage, which forces landing. Sierra denominated it in results, which forces landing and quality at once.
How to steal it. You do not need to convert your whole price list. Find one workflow in your product with a countable, attributable outcome, a resolved ticket, a completed onboarding, a recovered payment, and offer one contract where a meaningful slice of the price is contingent on it. Watch what happens inside your own company: the moment revenue depends on the outcome, every internal argument about whose job adoption is dissolves, because the answer became everyone's. Be honest about the boundary, though. Taylor concedes there is "not a great way to do it for every type of agent," and falls back to usage where outcomes are fuzzy. Do not fake an outcome metric where none exists; a gameable outcome is worse than an honest usage meter.
Choice 2: Own the last mile, on purpose
The standard SaaS playbook of the last two decades treated high-touch implementation as a margin disease. You productized, you documented, you handed off to systems integrators, and you booked the license revenue whether or not the software ever went live. Sierra rebuilt the opposite model on purpose: Sacra describes it as a "productized BPO," and Sierra's own hiring materials describe agent engineers who own "the full lifecycle of agents in production: building them, deploying them, and ensuring they deliver measurable outcomes."
That is a forward-deployed model. Sierra engineers sit with the customer's systems, wire the integrations into the CRM and billing and ERP, and stay accountable for the resolution rate after go-live, because, per choice 1, Sierra does not get paid otherwise. My favorite illustration is Taylor's three-CRM story: a client that had acquired three companies had three identity systems and was planning the usual multi-year unification project before the agent could launch. Taylor's response: "Why don't you just have the AI agent go in all three of them and just think? What does a person do? They think about it. Let's just do that." The agent shipped against the messy reality instead of waiting for the clean one. That is what owning the last mile looks like in practice: the vendor absorbs the customer's mess instead of writing a prerequisites document about it.
Notice what this does to the classic objection. Sacra flags implementation intensity as Sierra's biggest scaling risk, and the concern is fair: high-touch deployment compresses margins and gates growth on engineer hiring. Sierra's answer is to productize the last mile from the inside. Agent Studio 2.0 moved workflow-building to a no-code interface that customer-experience teams operate themselves. Ghostwriter, their agent-building agent, generates production-ready agents from SOPs, transcripts, and process docs. The sequence matters enormously: they did the work manually with forward-deployed engineers first, learned what the last mile actually contains, then built products to compress it. Most companies productize first, based on a guess, and then wonder why the product misses the mess.
How to steal it. Take your most recent enterprise deal and write down every step between signature and the customer's first realized outcome. Every integration, every data cleanup, every approval, every workflow decision. That list is your last mile, and today most of it is the customer's unpaid, unowned job, which is exactly why your launches do not land. Assign your own people to the three hardest steps for your next three deals, eat the cost, and log everything they do. Six months later, that log is your product roadmap for compressing the last mile, and it will be more accurate than any discovery interview you have ever run.
Choice 3: Invent the engineering discipline the product actually needs
This is the part of the Sierra story that product and engineering leaders should study line by line, because it is the most complete public answer to a question I have been hammering all year: how do you ship non-deterministic software responsibly?
Sierra's framing, from the Agent Development Life Cycle essay Taylor and Bavor published in mid-2024: agents break the software development lifecycle. Traditional software is deterministic, structured, fast, and cheap. Agents are goal-based, non-deterministic, conversational, slow, expensive, and, worst of all, have no change management story, because a model upgrade silently changes behavior. Their response is a four-part discipline.
Declarative goals and guardrails. Developers express what the agent should achieve and the hard limits it cannot cross, "orders can only be returned within 30 days," rather than scripting steps. The agent gets creativity in the middle and determinism at the edges that matter, and because the declaration is abstracted from the underlying LLM, new models slot in without rewriting the agent.
Immutable releases. Every agent release snapshots everything that shapes behavior: code, prompts, model versions, and the knowledge base, atomically. Roll back instantly, A/B test releases against business goals. This sounds like plumbing until you remember most teams shipping on LLMs today cannot even reproduce yesterday's behavior, let alone roll back to it.
Structured human QA. Customer-experience staff, the people who actually know the rules of the business, annotate samples of real conversations every day in Sierra's Experience Manager. Not engineers guessing at quality. Subject-matter experts grading it, continuously.
Conversations become regression tests. This is the masterstroke. Every annotated failure becomes a simulated conversation test against mock APIs, reproducible and runnable in parallel. Thousands of annotations become thousands of regression tests, so "your AI agent never makes the same mistake twice." And before Sierra upgrades its own platform, it runs every customer's regression suite. Model upgrades stop being a leap of faith and become a gated release.
On top of this sits the runtime reliability technique Taylor calls a "constellation of models": supervisor agents observe the reasoning of the primary agent and send it back with notes when it violates a policy. His math: chain a 90% reliable reasoner with a 90% reliable supervisor and you get to 99%. The customer never sees any of this machinery; they express goals and guardrails, and Sierra absorbs the complexity.
Readers of this site will recognize the shape: this is the eval is the spec, industrialized. The annotated-conversation-to-regression-test loop is exactly the discipline that caught what my demo missed, running at enterprise scale with the customer's own staff as the graders.
How to steal it. Three moves, in order of return. First, make releases reproducible: version your prompts, pin your model versions, snapshot your knowledge sources, and do not ship an agent you cannot roll back. Second, put your domain experts, not your engineers, in a daily annotation loop over real outputs, even if the tooling is a spreadsheet. Third, and this is the one almost nobody does, convert every annotated failure into a permanent regression test before you fix it. The fix without the test is how the same failure ships three times. If you run this loop for a quarter, you will have the beginning of what Sierra has: an asset that compounds, because every mistake becomes infrastructure.
Choice 4: Build ahead of the models, and throw the code away
Every applied-AI company faces the same strategic question: do you wait for the frontier models to make your feature possible, or build scaffolding now that the models will eventually obsolete? Sierra's answer is unambiguous and unsentimental. Build it now, know it is temporary, and throw it away without grief.
Taylor: "So much of what we write, we plan to throw out later. It's a very weird way to build a company." Sierra built chain-of-thought orchestration before OpenAI's o1 made it native, then deleted it. When a large Hong Kong bank needed Cantonese voice and no model did it well, Sierra built what Taylor calls the best Cantonese support on the market, while stating plainly it "will certainly be commoditized in three years." The principle underneath: "You can't have the luxury of waiting for all the models to catch up with your aspirations. But you know they will. You have to have the best technology and be comfortable with throwing it out. It's a momentum and pace of innovation game rather than thinking of this as precious intellectual property."
And the warning, which I would frame and hang in most product orgs: "Teams that start to treat the code that they wrote as precious, that has been obviated by a general-purpose AI model, will fundamentally fall behind."
This inverts how most companies account for engineering work. The scaffolding Sierra throws away was not waste; it won the Hong Kong bank, and the deal outlives the code. The asset is the customer relationship and the pace, not the artifact. This is the repricing logic applied from the inside: Sierra treats its own effort-denominated assets as depreciating, because they are, and refuses to let the depreciation schedule of its code dictate its strategy. Companies that got repriced this cycle, Chegg, Stack Overflow, were the ones that mistook a depreciating stock of past work for a durable moat. Sierra runs the opposite book.
How to steal it. Do the two-column exercise on your own codebase: what you have built that a frontier model will likely absorb within 24 months, versus what it cannot, your integrations, your customer's trust, your regression suites, your data. Budget the first column as opex with an expiry date and the second as the actual balance sheet. And when a model release obviates something your team is proud of, celebrate the deletion publicly. The teams that mourn are the teams that slow down.
Choice 5: Land one channel at the biggest brand that will say yes
Sierra's go-to-market inverts the standard startup wisdom of starting with easy logos. They went straight at enterprises with billion-dollar revenues, and the wedge is narrower than the ambition: land one channel, a few use cases, prove it, expand.
Taylor describes the pattern concretely. A healthcare company starts with a few types of phone calls: "Let's have the AI agent take them and see how it does. Do people like it? Does it lower our cost? Usually, it's customer satisfaction." A car insurer starts with first notice of loss, the fender-bender call. Digitally-native companies start with chat. Then the expansion: "Almost all of our clients will do both," phone and chat, and the act of unifying them collapses a wall inside the customer's own org, because the call center and digital teams were separate departments until the agent digitized "the last remaining analog channel, which is the telephone."
Three structural details make this motion work. The buyer is deliberately different from the model-vendor's buyer: Sierra sells to the CFO, the Chief Customer Officer, the Chief Digital Officer, which Taylor argues protects applied-AI companies from the labs, because "software companies orient around individual buyers within companies." The proof is empirical and immediate, resolution rate and CSAT within weeks, which matters because, in his words, today's clients buy "because our technology works," and only later will they buy because of product breadth. And social proof ladders within verticals: "You want to be maybe the first healthcare insurance company to adopt. There's another insurer who says, I want to be the fifth." Winning Cigna is not one logo. It is the license to win the next four insurers who were waiting for someone else to go first.
The partner motion follows the same grain. Blue Shield of California came through Stellarus, a spun-out services partner. R1 brings Sierra into revenue-cycle management, where Sierra-powered payer agents and provider agents literally call each other, English over the public telephone network, which Taylor gleefully describes as the current state of agent-to-agent interop. Japan came via a SoftBank Vision Fund 2 strategic investment plus the acquisition of Opera Tech in Tokyo; France via acquiring Fragment in Paris. And the Salesforce-shaped scaling hire, Eric Eyken-Sluyters from Agentforce as President of Field Operations, tells you they are building a classic enterprise field machine under the novel pricing.
How to steal it. Pick your wedge by volume and measurability, not by strategic importance. The fender-bender call is not glamorous; it is frequent, bounded, and instantly measurable, which is why it converts a skeptical enterprise. Write down the one workflow in your target account that runs a thousand times a week with a countable success condition, and refuse to pitch anything wider until that number is public inside the account. Then price the expansion into the relationship rather than the deal: Sierra's median customer grows by adding channels and use cases, which under outcome pricing means revenue grows exactly as fast as trust does.
Choice 6: Market with credibility artifacts, not category slogans
Sierra's marketing budget, as far as I can tell from the outside, goes into things that would survive peer review. They published τ-bench, a public benchmark for evaluating agents on real-world tasks, and followed it with τ-voice for real-time voice agents, 278 grounded customer-service tasks with deterministic end-to-end scoring. They publish engineering doctrine, the Agent Development Life Cycle essay is a genuine contribution to the field, not content marketing wearing a lab coat. Their case studies lead with auditable numbers: SoFi NPS up 33 points, Ramp at 90% automation, Redfin users viewing nearly twice as many listings and 47% more likely to request tours. And the founders carry a narrative with actual explanatory content: Taylor's "the atomic unit of productivity in AI is a process, not a person," and his line that companies "ship our org charts," are ideas people repeat in meetings Sierra will never attend.
Contrast this with the dominant mode of AI marketing right now, which is category invention and superlative inflation. Sierra rarely argues it is the best. It publishes the benchmark and the discipline and lets the Fortune 50 logos argue for it. For a company selling trust to CFOs, the medium is the message: everything they publish is itself evidence of the operational seriousness the product claims.
How to steal it. Replace one piece of planned category-marketing content this quarter with a credibility artifact: a benchmark on your problem domain, a published eval methodology, a case study with a number the customer signed off on. One real artifact outperforms a quarter of thought leadership, because artifacts get cited and slogans get scrolled past. This is the same bet I make with the answer library on this site: the durable, narrow, checkable page beats the essay in retrieval, and retrieval is where buying decisions increasingly start.
Where the playbook could break
I do not write hagiography, so here is the honest risk ledger.
The margin question is unresolved. Outcome pricing plus forward-deployment means Sierra carries compute cost on failures and engineering cost on every deployment. Ghostwriter and Agent Studio are the bet that the last mile compresses faster than competitors productize; if that race goes badly, growth gates on hiring and margins compress exactly as Sacra warns. The blast radius is growing: Level 1 PCI compliance and healthcare deployments mean a single payments breach or benefits error carries consequences a misconfigured support bot never did, and Sierra has already disclosed a coordinated jailbreak attempt against more than a dozen customer agents and a guardrail misconfiguration that let Gap's agent chat off-script. Outcome revenue is also harder to forecast than subscriptions, the same discipline-for-a-free-lunch trade Snowflake lives with. And the competitive question Taylor himself refuses to dismiss: whether the frontier labs' agents eventually swallow the applied layer. His defense, different buyers, bespoke integration, and the observation that "coding the software was never the hard part," is the right argument. It is not a proof.
But notice what every one of those risks has in common: they are the costs of accountability. Sierra chose a model where it is exposed to its own failures, financially and reputationally, in a way the seat-license world never was. That exposure is the moat and the risk at once, and I think they chose correctly, because the alternative, shipping software whose failure is the customer's problem, is exactly the model the market is currently repricing.
What I am stealing for Heidi
I apply everything I write or the post is spectating, so, concretely. I am restructuring one Heidi workflow, the one with the cleanest countable outcome, to price on completion rather than access, this quarter, to force the internal accountability shift before I scale it. I am converting our annotation backlog into regression tests using exactly Sierra's loop, failure annotated, test created, then fixed, in that order. And I am writing off two pieces of scaffolding I have been treating as assets that the next model generation will absorb, because I would rather delete them on my schedule than defend them on the market's.
Try this week
Pick the one that matches your seat.
If you own pricing: find the single most countable outcome in your product and draft one contract where 20% of the price rides on it. You will learn more about your own product's reliability from that negotiation than from a quarter of dashboards.
If you own product or engineering: take your last five agent failures, turn each into a reproducible test before anyone fixes anything, and make "no fix without a test" the rule. That is the whole ADLC in embryo.
If you own go-to-market: write down the highest-volume, most countable workflow inside your best prospect, and rebuild the pitch around landing only that. The platform pitch can wait until the number exists.
Sierra's growth looks like magic from the outside because six mutually reinforcing choices compound. From the inside it is just the same decision made six times: put yourself on the hook for the outcome, everywhere, on purpose. That is available to any company willing to be accountable for whether its product actually lands. Most are not. That is the opportunity.
Sources: Sierra revenue, valuation and funding, Sacra. Bret Taylor's Sierra reaches $100M ARR in under two years, TechCrunch, November 2025. Bret Taylor of Sierra on AI agents, outcome-based pricing, and the OpenAI board, Cheeky Pint interview with John Collison, March 2026. The Agent Development Life Cycle, Bret Taylor and Clay Bavor, Sierra blog. Outcome-based pricing for AI agents, Sierra blog. τ-voice benchmark, Sierra blog.
Also on Medium
Full archive →AI Agents and the Future of Work: A Pixar-Inspired Journey
What product managers can learn about AI agents from how Pixar runs a film team.
Many AI Agents Are Actually Workflows or Automations in Disguise
How to tell agents from workflows from cron jobs, and why it matters for what you ship.
Frequently asked
Why is Sierra growing so fast?+
Sierra went from launch in February 2024 to $100M ARR in seven quarters, $150M in eight, and roughly $200M by May 2026, at a $15.8 billion valuation. The growth is not primarily a model advantage. It is an operating-model advantage: outcome-based pricing that puts the vendor's revenue behind the customer's result, a forward-deployed engineering motion that owns the last mile of implementation, an engineering discipline built for non-deterministic software, and a land-one-channel-then-expand enterprise motion aimed at the biggest brands rather than the easiest logos.
How does Sierra's outcome-based pricing actually work?+
If the AI agent resolves the case with no human intervention, the customer pays a pre-negotiated rate, reportedly around $1.50 per resolution in support contexts. If the agent escalates to a human, it is free. Contracts are custom-negotiated, typically starting around $150,000 a year, blending volume-based pricing for routine interactions with pay-per-resolution for complex ones. Taylor explicitly distinguishes this from usage-based pricing: tokens do not correlate with business value, outcomes do.
What is Sierra's Agent Development Life Cycle?+
Sierra's answer to the fact that agents break the traditional software development lifecycle. Four practices: declarative goals and guardrails instead of scripted flows, immutable agent releases that snapshot code, prompts, model versions, and knowledge so any release can be rolled back or A/B tested, continuous QA where customer-experience staff annotate real conversations daily, and those annotated conversations becoming regression tests, so an agent never makes the same mistake twice. Sierra runs every customer's regression suite before each platform upgrade.
Who does Sierra sell to and how?+
The biggest brands first: roughly a third of clients have over $10 billion in revenue, more than half over $1 billion, and 40% of the Fortune 50 are customers, including Cigna, Nordstrom, ADT, SiriusXM, Ramp, and Rocket Mortgage. The buyer is the CFO, Chief Customer Officer, or Chief Digital Officer, not the IT department. The motion lands one channel and a few use cases, proves resolution rate and CSAT, then expands, with almost all clients eventually running both voice and chat.
What does 'build to be commoditized' mean at Sierra?+
Sierra deliberately builds capabilities ahead of the frontier models knowing the models will absorb them. They built chain-of-thought scaffolding before o1 existed and best-in-market Cantonese voice for a Hong Kong bank knowing it will be commoditized within three years. Taylor's rule: teams that treat code obviated by a general-purpose model as precious will fall behind. The advantage is pace, not accumulated IP.
What should other companies steal from Sierra?+
Five things: denominate at least part of your price in the customer's outcome so your revenue depends on landing, own the last mile of implementation instead of throwing software over the wall, build the eval-and-regression discipline before you scale the agent, land one narrow high-volume workflow before you pitch the platform, and market with credibility artifacts like benchmarks and engineering essays instead of category slogans.

Comments (0)
Sign in with LinkedIn to leave a comment.
Sign in with LinkedIn