
Rich Holmes at Department of Product wrote up Stripe's Harbor this morning, and it is the best internal prototyping story I have read. I want to argue with the number in the headline.
Not because it is wrong. Because it is the wrong number.
The short version
Stripe's Harbor, per Rich Holmes at Department of Product, was built by one engineer in about 40 days and has produced more than 12,000 prototypes since May, used by a quarter of the company. That count hides the outcome that matters. In my own Instant Prototype Agent, about a third of prototypes became production after engineering cleanup, a third triggered a better idea, and a third were killed in 20 minutes instead of a three-week spike. The rewrite: report the split and the median time from first render to the kill decision, not prototypes created. And when a prototype stops being thrown away, as Harbor's finance and risk dashboards have, it is production without an owner. Across my 39 agents, 13 are orphaned, and a named owner was the best predictor of survival.
What Stripe built
The facts, from the free half of the piece (the deep dive is paywalled past the section on internal tool connections, and I am not claiming anything past it).
Harbor is Stripe's internal prototyping platform. One engineer, Cristian Rivera, built it in about 40 days. Since May it has produced more than 12,000 prototypes, and about 25% of the company uses it. You pick a design system, describe the interface in your own words, and an agent writes a multi-file project with Stripe's own components and realistic mock data. The agent can read the rendered page, switch the viewport between tablet and desktop, and write up the errors it found so it can come back and fix them. You share it, coworkers comment, and agents read the comments and mark them done or archive them. Stripe says the comment overhead was one of the biggest problems, and linking comments to versions is what tamed it.
The original scope was designers turning static designs into something a PM could click. It has since spread into finance, risk, and sales, where teams build the internal apps and dashboards that used to sit in a prioritisation backlog. Stripe's engineering lead says it changed how they work "seemingly overnight."
I believe all of it. I run a smaller version of the same loop.
My version of the loop
The Instant Prototype Agent: Customer Request to Prototype in Minutes listens for feature requests in Slack, Zendesk, Gong, and Salesforce, and turns a validated one into a working prototype on a Vercel preview URL in 7 to 10 minutes, with a branch, a Linear ticket, and a Notion doc behind it. The customer gets a clickable thing the same day.
Because I have run it for a while, I know where the prototypes go. About a third became the production implementation after engineering cleaned them up. Another third triggered a better idea than the one the customer asked for. The rest were rejected in 20 minutes.
Twenty minutes. The old price of that no was a three-week engineering spike, or worse, a feature that shipped because nobody had a cheap way to find out it was wrong.
That third is the one a prototype count hides. All 12,000 look the same in the headline. The ones that shipped, the ones that changed the question, and the ones that died in the first meeting are all one number. Nobody puts kills on a slide, so nobody reports them, so the tool gets judged on volume, and volume is the one thing a prototyping tool can always produce more of.
The rewrite
Do not report prototypes created.
Report the split. Shipped after cleanup, replaced by a better idea, killed. And report one more number, the median time from first render to the kill decision.
That second number is the health check. When a prototype is doing its job, the kill comes fast, because the whole point was to find out. When the median time to kill starts climbing, prototypes are lingering, and a lingering prototype is not learning anything. It has become something else.
(My handbook chapter Prototype Before You Spec says prototypes are for learning, not for polish. The kill time is how you measure whether that is still true.)
Where I get off the Harbor story
Here is the part of the write-up I would push on, and it is the part Holmes presents as the win.
Harbor's growth is in finance, risk, and sales. Teams that used to wait on product and engineering now build their own internal apps and dashboards. Holmes pairs it with a Lovable report saying internally built applications went from 11 in April to over 80 in August.
A dashboard a finance team builds and then uses every month-end is not a prototype. It is production. It just has no owner, no maintenance budget, and nobody whose job breaks when it drifts.
I have the bill for that one. This spring I shipped 39 PM AI Agents Deployed: What Stuck, What Died, and Why, 39 agents in 80 days. Thirteen are orphaned, nothing else I have written refers to them. The best predictor of which agents survived was whether a named human owned them. The ones that lasted a year of model migrations cost about two hours a month each to keep true. Most of the ones that died, died because nobody budgeted those two hours.
Same handbook chapter, one more rule: do not prototype financial systems where mistakes are expensive. I would say that out loud about a tool whose fastest-growing users are finance and risk. Not because they should stop. Because the moment one of those dashboards stops being thrown away, it needs a name on it.
What to do this week
If you run an internal prototyping tool, add two columns to whatever dashboard reports on it. Outcome, with three values. And time to kill.
Then pull the list of prototypes older than 30 days that someone still opens. Every one of those is production. Give each a named owner and two hours a month, or turn it off before it becomes the next thing that quietly stops being true.
This one belongs to the running argument on AI Product Management: AI collapsed the cost of the parts of the PM job we were trained to be good at, and building the prototype was one of them. What is left is deciding what the prototype was for, and killing it when it has answered.
Related answer: What should a prototyping tool report instead of prototypes created?
Sources: How Stripe built an internal prototyping tool that's created 12,000 prototypes so far, Rich Holmes, Department of Product, September 2026.
Also on Medium
Full archive →Frequently asked
What is Stripe's Harbor prototyping tool?+
Harbor is Stripe's internal AI prototyping platform, as written up by Rich Holmes at Department of Product. One engineer, Cristian Rivera, built it in about 40 days. An author picks a design system, describes an interface in plain language, and an agent writes a multi-file project with Stripe's own components and mock data, reads the rendered page, and fixes the errors it finds. Coworkers comment on the shared prototype and agents work the comments. Since May it has produced more than 12,000 prototypes and is used by about 25% of the company.
Why is the number of prototypes created the wrong metric?+
Because a prototype's best outcome is often its own death. In my own Instant Prototype Agent, about a third of prototypes became the production feature after engineering cleanup, a third triggered a better idea than the one requested, and a third were rejected in 20 minutes instead of a three-week spike. A raw count treats all three the same and hides the third that saves the most engineering time.
What should a prototyping tool report instead?+
The split of outcomes, shipped, replaced by a better idea, or killed, plus the median time from first render to the kill decision. If that time is rising, prototypes are lingering, which usually means they have stopped being prototypes and started being unowned production.
Is it a problem that non-engineers build internal apps in a prototyping tool?+
It is a problem when nobody owns what they built. Stripe's write-up says Harbor has spread into finance, risk, and sales, where teams build internal apps and dashboards that used to sit in the backlog. A dashboard a team keeps using is production. Across 39 agents I shipped in 80 days, 13 are orphaned, and the best predictor of survival was a named human owner. Each survivor costs about two hours a month to keep true.
When should you not prototype?+
When you already know exactly what to build and engineering is waiting, when the problem is purely backend, when the feature exists and you are optimizing it by ten percent, and when mistakes are expensive, which is my handbook's rule for financial systems. Prototypes are for learning, not for polish. The moment one is kept, it needs an owner and a maintenance budget.

Comments (0)
Sign in with LinkedIn to leave a comment.
Sign in with LinkedIn