Undo Is a Design Primitive
Undo used to mean the last thing you did. It has to mean the thing the system did while you slept, which is three design problems instead of one.
The short version
Undo is a design primitive again, and the version most products ship is too weak for a system that acts on its own. When an agent takes an action overnight, undo splits into three problems: the reversal, the notification, and the window between them. Teams build the first, add the second late, and almost never name the third. The window is where the trust lives, because an action that is technically reversible but whose window closed while the user slept is not reversible in any way that matters to them. Classify every autonomous action before designing any of it, and accept that one class does not get to run autonomously at all.
Undo was simple when the person did the thing themselves. They acted, they regretted it, they pressed the other thing. The stack lived in their head as much as in the software, because they had been present for every entry.
Now the stack has entries the user never made.
An agent sorted the inbox at 2am. A pipeline drafted eleven replies and queued four. Something got tagged, merged, escalated, rescheduled. The person wakes up to a product that has been busy on their behalf, and what they need first is not a summary, it's the ability to take back whatever they disagree with, cheaply, without reading everything first.
That makes undo a primitive in the strict sense. Everything else on the surface depends on it. Confidence signals, disclosure, tone, how assertive the product gets to be, all of it is downstream of whether the user believes they can put things back. A product with a real undo can afford to be bold. A product without one has to hedge on every screen, and hedging on every screen is how a capable system ends up feeling timid.
Three parts, and most teams build one
Reversal is whether it can be taken back, by whom, at what cost.
Notification is how the person finds out it happened.
The window is how long they have, and what happens when it closes.
Engineering will build the first because it looks like the whole problem from inside the code. The second usually arrives late, wired into whatever notification system already exists, which is why it so often says that something changed without saying what changed. The third almost never gets named at all. Nobody writes down the number. It emerges from a cache expiry or a cron job, and the user meets it on the worst possible day.
Ask your team what the undo window is for your most common autonomous action.
If the answer takes more than ten seconds, or comes back as a question, that is the finding.
Classify before you design
What goes in the checklist depends on the kind of action, so sorting comes first. Four classes.
| Class | Definition | Example shape |
|---|---|---|
| Reversible, private | Undoing restores the previous state, nobody outside saw it | Drafted, sorted, tagged, organized |
| Reversible, witnessed | State restores, but somebody already saw the original | A message edited after delivery, a status somebody read |
| Irreversible, contained | Cannot be undone, effect stays inside the account | Data permanently deleted, a credit consumed |
| Irreversible, external | Cannot be undone, effect left the building | Payment sent, email delivered, external write, permission granted |
Number two is the interesting one. Reversible and witnessed is where teams get lazy, because the database restores cleanly and the ticket closes. But somebody read the original. The system state went back and the social state did not, and the interface has to say so rather than pretending the edit never happened.
The fourth class carries the rule that everything else hangs off. Irreversible external actions do not run autonomously. Money leaving, a message delivered outside the account, a permanent deletion, a permission granted. A human confirms in-session or it does not happen. Call that conservative if you like. It remains the one place where "the agent handled it" costs more than it saves, because there is no interface that can call back a sent payment.
I have watched teams argue this one for weeks and land in the same place every time, usually after somebody sketches the apology email.
The window is the design decision
Here is where I would spend your review time.
There is no universal window length and any table that hands you one is guessing. Two things set it, and both are observable this week.
How often the user is actually present. If people open the product on weekday mornings, a four-hour window is a fiction. It expires inside the gap between sessions, which means the reversal exists in the code and not in the user's life.
What it costs you to hold the action reversible. Storage, complexity, external systems that will not wait. Where that cost is near zero, be generous enough that nobody ever bumps into the limit. Where it is real, name the number and defend it.
Write both down per action, pick, and record it. The point of recording it is that the next person does not re-derive it from scratch and land somewhere different, which is how a product ends up with five inconsistent windows nobody chose.
Then three properties the window has to have. It is stated in the interface before the action runs, not discovered afterwards. It is stated again in the notification, in the user's own terms, so "until Friday morning" rather than a timestamp. And it survives the user being away, because a window that expires overnight for an action taken overnight is a window in name only.
Something has to happen when it closes. A summary, a confirmation, a visible state change. Silence at the close is how people learn the window existed only after it mattered.
The copy is the deliverable
Three sentences per autonomous action, written early rather than in the last week before launch.
What the product will do, and how long it can be taken back. What it did, and the control to take it back. That the window has closed, and what the state is now.
That third sentence is the test. If you cannot write it without it sounding alarming, the action belonged in a class that needed confirmation, and the copy just told you something the design review did not. That is the same move as You Cannot Mock a Distribution: you find the boundary by writing the worst acceptable case out loud, not by drawing the good one more carefully.
Two more things worth putting in the spec. Reversal costs no more actions than the original did, and partial reversal is possible when the action was a batch. Undoing forty changes because one was wrong is a punishment for using the feature, and people only need to be punished once.
What this buys you
A working undo changes what the rest of the product is allowed to do. It is the precondition for autonomy, and teams that skip it end up building an agent that asks permission for everything, which is a worse product than no agent at all.
It also changes what happens when the system is wrong, which is the subject of Recovery and Trust Repair. Repair is much cheaper when the user still has a path to unwind whatever they acted on. Without one, the apology is all you have, and apology alone repairs very little.
Some of this is a safety conversation as much as a design one, and the framing in Trust and Safety is the version to bring to a room that includes legal.
This week: list every action your product takes without a human in the loop. Assign each one a class from the table above. Any that land in irreversible external are the conversation to have on Thursday, and the list is almost always longer than anyone on the team expects.
The full checklist, with the reversal, notification, and window sections and a table for setting window length, is in The Undo Checklist. The short version to send someone who does not want the argument is What does undo mean when an agent acted on your behalf overnight?
Chapter 7 of a series on the design operating model. Next: how much of the system's reasoning to show, to whom, and when showing it costs trust instead of building it.
Take the template
The undo checklist
Frequently asked
What does undo mean when an AI agent acted on your behalf overnight?+
It means three things at once, and they have to be designed separately. The reversal is whether the action can be taken back, by whom, and at what cost. The notification is how the person finds out it happened at all. The window is how long they have and what happens when it closes. Most teams build the reversal, bolt on a notification late, and never name the window, which is the part users actually feel.
Which autonomous actions should never run without a human confirming?+
Anything in the irreversible external class: money leaving, a message delivered outside the account, a permanent deletion, a permission granted, a write to a third-party system. Those effects have left the building and no interface can call them back. Everything else can be classified as reversible private, reversible witnessed, or irreversible contained, and each of those can run autonomously with the right window. The irreversible external ones get a human in-session, every time.
How long should an undo window be for an autonomous action?+
There is no universal number, and any table that gives you one is guessing. Set it from two things you can observe: how often the user is actually present, and what it costs you to hold the action reversible. A window shorter than the gap between a user's sessions is decorative. Where holding it open is close to free, make it generous enough that nobody ever bumps into the limit, then write the number in the decision record.
Why is the undo notification part of the undo and not a separate feature?+
Because an undo the user never learns about does not exist. If the notification says something changed without saying what, they have to go find it. If the notification does not carry the undo control itself, they have to go looking for that too. The notification is where the reversal is actually delivered, so the control lives in the notification, and the notification names the change, not just the fact of a change.
How do you handle undo for a batch of autonomous actions?+
Partial reversal has to be possible. Undoing forty changes because one of them was wrong is a punishment for using the feature, and users learn from it fast. Batch the notification so people are not trained to ignore forty separate alerts, but keep the reversal granular enough that a single bad item can be pulled out. One summary, many undo targets.
What should happen when an undo window closes?+
Something visible. A summary, a final confirmation, or a state change the user can see. Silence at the close is how people discover the window existed only on the day it mattered, and that discovery costs more trust than the original action ever earned. If you cannot write the closing sentence without it sounding alarming, the action probably belonged in a class that required confirmation.
Related reading
Chapters and essays on the same thread, across both handbooks.
One Problem a Week
Marc Andreessen says Musk finds the biggest problem at each company every week, fixes it, and repeats 52 times a year. Strip out the hero and what is left is the only leadership loop that survives AI, at every level of the org.
I Gave My AI Agents a Performance Review. Three Got Fired.
If an agent does the work of a team member, manage it like one. The scorecard, the coaching loop, and the firing criteria for the agents on your product team.
Agent-to-Agent Dispatch: The Product Org Chart Nobody Is Designing
Everyone is building single AI agents for PMs. The real shift is agents handing work to other agents, with the PM as dispatcher. Here is the architecture.
Design Just Got Promoted
The claim that design is lost contains a category error. Drawing got cheap. Deciding got more valuable, and there is more to decide than in twenty years.
Recovery and Trust Repair
The thirty seconds after your product is confidently wrong decide whether the person keeps using it. Almost nobody designs that moment.