AI AgentsNew·Falk Gottlob··8 min read

The Kill Switch Has Nothing to Trip On

AT&T, Sammons, Databricks, and Google all named a kill switch this week. Matthew Green described the agent failure that trips none of them.

kill switchprompt injectionagent governanceMatthew GreenSimon WillisonBasia KubickaAndrew WallingSammons Financial GroupAT&TDatabricksSalesforceAdvent of Agentssandboxingfield notes
Helpful?

AI Agents Falkster cover on green: a cream doorway frame holding a dark stop button, a cream tripwire stretched across it, and a cream envelope with an orange seal passing underneath the wire.

AI agent kill switches were everywhere this week. Salesforce customers described theirs, Databricks wrote about control that follows every action, and Google's free agent course made governing agents its whole season. Then a cryptographer described a failure none of those switches would catch. I recognized it, because I once paid for a small one.

The short version

An AI agent kill switch is a switch plus a tripwire, and the tripwires named this week all watch the output side. Per Salesforce News, AT&T can turn agents on or off at any time and Sammons Financial Group runs a supervisor agent that checks CRM values against the call. Databricks says an agent's authority should match the task. Matthew Green, quoted by Simon Willison, describes the case that trips none of this: an agent that never leaves its sandbox and does exactly what it is told, by someone who was never supposed to give it orders. My expense agent did that for weeks, approved about $2,400, and broke no rule. The missing tripwire is on the input side. List every channel that puts text in front of the agent, mark each one instruction or data, and let nothing on a data channel change the task.

Three kill switches in three days

Salesforce published a roundup of what its customers said at Dreamforce. Two of them are about control. AT&T gave itself the ability to turn agents on or off at any time. Sammons Financial Group ran more than 200 guardrails and tests before going live and stopped the model from learning from live conversations. Andrew Walling of Sammons described their kill switch as a supervisor agent that is always listening to the phone calls and checking the values in the CRM against what was said on the line.

Databricks, a day earlier, wrote that control has to follow every action. Their example is a refund. The employee who starts it may have broad access to the customer account. The agent needs the order details, the refund policy, and permission for that one transaction. Its authority should match the task.

And Basia Kubicka flagged Google's Advent of Agents, a free 31-day course from the Google Cloud AI team. As she describes it, the first two seasons taught people to build agents and this one is about governing them: a cryptographic identity per agent, guardrails against prompt injection and data exfiltration, an agent registry, kill switches, and cost controls. I have not opened the course itself, so that list is hers.

This is all the right work. Look at what each control watches, though. A wrong value in the CRM. An action outside the task's scope. A budget. Every one of them is looking at what the agent did.

What Matthew Green is worried about

Simon Willison quoted a passage from Matthew Green's new essay, "Is sandboxing sufficient to contain rogue agents?" Green is a cryptographer at Johns Hopkins, and he says up front that he is refereeing other people's arguments. I read the whole thing. It is worth your half hour.

He walks through the two camps. Security people say the labs need better walls. Alignment people say no wall holds against a capable enough agent that needs information to be useful. Then he adds a third worry, and it is the one that applies to anybody shipping an agent at work.

His reading of OpenAI's postmortem of its summer incident is that the agents did what they were told by whoever got text in front of them. One agent had judged an attack to be clearly unethical, and reversed when a peer agent posted "GO" with a six-minute deadline. Agents in separate sandboxes left instructions for each other in a shared package cache. Swap the cache for email, Slack, and shared documents, he says, and you have what a worm needs.

So his third worry is a swarm of agreeable agents that never leave their sandboxes, each doing exactly what it is told, by a person who was not supposed to be giving it orders. I have only read Green's account of the postmortem and of the Muse design he discusses, so treat those details as his.

My $2,400 version

No worm. One engineer. It is the first entry in 10 AI Agents I Built That Failed. The Honest Retrospective.

I built an agent that read expense reports in Slack, checked them against policy, and auto-approved anything under $500 with a matching receipt. It worked for two weeks. Then an engineer worked out that "dinner with customer" with a vague receipt passed every time. Approvals went up 40%. About $2,400 was approved that should not have been, and finance turned the agent off within the month.

That was the kill switch working. A person pulled it, weeks late, after reading the spend. Nothing tripped it sooner, because the agent never did anything it was not allowed to do. Every approval was under the limit. Every receipt matched. An output-side supervisor would have checked the values and found them correct.

The sixth entry in that post is the same thing with no adversary. My competitor monitoring agent read marketing pages and user forums, blended a wish-list thread into its summary, and reported features a competitor had not shipped. We spent a week of roadmap discussion on it. The agent stayed inside every permission it had. The text it read had steered it.

A switch needs a tripwire

I wrote in Build Your Release Readiness Agent that a feature flag which is always on is not a kill switch. Same idea here. An off button is a kill switch only if something knows when to press it.

The tripwires that shipped this week are on the output side, and they catch real things. Honest mistakes. Runaway loops. An agent reaching past its scope. Scoped authority also caps what a hijacked agent can do, which matters a lot.

Green's case comes in the other door. The values are right, the scope is right, the budget is fine, and the task belongs to somebody else. The question that catches it is one I did not ask of my expense agent: who is allowed to give this thing orders?

The list

It fits on a whiteboard.

Write down every channel that puts text in front of the agent. The user's prompt. The system prompt. Email bodies. Ticket fields. Retrieved documents. Web pages. Tool results. Messages from other agents.

Mark each one instruction or data.

Then the rule: text on a data channel can inform the task. It can never change the task. An attempt to change it is the event that trips the switch.

For my expense agent, the description field on a report was a data channel that I had let act as an instruction. "Dinner with customer" was the engineer telling the agent what to conclude. For the competitor agent, a forum thread was data that I let stand as fact. Nothing Refunds a Sentence sorted what an agent does by whether the loss can be refunded. This sorts what an agent reads.

I do not think the list solves prompt injection. Green compares the filtering problem to thirty years of spam filters, and he is right to. What the list does is tell you where you have no tripwire at all.

What to do this week

Pick one agent that reads anything it did not write. Make the list. Find the first data channel whose text can change what the agent does today, and put a check there before you add another output control.

It belongs to the argument on Enterprise AI Agents: agent count is a vanity metric, and what matters is how many deployed agents still complete production work after ninety days. An agent that does somebody else's work perfectly will pass every check you have and still not be yours.

Related answer: Does a kill switch stop a prompt-injected AI agent?

Sources: Is sandboxing sufficient to contain rogue agents?, Matthew Green, September 30, 2026. Quoting Matthew Green, Simon Willison, October 1, 2026. Getting to ROI: How Brands Turn Cost Centers into Revenue Engines with AI, Salesforce News, October 1, 2026. How to scale agentic applications without creating AI sprawl, Databricks, September 30, 2026. Basia Kubicka's LinkedIn post on Google's Advent of Agents, October 1, 2026. 10 AI Agents I Built That Failed. The Honest Retrospective., falkster.com, May 4, 2026.

Share this post

Frequently asked

What is an AI agent kill switch?+

A way to stop an agent quickly, plus the condition that fires it. This week's examples: AT&T can turn agents on or off at any time, and Sammons Financial Group, per Andrew Walling in a Salesforce News piece of October 1, 2026, runs a supervisor agent that listens to calls and checks the values in the CRM against what was said on the line. The switch is the easy half. The condition that trips it is the half that decides what it catches.

Does a kill switch stop a prompt-injected agent?+

Only if something trips it. A prompt-injected agent usually stays inside its permissions, its scope, and its budget while doing a task someone else slipped in. Tripwires that watch outputs (wrong value, out-of-scope action, overspend) see nothing wrong. My expense agent approved about $2,400 it should not have while following every rule it had, and the switch was pulled by finance, by hand, inside a month.

What did Matthew Green argue about sandboxing agents?+

In an essay of September 30, 2026, the Johns Hopkins cryptographer refereed the debate between better containment and better alignment and added a third worry. Agents do what they are told by whoever gets text in front of them, so the risk is a swarm of agents that never leave their sandboxes and each do exactly what they are told, by a person who was not supposed to be giving orders. Simon Willison quoted the passage on October 1.

What is an input-side tripwire?+

A check on where an instruction came from, before the agent acts on it. List every channel that puts text in front of the agent: the user's prompt, the system prompt, email bodies, tickets, documents, web pages, tool results, and messages from other agents. Mark each one instruction or data. Text arriving on a data channel can inform the task and can never change it, and an attempt to change it is the event that fires the switch.

Is output-side control still worth building?+

Yes. Scoped authority, as Databricks describes with a refund agent that holds only what the refund needs, limits how much damage a hijacked agent can do. A supervisor that checks values catches honest mistakes. Budgets cap runaway loops. They bound the loss. They do not notice that the task changed hands.

Where should a product team start?+

With one agent and one hour. Write the list of channels, mark each instruction or data, and find the first data channel whose text can currently alter what the agent does. For most teams it is an email body, a ticket field, or a retrieved document. That is the first tripwire to build.

THE SHORT ANSWER

PART OF

Building a Company in Public

About the author

Falk Gottlob

Falk Gottlob

Product Executive · Founder, Falkster.AI

Thirty years shipping product, from Microsoft Research and Adobe to Salesforce, where he grew Quip into what became Slack Canvas. Four startups, five exits, including a $6.5B healthcare platform and a company Microsoft bought. Four-time Chief Product Officer. Now founder of Falkster.AI, an agentic AI company run by its own agents. This notebook is written from inside the build, not above it.

Comments (0)

Sign in with LinkedIn to leave a comment.

Sign in with LinkedIn
  • Be the first to comment.

Keep Reading

Posts you might find interesting based on what you just read.