Does a kill switch stop a prompt-injected AI agent?

THE SHORT ANSWER

Only if something trips it, and most tripwires watch the wrong side. A kill switch is a switch plus a condition that fires it, and the conditions teams ship (a wrong value, an out-of-scope action, a blown budget) all watch what the agent does. A prompt-injected agent usually stays inside its permissions while doing a task somebody else slipped in. My expense agent auto-approved anything under $500 with a matching receipt; one engineer learned that 'dinner with customer' passed every time, approvals rose 40%, about $2,400 went out, and finance pulled the switch by hand inside a month. Add an input-side tripwire: list every channel that puts text in front of the agent, mark each one instruction or data, and treat any attempt by a data channel to change the task as the event that stops the agent.

Three kill switches were described in public in one week, and The Kill Switch Has Nothing to Trip On is my note on the failure none of them is watching for. The expense agent and the competitor monitoring agent behind it are entries one and six in 10 AI Agents I Built That Failed. The Honest Retrospective. The companion sort, for what an agent does instead of what it reads, is in Nothing Refunds a Sentence.

It belongs to the Enterprise AI Agents argument: what matters is how many deployed agents still complete production work after ninety days, and an agent doing somebody else's task is not completing yours.

SOURCES

THE LONG VERSION

RELATED ANSWERS

Last reviewed 2026-10-01 · 1 min read