AI AgentsNew·Falk Gottlob··8 min read

The Stop Rule

Salesforce's Agent Designer repairs an agent team twice, spends at most $3, then escalates. LangSmith Engine v2 tests fixes until one passes. Mine had no cap.

self-healing agentsAgent DesignerSalesforce EngineeringSohini AryaManish Kumar JhaLangSmithLangChainEngine v2Jacob TalbotGergely Oroszautonomy gatesoft failuresfield notes
Helpful?

AI Agents Falkster cover on green: a thick cream loop of two chasing arrows with a dark octagonal stop sign planted on its right edge and a small teal price tag hanging from the sign.

Two systems that repair their own agents shipped this week, and the difference between them is one line. I know what that line costs when it is missing, because I ran the loop without it.

The short version

Salesforce Engineering's Agent Designer, by Sohini Arya and Manish Kumar Jha, has agents design, verify, test, and repair teams of other agents inside Marketing Cloud: 15 to 30 minutes from request to a tested team instead of two to four hours, verifiers that return PASS, WARN, or FAIL, a meta-check across them, and self-healing capped at two fix attempts and a $3 ceiling before it stops and escalates. LangChain's Engine v2 for LangSmith, by Jacob Talbot, adds red teaming and, for deployed agents, tests candidate fixes against a broader eval set until it finds one that resolves the issue, then offers a one-click PR; the post states no cap on attempts or spend. My own autonomous pricing agent had no stop rule, drifted 14% under benchmark in five weeks, and took two quarters to recover. A stop rule has three fields, attempts, spend, and the name it escalates to, and it is a product decision written before the loop runs. "Until a fix passes" is the absence of one.

Salesforce published the stop rule

I ran product at Marketing Cloud, so I read this one closely. The manual workflow Agent Designer replaces is the one that eats the productivity gain: define the YAML, the prompts, the profiles, the tools, the budgets, and the turn limits; run the team; read every agent's output; diagnose; edit; run again. The team built an orchestrator and seven specialists, each with its own model, budget, and turn limit. Architecture, failure-mode, and compliance analysts work first, the orchestrator reconciles their contradictions and pauses for a human to approve the design, and only then do other agents write and verify the files, run the test scenario, diagnose failures, and attempt repair. The orchestrator can dispatch and verify. It is prohibited from writing an agent file itself.

Three details are worth keeping.

Verification is structured because the auditor can be wrong. Each verifier returns PASS, WARN, or FAIL rather than prose, with rules like "warn rather than fail when uncertain," and because no verifier is trusted in isolation the orchestrator runs a meta-check across all of them and asks whether they agree with each other and with the design.

Budgets are enforced by topology. A flat check for a single agent, items times cost times steps for a swarm, a whole-roster ceiling for a team, all in one configuration file, inside the orchestrator's own 500-turn budget, which is split across analysis, human approval, file generation, testing, and repair so an early phase cannot starve verification.

And the stop rule. A test doctor classifies each failure as a violation, a permission block, a tool error, a boot error, a soft failure, an infinite loop, or state corruption. Soft failures execute normally and still return the wrong result, so detecting one means comparing the output point by point with expected behavior. Because diagnosis is not always deterministic, the post says, stopping behavior became as important as repair behavior. Self-healing is limited to two fix attempts and a $3 spending ceiling. If the team still fails, the system stops and escalates.

LangSmith published the loop

Engine v2 is the in-platform agent LangChain built to work every step of the improvement cycle, and it has analyzed more than 60 million traces since May. The new release adds red teaming, which generates hypotheses about issues not yet seen in production from traces and repos, tests them, and surfaces confirmed failures for review. It detects more issue types, including error-rate, latency, and cost trends, repetitive tool calls, and unnecessarily long trajectories.

Then the repair loop. For agents on LangSmith Deployment, Engine reruns the offending inputs to confirm the issue, tests candidate fixes against a broader eval set until it finds one that resolves it, and presents the fix to the user with a one-click PR.

That is a good design, and I would want it. What the post does not state is a cap. No maximum number of candidate fixes, no spend ceiling per incident, no named escalation. Maybe it exists in the product. It is not in the announcement, and the announcement is what a buyer reads.

Gergely Orosz's Pulse this week, paywalled past the standfirst, carries the direction of travel in its first sentence: 37signals moving to agents generating nearly all its code, and code reviews will probably also go away. The human review step is the one being priced out. Which is exactly the step the stop rule protects.

What the loop costs without one

Agent number four in 10 AI Agents I Built That Failed. The Honest Retrospective. is an autonomous pricing tester. It ran experiments, read the results, and adjusted. No cap on iterations, because the whole point was that it would keep improving.

It found a local maximum that was not the global one and optimized toward it with compounding effect. Within five weeks we had anchored on a price 14% below what our benchmark analysis said was correct. Recovering the anchor took two quarters of careful re-pricing.

Every iteration passed its own test. That is what a repair loop without a stop rule looks like from inside: a run of green checks that walks a number somewhere nobody chose. The rule that came out of it is the one I still use. Anything that touches money, contracts, customers individually, or decisions that compound over time gets a human review layer that cannot be removed by configuration, and autonomous optimization in a long-feedback-loop domain gets an approval gate at every iteration, not at the end.

Salesforce's two attempts and $3 is that rule with numbers in it.

Who writes the expected behavior

The soft failure is the reason the stop rule is a product decision and not an engineering one. A test that compares output to expected behavior point by point only works if someone wrote the expected behavior, and Salesforce's post is careful about where that comes from: engineers keep the control points that matter, defining the outcome, approving the architecture, reviewing verification results, and deciding when repair must stop.

The Wince Is the Spec: Bottom-Up Evals a Model Can't Write is where I argued the same point from the other direction. Top-down evals, the ones you can write from the task description, a model drafts well. Bottom-up evals come from reading fifty real outputs and noticing what bothers you, and no model produces them, because they require having read the outputs and caring whether they are good. On Heidi that loop caught actions that were permitted, reversible, logged, in budget, and still wrong, which became a three-axis autonomy gate: confidence, reversibility, and evidence grade together. A repair loop that tests "until a fix passes" is testing against whichever of those two kinds of eval exists. If only the top-down kind does, the loop will pass a soft failure and open the PR.

The Receipt Gate is the same rule at the end of the pipeline: nothing asks the agent whether it did the work. And the handbook chapter The Eval Is The Spec is where the expected-behavior file lives before any of this runs.

The three fields

A stop rule has three fields, and all three are decisions a PM makes before the loop is switched on.

Attempts. How many repairs before the loop stops. Salesforce chose two.

Spend. The ceiling per incident. Salesforce chose $3.

Escalation. The name of the person who gets the failing case, with the diagnosis attached. Not a queue, a person.

Underneath them, the domain rule: money, contracts, individual customers, and compounding decisions get a human gate that configuration cannot remove.

Two vendors shipped self-repair in one week. One published all three fields. The other published a loop. The difference is not engineering quality; both teams are better at this than most. The difference is whether the stop rule was a product decision somebody wrote down before the loop ran, and 14% for two quarters is what it costs when nobody did.

What to do this week

Find the one repair or optimization loop in your stack that runs without a human in it. Write down its attempts, its spend ceiling, and the name it escalates to. If any of the three is blank, that loop is currently running my pricing agent's rule, and it is a matter of time before it finds a local maximum.

This one belongs to the argument on Enterprise AI Agents: agent count is a vanity metric, and what matters is how many deployed agents still complete production work after ninety days and what each outcome costs. A loop with no stop rule has no cost per outcome. It has a bill.

Related answer: What stop rule should a self-healing AI agent have?

Sources: Engineering Multi-Agent AI Teams That Build and Test Themselves, Sohini Arya and Manish Kumar Jha, Salesforce Engineering, September 28, 2026. New in LangSmith: Engine v2, Managed Deep Agents, Fine-Tuning, and More, Jacob Talbot, LangChain, September 25, 2026. The Pulse: RoR creator sparks new "death of coding by hand" debate, Gergely Orosz, The Pragmatic Engineer, September 24, 2026, standfirst only. 10 AI Agents I Built That Failed. The Honest Retrospective., falkster.com, May 4, 2026.

Share this post

Frequently asked

What is Salesforce's Agent Designer?+

Per Sohini Arya and Manish Kumar Jha on the Salesforce Engineering blog, it is a system built for Marketing Cloud in which agents design, verify, test, and repair teams of other agents. An orchestrator and seven specialists each carry their own model, budget, and turn limit; the orchestrator can dispatch and verify but cannot write an agent file. It turns an engineer's request into a tested multi-agent team in roughly 15 to 30 minutes instead of two to four hours, and about 200 engineers began adopting the model.

What stop rule does Agent Designer use for self-healing?+

Two fix attempts and a $3 spending ceiling, inside a 500-turn orchestrator budget that is divided across analysis, human approval, file generation, testing, and repair. If the team still fails after that, the system stops and escalates rather than pursuing a possibly wrong diagnosis. The post says stopping behavior became as important as repair behavior.

What does LangSmith Engine v2 do?+

Per Jacob Talbot's LangChain post, Engine v2 adds red teaming that generates hypotheses about issues not yet seen in production and tests them, detects error-rate, latency, cost, and repetitive-tool-call trends, and for agents on LangSmith Deployment reruns the offending inputs to confirm an issue, tests candidate fixes against a broader eval set until it finds one that resolves it, and presents the fix for a one-click PR. The post states no cap on attempts or spend for that loop.

What happens when a self-repairing agent has no stop rule?+

In my case, an autonomous pricing tester found a local maximum and optimized toward it with compounding effect. Within five weeks the price was 14% below what the benchmark analysis said was correct, and recovering the anchor took two quarters of careful re-pricing. The rule since: autonomous optimization in long-feedback-loop domains gets a human approval gate at every iteration.

What three fields should a stop rule have?+

The maximum number of repair attempts, the spend ceiling per incident, and the name of the person the loop escalates to when either is hit. Add the domain rule underneath: anything that touches money, contracts, individual customers, or decisions that compound over time gets a human review layer that cannot be removed by configuration.

What is a soft failure in an agent test?+

Salesforce's test doctor classifies failures as violations, permission blocks, tool errors, boot errors, soft failures, infinite loops, or state corruption. A soft failure executes normally and still returns the wrong result, so it gives no error to catch. Detecting one means comparing the output point by point with expected behavior, and the expected behavior is a bottom-up eval that a human writes by reading real outputs.

THE SHORT ANSWER

PART OF

Building a Company in Public

About the author

Falk Gottlob

Falk Gottlob

Product Executive · Founder, Falkster.AI

Thirty years shipping product, from Microsoft Research and Adobe to Salesforce, where he grew Quip into what became Slack Canvas. Four startups, five exits, including a $6.5B healthcare platform and a company Microsoft bought. Four-time Chief Product Officer. Now founder of Falkster.AI, an agentic AI company run by its own agents. This notebook is written from inside the build, not above it.

Comments (0)

Sign in with LinkedIn to leave a comment.

Sign in with LinkedIn
  • Be the first to comment.

Keep Reading

Posts you might find interesting based on what you just read.