
Most companies I talk to have one internal agent with fans. Support uses it, or sales ops does, and somebody in a planning meeting says the obvious thing: customers would love this, can we just expose it?
Snyk did expose theirs, to every paying customer, and wrote up how. It's a good ship story and I think the headline lesson people will take from it is the wrong one.
The short version
Snyk Assist ran for about a year as an internal tool for Snyk's support team, went to customers in the support portal in April 2026, and on September 1 became a panel on every page of the product. The case study says the runtime stayed the same and "the only thing that changed was the surface in front of it." Two more things changed: who reads a wrong answer, and what counts as resolved. The year inside isn't what made the promotion safe, and my own internal docs agent is the proof that internal errors aren't free. What made it legitimate is a gate that's described in the same post: three things proven before rollout, tools attached per user, eval thresholds in the repo that block a pull request, and every internal trace turned into a test. Copy the gate.
What Snyk shipped
The case study is on the LangChain blog under Oliver Armstrong's byline, written in Snyk's voice.
Snyk's support team handles thousands of cases every couple of weeks. Each had to be read, classified, and routed. They built a conversational agent on LangChain and LangGraph that answers in plain language and can act: look up open issues, check a package for known vulnerabilities, open a support case, log a feature request.
Three phases. For roughly a year it was internal, with Snyk staff as the only users. In April 2026 it launched in the support portal. On September 1 it moved into the core product for every paying customer.
Since April, by their numbers: 60,000+ queries, 500+ customer accounts, 85%+ of sessions resolved without a support ticket, and 250+ cases the agent detected and escalated to the right team.
And the sentence I want to look at: "We kept the same agent runtime through each phase. The only thing that changed was the surface in front of it."
As an architecture statement, that's true and it's a real achievement. One runtime, several front doors, one place to fix bugs. As a description of what happens when an internal agent meets customers, it leaves out the two things that bite.
Two things that changed besides the surface
The reader. For a year, every answer the agent gave was read by someone who works in Snyk support. That person knows when an answer is off. They discard it, maybe flag it, and carry on. The post says the internal feedback loop was measured in hours. Of course it was. The users were the experts.
A customer isn't a filter. A customer reads the answer in the middle of real work and acts on it.
The count. Internally, nobody needs a resolution rate. Externally, Snyk reports 85%+ of sessions resolved without a support ticket. That's one honest definition. It's also the loose one. A customer who gave up doesn't file a ticket. I went through this in One Support Agent, Three Resolution Rates: the same agent has a zero-touch rate, an invoice rate, and a no-follow-up rate, and they aren't the same number. To Snyk's credit, the post describes grading every production run on whether the response answered the question, and calls the result a measured deflection rate. I'd want that graded number next to the 85.
The year isn't the lesson
The line that will get copied is "run it internally for a year." The post gives the reasoning: starting inside meant any initial errors remained within Snyk and didn't reach paying customers.
Errors that remain inside still land on somebody.
Number 5 in 10 AI Agents I Built That Failed. The Honest Retrospective. is an internal docs agent. Employees only. Classic retrieval with citations. It answered confidently from outdated documents, a 2021 page and a 2024 page with equal weight, and people got wrong answers about current policies. Two HR incidents traced back to it, plus legal review time.
That agent was internal the whole time. Being internal didn't catch the problem. The incidents did. Time inside is only worth what you extract from it.
The gate
Here's what Snyk extracted. All four are in the post.
- Three proofs. Before anything went in front of customers they set out to prove that the agent answered correctly, refused what it should refuse, and never returned anything the user wasn't already allowed to see.
- Tools per user. Tools are attached at request time based on the permissions the signed-in user has. This is the change the "same runtime" line hides. Internal staff can see across accounts. A customer must see only theirs.
- A blocking check. Every pull request runs the real agent against the eval suites, one of real questions with known good answers and one of red-team attempts, and blocks on thresholds committed to the repo.
- Traces become tests. Every trace from the internal sessions fed the evaluation sets.
Bailey Millns, an AI engineer at Snyk, is quoted on the principle: "if it doesn't clear the bar, it doesn't ship."
That's a promotion gate. It's written down, it's in the repo, and a pull request can fail it.
Where this meets my own rule
In May I published PMs Should Vibe Code Internal Tools, Not Customer Features. The trap I named there was promotion by inertia. A thing works in the demo, everyone's excited, and somebody asks why rebuild it, can we just ship this. I called the working prototype the bait.
Snyk's agent was built by engineers, so it isn't the case I was warning about. But it answers the question that post left open: what does a legitimate promotion look like? Same artifact, different standard, and the standard is checked by a machine on every change.
So if you own an internal agent and you've been asked to put it in front of customers, skip the question of how long it has run. Ask four things on Monday. Can it prove correct, refuse, and never-leak on a test set? Are tools scoped to the signed-in user? Does a failed eval block a merge? Did last quarter's bad internal answers become test cases?
If it's one out of four, you have an internal tool people like. That's fine. Keep it internal.
This belongs to the Enterprise AI Agents argument: the number that matters is how many agents still do production work after ninety days and what each good outcome costs. Promotion is the day the cost of a wrong answer changes hands, from your support rep to your customer.
Related answer: When is an internal AI agent ready for customers?
Sources: Oliver Armstrong, "How Snyk turned an internal support agent into a customer feature," LangChain blog, October 8, 2026; all Snyk figures and quotes are from that post.
Also on Medium
Full archive →AI Agents and the Future of Work: A Pixar-Inspired Journey
What product managers can learn about AI agents from how Pixar runs a film team.
Many AI Agents Are Actually Workflows or Automations in Disguise
How to tell agents from workflows from cron jobs, and why it matters for what you ship.
Frequently asked
When is an internal AI agent ready to be put in front of customers?+
When it passes a written gate, not when it has run for a set time. Snyk's case study gives a usable one: prove the agent answers correctly, refuses what it should refuse, and never returns anything the user was not already allowed to see; attach tools per request from the signed-in user's permissions; block every pull request on eval thresholds committed to the repo; and feed every internal trace into the eval sets.
What did Snyk do with Snyk Assist?+
According to the case study on the LangChain blog, Snyk Assist ran for roughly a year as an internal tool for the support team, launched to customers in the support portal in April 2026, and moved into the core Snyk product on September 1, 2026, as a panel on every page for every paying customer. Snyk reports 60,000+ queries, 500+ customer accounts, 85%+ of sessions resolved without a support ticket, and 250+ cases auto-detected and escalated since April.
What changes when an internal agent becomes a customer feature?+
More than the surface. The reader changes: a trained support rep recognizes a wrong answer and discards it, and a customer acts on it. The permission model changes: staff can see across accounts and a customer must see only their own. And the success count changes: a session with no support ticket can be a resolved customer or one who gave up.
Is running an AI agent internally first a safe way to test it?+
Safer, not safe. I ran an internal documentation agent for employees that answered from a 2021 page and a 2024 page with equal weight, and two HR incidents traced back to it. Internal use shrinks the blast radius and shortens the feedback loop. It does not make errors free, and time spent inside proves nothing unless the traces become tests.
Does 85% of sessions resolved without a ticket mean the agent resolved 85% of issues?+
Not necessarily. No ticket filed is one definition of resolved. A stricter one is zero human touch and no related follow-up within a set window. I would ask for the same sessions scored under more than one definition before quoting the figure, which is the argument of One Support Agent, Three Resolution Rates.
What is promotion by inertia?+
It is my name for shipping a prototype or internal tool to customers because it already works. Somebody asks why rebuild it, and the question skips the part that makes it safe. Snyk's agent is the counterexample: the runtime stayed the same, and it was promoted through permissions, refusal tests, and a CI gate.

Comments (0)
Sign in with LinkedIn to leave a comment.
Sign in with LinkedIn