
Salesforce signed an agreement this week to acquire Listen Labs, and one word in the announcement carries more weight than the rest. The word is simulate. I shipped a feature that a simulation of my users would have waved through, so I have a view on where that word belongs in a research plan.
The short version
Salesforce says Listen Labs' agents design studies, recruit from more than 50 million participants, interview in more than 120 languages, and synthesize, taking research from months to days. It also sells digital twins, simulations grounded in real customer behavior, so companies can understand and even simulate what customers will do. Basia Kubicka showed the small version the next day in a sponsored post on Aha! Builder: a persona generated from her app, then virtual user testing against it. My onboarding agent found real stuck points, offered help, was read as surveillance, and left activation below the control group for a quarter. Nothing in the behavior data held that reaction. The rewrite: simulate to test the instrument, then run a small real cohort and measure how it felt. A twin's answer never goes in the evidence column about a person.
Two products in one announcement
The release describes two different things, and it helps to pull them apart.
The first is real research, run by agents. They design the study, recruit the participants, conduct the interviews, and synthesize the findings on one platform. The numbers Salesforce gives are a network of more than 50 million participants and interviews around the clock in more than 120 languages. Listen Labs' CEO, Alfred Wahlforss, frames it as the difference between talking to a handful of customers and having thousands of real conversations at once.
I would take that half today. One of my failures in 10 AI Agents I Built That Failed. The Honest Retrospective. was a sentiment classifier that scored 89% on a held-out set and 61% in production within six weeks, because the training set was English-heavy and non-English customers were scored as angry no matter what they wrote. More real conversations with the people you were not hearing is the direct fix for that.
The second product is the twin. Salesforce describes digital twins as AI simulations grounded in real customer behavior, which generate simulated responses so teams and agents can compare approaches and refine their plans before further customer or user testing. The headline version is that companies can move past surveys to understand, and even simulate, what customers will do and why.
The small version is already in your builder
Basia Kubicka posted a walkthrough of Aha! Builder today. It is a paid partnership and she marks it as one. She built a draft checker for her LinkedIn posts. The part she found interesting was what happened after the app existed: the assistant defined a user persona, goals, and a vision from the app, and she ran virtual user testing against that persona. She had a first round of feedback before anyone else saw it. Real user testing, through a feedback widget, is the step after.
That is the same idea at the scale of one PM and one afternoon. And her sequence is the right one. The simulated round comes first and the real round follows.
What the twin would have told me
The onboarding walkthrough agent is number seven in that retrospective. It watched new users in their first session, noticed where they got stuck, and offered contextual help.
It was technically correct. The stuck points were real. If you had built a simulation of those users from their behavior, the simulation would have had the same stuck points, and it is hard to see how it would have objected to help arriving at exactly those moments.
Users hated it. They read it as surveillance. Activation among users who got the interventions was lower than the control group, and it cost a quarter plus the engineering to build and remove it.
The meeting summary agent, number nine, failed the same way from the other side. Accurate summaries. People stopped raising sensitive topics once it was in the room, and the hard conversations moved to hallways.
Three of those ten failures had this shape: the agent did what it was asked, and what users cared about was something it was not measuring. A twin grounded in past behavior is grounded in a world where the feature did not exist yet. The reaction I needed to know about was a reaction to the feature.
The rewrite
"Simulate what customers will do" becomes two lines.
Simulate to test the instrument. Run the interview guide past a twin to find the leading question. Generate the edge cases. Rehearse the one session you cannot afford to waste. Check whether a flow breaks for a segment you cannot recruit. All of that tests your process, and a simulation is good at it.
Then a small real cohort, and measure perception next to behavior. That is the rule I took from the onboarding agent: prototype the intervention with a small group and ask how it felt before scaling it. Behavior tells you what they did. It does not tell you that they felt watched.
The release's own sentence already has the order right. Twins help teams refine plans before further customer or user testing. Keep the word before, and write it into the research plan so nobody can quietly drop the second step when the simulated chart looks convincing.
For the case where the customer is software, Customer Discovery When Your Customer Is an Agent covers scripted agent sessions against your own endpoints, which is a different and more honest use of simulation. The weekly human cadence this sits on top of is in the handbook chapter Continuous Discovery.
What to do this week
Open your evidence log, or whatever holds the reasons behind the roadmap. Add a column that says where each row came from: observed, said by a customer, or simulated. Then write one rule at the top. Simulated rows can change a question. They cannot justify a decision.
It belongs to the argument on AI Product Management: AI collapsed the cost of the parts PMs were trained to be good at, and left the parts nobody trained for. Running the study just got cheap. Knowing which answers count as evidence did not.
Related answer: What can a customer digital twin tell you before launch?
Sources: Salesforce Signs Definitive Agreement to Acquire Listen Labs, Salesforce News, September 29, 2026. Basia Kubicka on Aha! Builder, LinkedIn, September 30, 2026, a paid partnership. 10 AI Agents I Built That Failed. The Honest Retrospective., falkster.com, May 4, 2026.
Frequently asked
What is Salesforce buying with Listen Labs?+
Per the Salesforce announcement of September 29, 2026, Listen Labs is an AI-powered customer research and human simulation platform. Its agents design studies, recruit participants from a global network of more than 50 million people, run interviews around the clock in more than 120 languages, and synthesize the findings, taking research from months to days. It also offers digital twins, simulations grounded in real customer behavior. The deal is expected to close in the fourth quarter of Salesforce's fiscal 2027, subject to regulatory clearance.
What is a customer digital twin?+
In Salesforce's description, an AI simulation grounded in real customer behavior that generates simulated responses, so teams and agents can compare approaches and refine plans before further customer or user testing. The smaller version is the generated persona in an app builder: in Basia Kubicka's sponsored walkthrough of Aha! Builder, the assistant defined a persona from her app and she ran virtual user testing against it.
What can a digital twin not tell you?+
How people will feel about something that does not exist yet. My onboarding agent found real stuck points in real first sessions, and a simulation built on that behavior would have had them too. Users read the help as surveillance and activation came in below the control group. The reaction was to being watched, and it was not in any behavior recorded before the feature shipped.
Where do simulated customers belong in a research plan?+
In instrument testing. Use them to find the leading question in an interview guide, to generate edge cases, to rehearse a session, and to check flows for a segment you cannot recruit. Then test with a small real cohort and measure perception next to behavior. A simulated answer is never recorded as evidence about a person.
Is the real-interview half of Listen Labs worth having?+
Yes, and it is the stronger half. Interviews at scale in more than 120 languages go straight at a failure I have had: a sentiment classifier that scored 89% on a held-out set fell to 61% in production within six weeks because the training set was English-heavy. More real conversations with the customers you were not hearing is the fix. A simulation of them is not.

Comments (0)
Sign in with LinkedIn to leave a comment.
Sign in with LinkedIn