LeadershipNew·Falk Gottlob··10 min read

100 to 3,000 in a Week: Why Microsoft's Best Agent Is Not Called Copilot

Microsoft's Scout agent, formerly ClawPilot, went from 100 to 3,000 users in a week with no mandate. What its naming, identity, and org design teach.

AI agentsMicrosoft ScoutClawPilotCopilotOmar ShahineProject Lobsteragent identitydogfoodingautopilotoutcome accountability
Helpful?

Illustrated cover for the falkster.com post: a steep adoption curve rising from a small dot labeled 100 to a large cluster labeled 3,000, with no push arrow behind it, only pull.

Someone asked me this week whether Microsoft runs something internally called Scout instead of Copilot. Close. The interesting part is what the question gets wrong.

The short version

Microsoft's most-adopted internal agent is not Copilot. It is Scout, formerly ClawPilot, an always-on desktop agent that went from about 100 to more than 3,000 daily users inside Microsoft in a single week, with no mandate and no campaign. The name is not marketing. Scout gets its own Entra identity and runs in a zero-trust runtime, so it is a separate principal in your directory, which forces a separate noun, a separate audit row, and a separate answer for your CISO. Copilot waits for you to ask. Scout acts on a schedule. The lesson for anyone shipping agents: internal pull is the only pre-launch eval that means anything, identity comes before capability, and if you measure an autopilot by daily active users you kill the exact thing that made it spread.

Scout is not a secret. Satya Nadella put it on the Build stage on June 2, and Omar Shahine, the CVP who owns it, published the launch post the same day. It is in Frontier private preview right now, gated behind program enrollment, an Intune policy, and an admin opt-in attestation. You need a Microsoft 365 Copilot license and a GitHub Copilot Business or Enterprise license to install it. Windows 11 or macOS 12 and up. Desktop only, no mobile.

Here is the part that is true, and better than the rumor. Before it was Scout it was ClawPilot, and inside Microsoft it went from roughly 100 daily users to more than 3,000 in a single week. GeekWire reported that number on May 4. There was no campaign behind it, no mandate, no enablement deck, no VP telling a skip level to drive attach rates. People downloaded a desktop app because they wanted it.

Now compare that to how Copilot arrived on those same employees' machines.

That contrast is the whole post. There are four or five decisions buried in it that I want to steal.

Where it actually came from

Shahine had been messing with OpenClaw at home earlier this year. He built a personal assistant called Lobster on top of it, gave it its own Apple ID and email address so he could text it from any device with iMessage, and split it into three agents, each with a separate security profile and its own tool access. Eight weeks later he was running nine always-on agents. They handle travel logistics and family reminders. Boring domestic work, done well.

He showed it to Microsoft's AI Accelerator group. That demo got him a new job: bring OpenClaw-style agents to Microsoft 365 as CVP of what became Project Lobster.

At the same time, and separately, a Member of Technical Staff named Jakob Werner was building a desktop-app version of the same idea, with enterprise security as the point rather than an afterthought. Internally it got called ClawPilot. Within a couple of weeks thousands of Microsoft employees had downloaded it.

Then the two efforts merged. Shahine assembled a deliberately small team he named Ocean's 11: a platform squad in Redmond, a Teams integration squad, a governance and identity squad, and a group in Oslo who owned the desktop runtime and gave the project engineering coverage around the clock. Which, for an always-on agent, is not a cute detail. It is the operating model matching the product.

The team runs without an executive assistant. Everyone uses the prototypes all day. Everyone writes code, including the CVP, and Shahine has said some of his own pull requests did not clear the bar. When he talks about what surprised him, it is the volume of unsolicited contribution: he has never seen a project inside the company where that many people showed up with their own ideas and their own code.

Werner has a word for this that I have not been able to stop thinking about. Gravity. Build something whose influence is large enough that good new ideas want to fall into it rather than spin off into their own orbit. Not feature surface. Pull.

Why it is not called Copilot

Four reasons. Only one of them is marketing.

The word Copilot is a promise about assistance. A copilot waits. You ask, it drafts, you edit. That is a specific contract with the user and Microsoft has spent three years and enormous money teaching 400 million seats to expect exactly that. Scout does something categorically different: it runs on a schedule you set, watches your inbox and your Teams messages, executes shell commands, drives a browser through Playwright, and finishes work while you are in a meeting. Microsoft named that category Autopilots. Ship the second thing under the first name and you corrupt both promises. Users who expected a chat box get an agent that moved their calendar. Users who wanted the agent go looking for it in a sidebar.

Identity forced the split, and this is the real reason. Scout gets its own Entra identity. It acts on your behalf as a distinct principal, with permissions that can be narrower than yours, inside a zero-trust runtime where the agent's own container is treated as untrusted and identity, tokens, and policy sit outside it. Every package comes through a signed Microsoft supply chain. Agent 365 gives admins one control plane. Purview gives security teams the same DLP and compliance signal they get from every other M365 surface.

An agent with its own identity cannot be a feature of your Copilot session. That is not a branding preference, it is an architectural fact expressing itself as a name. The moment the agent is a separate principal in your directory, it needs a separate noun in your product line, a separate row in your audit log, and a separate answer for your CISO. If it stayed a ghost inside the user's session, audit logs would blur human and machine, least privilege would be aspirational, and incident response would be archaeology.

Packaging. Two license requirements, a Frontier gate, and an attestation. You cannot upsell a feature. You can upsell a category.

Blast radius. An always-on agent with shell access and browser automation will eventually do something expensive on someone's behalf. When that happens, Microsoft would very much like the headline to name a preview product with 3,000 users rather than the brand attached to the entire productivity suite.

Why it spread when Copilot had to be pushed

This is the part product people should sit with.

It shipped as a download, not a feature. Copilot arrives in your ribbon whether you asked or not, which means installing it is not a decision anyone made. ClawPilot required a person to choose it. That single difference converts adoption from a compliance metric into a real one, and a real one is the only kind you can learn from.

It was built by people running it as their primary interface. Not dogfooding as a scheduled exercise. Shahine's agent is named Sebastien and it is how he works. When the builders are also the only users who matter, the feedback loop is hours long instead of quarters long.

The curve itself is the eval. 100 to 3,000 in a week is not something a comms plan produces. And it happened after Nadella had publicly compared this class of technology to a virus a few months earlier, which makes the pull more credible, not less. Nobody was rewarded internally for adopting it.

I want to be honest about the asterisk here. Microsoft employees face zero procurement friction, are unusually tolerant of preview software, and are the most motivated dogfooders on earth. A curve like that inside the house does not reproduce at a customer who needs an Intune policy change and two license SKUs before anyone can try it. The signal is real. The magnitude will not survive contact with enterprise IT.

The part I do not like

404 Media obtained an internal document titled "ClawPilot: Overview and Plan with Project Lobster." It lays out a three-phase strategy described as going from addictive app to agentic platform. Phase one is labeled, in plain language, making people addicted.

I am not going to pretend to be shocked that an enterprise software company wants habitual use. Every one of us has written a retention goal. What bothers me is the choice of metric, because it reveals a confusion about what the product is.

If an autopilot is working, your engagement with it should go down. That is the entire proposition. It triages the inbox so you do not open the inbox. It prepares the meeting so you do not spend Sunday preparing the meeting. Measuring that with daily active users is the same category error as measuring a PM by tickets closed or an SRE by incidents handled. You will get exactly what you measure, which in this case is an agent optimized to keep pinging you.

The right measurement for an always-on agent is two-layer, the same as everything else I have written about outcome accountability. A fast layer that tracks whether the agent's actions were accepted, reverted, or silently ignored. A slow layer that tracks whether the human's calendar, inbox depth, and decision latency actually improved over a quarter. Neither of those is DAU. If Microsoft ships Scout against an engagement target, the thing that made ClawPilot spread internally is the first thing to die.

What I am taking into Falkster

Five things, and I am already applying three of them.

Internal pull is the only pre-launch eval that means anything. If the people building the product will not voluntarily install it, the GA number you eventually report is a mandate, not a market. I now ask this about every agent surface we build: would I install this if nobody told me to?

Ship a runtime, not a feature. Features live inside someone else's product and inherit its permission model and its expectations. A runtime is a thing a person chooses, and choosing is where you get signal.

Identity before capability. Most agent products I see do the opposite. They build impressive capability, demo well, and then spend nine months in security review bolting on governance. Scout is boring in exactly the right place: the agent is a principal that can be provisioned, scoped, reviewed, disabled, and investigated. Do that first and capability ships behind it. Do it second and capability never ships at all.

Memory has to forget. Shahine's example is the one I keep repeating. He told ChatGPT his daughter was 17 and his son was 13. A year later it still believed that. The system had no concept that some facts decay and others never do. Werner built layered memory where relevance strengthens with use and fades without it. Almost nobody is building this. It is the most underserved problem in agent products right now and it is not a model problem, it is a product design problem.

Name the category when the architecture changes. Not before, and not to be clever. Copilot and Scout have different identity models, different permission surfaces, and different failure modes. The name follows the architecture. If you find yourself arguing about a name, check whether you are actually arguing about an architecture.

Pick one thing to try this week: take the last agent feature you shipped and ask whether anyone on your team would install it if it were a separate download that nobody required. If the answer is no, you do not have an adoption problem yet. You have a product problem that will look like an adoption problem in two quarters.

Sources: Microsoft's OpenClaw team takes on the personal assistant challenge, GeekWire; Microsoft Wants to 'Make People Addicted' to its New AI Assistant, Internal Documents Reveal, 404 Media.

Share this post

Also on Medium

Full archive →

Frequently asked

What is Microsoft Scout?+

Scout is Microsoft's always-on desktop agent for Microsoft 365, formerly called ClawPilot and part of Project Lobster. It runs on a schedule you set, watches your inbox and Teams messages, executes shell commands, and drives a browser through Playwright to finish work while you are in a meeting. It is in Frontier private preview, gated behind program enrollment, an Intune policy, and an admin attestation, and needs both a Microsoft 365 Copilot license and a GitHub Copilot Business or Enterprise license.

Why is it called Scout and not Copilot?+

Because the architecture changed. Copilot is a promise about assistance: it waits, you ask, it drafts. Scout runs on its own, so Microsoft named that category Autopilots. The deeper reason is identity: Scout gets its own Entra identity and acts as a distinct principal with its own permissions, which cannot be a feature of your Copilot session. A separate principal needs a separate noun, a separate audit row, and a separate answer for your CISO.

How fast did ClawPilot grow inside Microsoft?+

From roughly 100 daily users to more than 3,000 in a single week, per GeekWire's reporting. There was no campaign, no mandate, and no enablement deck. People downloaded a desktop app because they wanted it, which is what makes the number meaningful, and it happened after Satya Nadella had publicly compared the technology to a virus.

How should you measure an always-on AI agent?+

Not with daily active users. If an autopilot is working, your engagement with it should go down, because it does the work so you do not have to. Use two layers: a fast layer tracking whether the agent's actions were accepted, reverted, or silently ignored, and a slow layer tracking whether the human's calendar, inbox depth, and decision latency improved over a quarter.

What is the difference between a Copilot and an Autopilot?+

A Copilot assists on request: you ask, it drafts, you edit. An Autopilot acts on its own schedule, watches your inbox and messages, and completes work without being prompted each time. They have different identity models, permission surfaces, and failure modes, which is why they need different names.

What can product teams learn from how Scout was built?+

Ship a runtime people choose to install, not a feature that arrives in the ribbon, because choosing is where you get real adoption signal. Build identity before capability, so the agent can be provisioned, scoped, reviewed, and disabled. And ask whether your own team would install the agent if nobody required it. If they would not, you have a product problem that will look like an adoption problem in two quarters.

About the author

Falk Gottlob

Falk Gottlob

Product Executive · Founder, Falkster.AI

Thirty years shipping product at Microsoft Research, Adobe, Salesforce (Marketing Cloud / Quip / Slack), and several startups including one $6.5B exit and one acquired by Microsoft. Now founder of Falkster.AI, previously CPO at Smartcat, writing this notebook from the boardroom, not the keyboard.

Comments (0)

Sign in with LinkedIn to leave a comment.

Sign in with LinkedIn
  • Be the first to comment.

Keep Reading

Posts you might find interesting based on what you just read.