FoundationNew·Falk Gottlob··7 min read

Five Sessions, Then the Prompting Habit Sets

Users settle on 2 to 4 prompt templates within about 5 sessions. My onboarding agent lowered activation. Put a person in the first week instead.

FoundationDesignAI onboardingAI adoptionprompting habitsAI enablementfirst-run experiencepaired shippingJakob NielsenUX TigersShengqi ZhuBasia Kubickathe rewrite
Helpful?

Foundation-pink Falkster cover: six dark footprints growing from small to large across a cream slab of cement, with a mason's trowel resting at the slab's right edge.

A team gets an AI tool in March. In June somebody schedules the training. That training is about three months late, and as of last week there's a number for how late.

The short version

Jakob Nielsen's October 9 roundup reports a study of 139,535 ChatGPT sessions from 7,955 users. People keep bringing new tasks, and they settle on 2 to 4 personal ways of phrasing a prompt within roughly 5 sessions. His advice is to treat sessions 1 to 5 as the most valuable design real estate in an AI product, and one of his five guidelines is to manufacture the feedback chat never gives. I built a proactive version of that, an onboarding agent that watched a first session and offered help, and activation for the users it helped came in below the control group. The window is real. What goes in it has to be something the user reached for, or a person they trust. Inside a company, the same finding retires the recurring AI training slot: move the paired shipping session into each person's first week with a tool, and run it again when a major model ships.

What the study found

The study is by Shengqi Zhu and co-authors at Cornell and Loyola University Maryland. I haven't read it. Everything here is Nielsen's account, from the section of his roundup titled "Users Freeze Their AI Habits After Only 5 Sessions."

They took donated chat logs from the WildChat corpus and split each message into two parts. What the user wanted (the task) and how they phrased it (the template, like "Please, [request]" against a bare stacked command).

The two parts behave nothing alike.

Tasks grow in a straight line. Users keep bringing new problems, about 0.6 new task types per session, with no sign of stopping.

Templates stall almost at once. Users converge on 2 to 4 of them within roughly 5 sessions. Phrasings tried in those first sessions recurred 5 to 50 times more often for the rest of that user's time on the product. People who typed "hi" early were 49 times more likely to still be greeting the machine months later.

And one number with money attached: every 0.01 increase in early expression diversity went with a 4 to 6% longer expected lifetime on the service.

Nielsen lists the caveats and I'll repeat them. The data is correlational. It has no success metrics. The logs are from 2023 and 2024, on far weaker models. He thinks a 50x reuse effect survives better controls. I think so too, but that's a belief.

His explanation for why it sets so fast is the part designers should read twice. A menu shows you the options you haven't tried. A prompt box shows you nothing. And a wrong menu command gives an error, while a mediocre prompt gives a plausible answer. So people conclude that this is what the tool can do.

The guideline I'd handle with care

He gives five. Treat sessions 1 to 5 as your most valuable design real estate. Vary the form of your example prompts, not only the topic. Keep alternatives visible with persistent controls. Manufacture the missing feedback. Time re-education to model updates.

The fourth is the one I've been burned on. His version: when a user repeats one template across unrelated tasks, offer a side-by-side of their prompt and a better one, with both answers.

I shipped something with the same move in it. It's number 7 in 10 AI Agents I Built That Failed. The Honest Retrospective. An agent watched a new user's first session, noticed where they got stuck, and offered contextual help.

It was technically correct. The stuck points were real. Users hated it. They read it as surveillance, and the activation rate for users who got an intervention was lower than the control group. That cost a quarter of decreased activation, plus the engineering to build it and then take it out.

Not the same feature as Nielsen's side-by-side. Same move, though: the system noticed something about a new user and said so, in the exact window where the user is deciding whether they like the product.

So I'd lean on his third guideline and go slow on his fourth. A tone picker, a format toggle, a one-tap rewrite. Things that sit there and get reached for. That's the logic of Disclosure Without Overwhelm too: decide what a first-time user sees by default and what sits one action away. If you do build the side-by-side, test it on a small cohort and ask how it felt, before you read the click-through.

The same finding, pointed at your own org

Here's where I think the study bites harder than Nielsen says.

Every company rolling out AI tools has an internal onboarding problem with the same shape. And most of them answer it with a calendar invite.

In Kill the AI Office Hours. They're 2026's Agile Transformation. I wrote up an audit of that format in seven product orgs. High attendance. High sentiment. Zero measurable change in shipping velocity, eval discipline, or agent integration. I blamed broadcast: telling people about a tool doesn't put it in their workflow.

The five-session number adds a second reason. Timing. The person in the Friday session got access months ago. Their two or three ways of asking set in the first week, alone, with nobody watching and no error message. You aren't teaching them now. You're asking them to unlearn, and Nielsen's own review of forty years of research says people rarely switch methods once one works.

The replacement I ranked first in May was the paired shipping session. Ninety minutes, two people, one who has shipped with the tool and one who hasn't, building a real thing for the newcomer's team. It only counts if the thing is in that team's workflow within 48 hours.

I never said when to run it. Now I would: inside the first five sessions. It's a person the newcomer trusts, showing a different way of asking, on the newcomer's own work. That's the side-by-side, delivered by someone who was invited.

Basia Kubicka gave the size of the problem last week. She asked about 100 people at Harvard Business School who had ever built a Claude skill, and about 1 in 10 hands went up. She'd expected the reverse. The other 90 aren't beginners. They have habits.

What to do this week

Pull the list of people who got access to an AI tool in the last 30 days. Pair them first. They're the cheapest people in the building to help.

Open your product's first-run experience and count the ways a new user can see a different kind of request without typing one. If the answer is a row of example chips that all read the same, that's one way.

Put a note against the next major model release. In the study, the 246 users who moved from GPT-3.5 to GPT-4 started exploring again. That's your second chance with everybody else.

This sits under the argument on AI Product Management: AI collapsed the cost of the parts we were trained for and left the parts nobody trained for. How a team learns to ask is one of those parts, and it gets decided in the first week whether anyone designs it or not.

Related answer: When should you train a team on a new AI tool?

Sources: Jakob Nielsen, "UX Roundup: Erik the Red, AI Thematic Analysis, Fast Asymptote in AI Skills," UX Tigers, October 9, 2026, the section "Users Freeze Their AI Habits After Only 5 Sessions"; the study by Shengqi Zhu and co-authors is as he reports it. Basia Kubicka, LinkedIn post on her Harvard Business School session, October 7, 2026. Disclosure Without Overwhelm, falkster.com, October 7, 2026.

Share this post

Frequently asked

How quickly do people form AI prompting habits?+

Within about five sessions, according to a study Jakob Nielsen reported on October 9, 2026. Shengqi Zhu and co-authors analyzed 139,535 ChatGPT sessions from 7,955 users and found that people converge on 2 to 4 personal prompt templates within roughly 5 sessions and then stop inventing, while the tasks they bring keep growing at about 0.6 new task types per session. The data is correlational and comes from 2023 and 2024 logs. I have read Nielsen's account, not the paper.

When should a company train people on a new AI tool?+

In each person's first week with the tool, not on a recurring calendar slot months after rollout. If phrasing habits set within about five sessions, a training that reaches someone at session forty is asking them to unlearn. I would use a 90-minute paired shipping session, one person who has shipped with the tool and one who has not, building something real for the newcomer's team.

Should an AI product proactively correct a new user's prompts?+

I would be careful. I built an onboarding agent that watched a new user's first session, noticed where they got stuck, and offered help. It identified real stuck points, users perceived it as surveillance, and activation for users who received the intervention was lower than the control group for a quarter. Keep the alternatives visible and one action away, and let the user reach for them.

What did Jakob Nielsen recommend for the first five sessions?+

Five things: treat sessions 1 to 5 as the most valuable design real estate, vary the form of example prompts and not only the topic, apply recognition over recall with persistent affordances such as tone pickers and one-tap rewrites, manufacture the missing feedback by showing a user's prompt beside an improved version, and time re-education to major model updates.

Why do AI office hours fail to change how a team works?+

They are broadcast, and they arrive late. I audited the format in seven product orgs and found high attendance, high sentiment, and no measurable change in shipping velocity, eval discipline, or agent integration. The five-session finding adds a timing reason: by the time most attendees hear the tip, their way of using the tool is already fixed.

Does a new model release change user habits?+

It can reopen them. In the study Nielsen reports, the 246 users who switched from GPT-3.5 to GPT-4 spontaneously re-explored, like novices again. He recommends spending the teaching budget at major releases. For an internal rollout that means scheduling a second round of paired sessions when a major model ships.

THE SHORT ANSWER

PART OF

AI Product Management

About the author

Falk Gottlob

Falk Gottlob

Product Executive · Founder, Falkster.AI

Thirty years shipping product, from Microsoft Research and Adobe to Salesforce, where he grew Quip into what became Slack Canvas. Four startups, five exits, including a $6.5B healthcare platform and a company Microsoft bought. Four-time Chief Product Officer. Now founder of Falkster.AI, an agentic AI company run by its own agents. This notebook is written from inside the build, not above it.

Comments (0)

Sign in with LinkedIn to leave a comment.

Sign in with LinkedIn
  • Be the first to comment.

Keep Reading

Posts you might find interesting based on what you just read.