FoundationNew·Falk Gottlob··11 min read

Context Is King. Written Context Is Rented.

Everyone is racing for the context layer. The quadrant nobody occupies is empty for a structural reason, and importing your wiki will not fill it.

context layerenterprise AIAI agentsEvan ArmstrongNotionGleanServiceNowmoatsHeidi
Helpful?

Foundation-pink Falkster cover: a row of filing cabinets rolling out the door on casters while one ledger stays bolted to the floor, its open pages still filling with fresh ink.

John Hurley at Notion posted a chart a few days ago that has been making the rounds. 567 reactions, Marty Cagan in the likes. Two axes: agent capability going up, depth of existing organizational knowledge going right. Anthropic and OpenAI top left, building agents first. Notion, ServiceNow, and Glean bottom right, knowledge first. LangChain and MCP down in the orchestration basement where plumbing goes to commoditize.

The top right box, the one actually labeled "the context layer," is empty. A question mark and the words "no clear winner yet."

I've been staring at that box for a week, because it's the box I'm building in. It isn't empty because the race is young. It's empty because nothing on the horizontal axis gets you there.

The short version

Evan Armstrong's Context is King is the sharpest piece written on enterprise AI this year, and the chart everyone is sharing comes out of it. His argument holds. Software is splitting into three layers, systems of record and point solutions are both commoditizing, and Christensen's Law says the margin migrates to the layer in between, the one holding organizational meaning and institutional judgment. Where I'd push is on what actually constitutes that layer. Evan writes one sentence about execution traces feeding back to make the next run smarter and files it under "best of all," a nice property of the category. That sentence is the whole category. Everything else being fought over on that chart is a snapshot, and snapshots port. The best objection to the thesis, raised in Evan's own comments, is that companies will just carry their context to whatever vendor is cheapest next quarter. That objection is right about written context and wrong about earned context, and the difference between the two is the entire moat. What doesn't port is the graded record of what agents did inside your tenant, which acts a human reversed, and which rules survived contact with your outcomes. Nobody on that chart has it yet. That's why the box is empty.

What Evan gets right

The economics aren't in dispute. Public SaaS growth halved from 36% to 17% since 2023, AI-first applications run at 50 to 65% gross margins instead of the 75 to 85% that justified SaaS multiples, and switching costs are collapsing because LLMs made data migration trivial. Growth, margin, and terminal value compressed at once. The repricing I wrote about earlier this year was the market noticing.

His three layers are the useful part. Systems of record survive as regulated, auditable databases that stop commanding premium multiples. Point solutions, the per-seat interface layer, are the kill zone. And the new middle holds what he calls the directing layer: the institutional knowledge that tells agents what to do, in what order, and whether they're allowed to do it.

He's also honest about the hard part, which most people writing about this skip. A markdown file can describe your sales process. It can't encode that deals over $500K stall when legal reviews before procurement, or that your best AE skips discovery for inbound leads from existing customers. That knowledge doesn't live in a document. It emerges from thousands of executions. His word for the document version is snapshot, and it's the right word.

Then he writes this and moves on: every time an agent executes a workflow, it generates traces that feed back into the context layer, making the next execution smarter. Best of all, he says.

Not best of all. That's the thing itself. Strip it out and the context layer is a wiki with retrieval, a product category that already exists and already commoditized.

The migration objection

The strongest pushback showed up in Evan's own comment section, from a reader named Artas Bartas. If it's trivial to move information between platforms now, what stops a company from carrying its context to the cheapest vendor of the moment? And if nothing stops them, how does value accrue to the layer at all?

Take it seriously. It's fatal to most of the positions on that chart.

Your wiki ports. Process docs port. The permission model ports, the Slack archive ports, the onboarding material ports, and by Evan's own argument LLMs are what made all of it portable. If your context layer is a very good pile of documents with a retrieval index on top, you've built the thing a competitor can import over a weekend and a foundation model can read directly without you. Which is the uncomfortable position Notion, Glean, and to a lesser extent ServiceNow are arguing from. Their asset is real and it's deep, and it's also the most liquid asset in the stack.

Here's the other list. Which entities in your business are actually the same entity, and how confident the system is in each of those joins. Which acts an agent took, which ones a human reversed, who reversed them, and how fast. Which rules got promoted from a hunch to a default, and which ones got retired because the outcome data stopped supporting them. Which verifier caught which class of failure. What your organization means by "resolved," written by you rather than assumed by the vendor.

None of that is a document. It's an observation record, keyed to your entities and your outcomes, and it only exists if something was acting in your business long enough to accumulate it. You can't export it in any form that's useful, because the value sits in the joins and the grading rather than the rows. A new vendor would have to re-earn all of it in production, on your traffic, while you watch and decide whether to trust it.

So switching costs didn't disappear. They moved. They used to live in the database and now they live in the decision record. Better place for them if you're a buyer. Much harder place to reach if you're a vendor, which is exactly why the quadrant is empty.

What I got wrong building this

I apply what I write or the post is spectating, so here's the version with my name on it.

Heidi shipped a Signals layer that was display only. The system noticed things, surfaced them cleanly, and stopped. It demoed beautifully. It compounded nothing, because signals sitting next to each other on a screen aren't correlated with each other, and a signal that never triggers an act never generates the trace that would make the next act better. Open loop with a nice front end.

We rewrote it rather than extending it. Which cost a sprint I didn't have.

What came out the other side: derived composite signals instead of raw ones, outcome labels that close the receipt loop on every act, and autonomy gated on three things at once. How confident the system is. How reversible the action is. How strong the evidence behind the rule actually is. High confidence on a weak rule doing something irreversible should not execute, and the single-axis version of that gate lets it through every time. This is the eval as the spec pushed into the runtime, where the grading never stops.

And the thing I killed matters more than anything I built. Statistical exploration is out of scope, permanently. No bandits, no cross-tenant learning, none of the standard machine learning answer to "which rule is better." Any one tenant's N is too small to mean anything, and pooling across tenants to get the sample size collides head on with the isolation promise, which is the product. So the design assumes sparse evidence and treats human confirmation as the evidence grade. Slower to learn. Honest about what it knows. It's the only version that survives a security review, and I'd rather ship the version that survives.

One line runs underneath all of it: rent the model, own the company brain. The model is a commodity I'll swap next quarter without ceremony. The observation record inside your tenant isn't.

Who takes the empty quadrant

Four credible paths, and each deserves a fair reading.

Notion comes from documents and wikis, with real depth in the companies that run on them and an agents product aimed directly here. ServiceNow comes from workflow and process logic, with the deepest enterprise penetration of anyone in the fight and partnerships with both OpenAI and Anthropic. Glean comes from retrieval, betting that "find what we already know" extends naturally into "know what we mean." OpenAI Frontier and Anthropic's Cowork come down from the top, betting that if you build capable enough agents the organizational knowledge accumulates to you.

The foundation-model bet is the strongest of the four, and not because the models are better. It's the only one that starts from the act rather than the archive, and acts are what generate the record. Everyone else has to bolt execution onto a library.

Notion is hosting Evan alongside Snowflake on the 16th to argue about this in public, which I'll be watching. Though consider who's in the room. A company whose asset is the wiki has a structural interest in the wiki being the answer. Everyone on that chart does, me included, which is why the falsification test at the end of this post is the part I'd hold me to.

What I'd say to all of them is the same. The empty quadrant doesn't get filled by the deepest archive or the smartest agent. It gets filled by whoever first runs a loop where acting generates evidence, evidence changes the rules, and changed rules change the next act, with the whole chain legible enough that a customer signs off on more autonomy next quarter than they did this one. That's an infrastructure problem before it's a knowledge problem. Receipts on every act, a reversibility taxonomy, entity resolution a human can inspect and correct, outcome definitions the customer writes.

Boring work. Nobody demos it. It's the only part that compounds.

Where this thesis could break

I don't write hagiography, including of my own arguments, so here's the risk ledger.

The record might be thinner than I think. If agents mostly execute short, bounded, well-specified tasks, the traces they generate may carry very little signal beyond what a good process doc already states, and the whole earned-context distinction collapses into a rounding error. That's an empirical question and I don't have the data to settle it.

Frontier model capability could eat the middle. If models get good enough at in-context reasoning over raw organizational data, the graded intermediate layer looks like scaffolding, and scaffolding built ahead of the models is scaffolding you throw away. That's the Sierra lesson pointed at me rather than at someone else.

And the incumbents can buy their way to the record. Nothing stops ServiceNow from instrumenting execution inside workflows it already owns. They start with the distribution and the process logic, and they only need the loop. If one of them ships it before the loop-first companies ship distribution, the archive wins on the merits.

Notice what all three have in common. They're arguments about whether the record accumulates fast enough to matter. None of them is an argument that written context is a moat, and I've yet to see a good version of that one.

Try this week

Pick the one that matches your seat.

If you're running agents in production already, take a single workflow and answer three questions about the last thirty days. Which acts got reversed. Who reversed them. Did anything in the system change because of it. The first two are logging and most teams have them. The third is the loop, and almost nobody does.

If you're buying in this category, ask the vendor what they'd hand you on the way out. If the honest answer is a document export, you're buying a very good wiki and should price it accordingly.

If you're building here, stop measuring your context layer by what it contains and start measuring it by what it revises. Contents are a library. Revisions are a loop.

If you can't answer the third question, you don't have a context layer. You have an agent deployment with good search.

Sources: Context is King, Evan Armstrong, The Leverage, February 2026. Context is King: How to build a system of context for AI, Notion webinar with Evan Armstrong (The Leverage), David Tibbitts (Notion), and Rajhans Samdani (Snowflake), September 16, 2026. Quadrant chart and framing via John Hurley, Notion, on LinkedIn.

Share this post

Also on Medium

Full archive →

Frequently asked

What is the context layer?+

Evan Armstrong's term for the new middle of the software stack, sitting between systems of record (the database layer) and point-solution applications (the interface layer). It holds organizational meaning: what agents should do, in what order, and whether they are allowed to do it. His argument is that code and databases are commoditizing, so by Christensen's Law of Conservation of Attractive Profits the margin migrates to the layer between them.

If context is easy to export, how can it be a moat?+

Written context is easy to export and therefore is not a moat. Documents, wikis, process descriptions, and permission models all port, and LLMs made porting them dramatically easier. What does not port is the earned record: which entities resolved to the same entity and with what confidence, which agent actions were reversed and by whom, which rules got promoted or retired against outcome data, and what the organization means by resolved. That record only exists if something was acting in the business long enough to accumulate it.

Who is winning the context layer right now?+

Nobody, which is the interesting part. Notion comes at it from wikis and docs, ServiceNow from workflow and process logic, Glean from enterprise search, and OpenAI Frontier and Anthropic's Cowork from agent capability downward. The foundation-model path is the strongest of the four because it starts from the act rather than the archive, but no one has shipped a closed loop where acting generates evidence and evidence changes the next act.

What is the difference between a context layer and an agent deployment?+

Whether the loop closes. An agent deployment reads context and takes action. A context layer takes action, records the outcome, grades the evidence, and changes what it does next because of it. Test it on one workflow: for the last thirty days, can you say which acts were reversed, who reversed them, and what changed in the system as a result. If the third answer is nothing, the loop is open.

Why not use bandits or reinforcement learning to grade agent rules?+

In a multi-tenant product, any single tenant's sample size is too small for statistical exploration to mean anything, and pooling across tenants to reach significance collides directly with the tenant isolation promise. In Heidi I ruled it out of scope permanently and designed for sparse evidence instead, with human confirmation as the evidence grade. Slower, and it survives a security review.

What should I do this week if I am running agents in production?+

Pick one workflow and answer three questions about the last thirty days: which acts got reversed, who reversed them, and did anything in the system change because of it. The first two are logging. The third is the loop. If you cannot answer it, you have an agent deployment with good search, not a context layer.

THE SHORT ANSWER

About the author

Falk Gottlob

Falk Gottlob

Product Executive · Founder, Falkster.AI

Thirty years shipping product, from Microsoft Research and Adobe to Salesforce, where he grew Quip into what became Slack Canvas. Four startups, five exits, including a $6.5B healthcare platform and a company Microsoft bought. Four-time Chief Product Officer. Now founder of Falkster.AI, an agentic AI company run by its own agents. This notebook is written from inside the build, not above it.

Comments (0)

Sign in with LinkedIn to leave a comment.

Sign in with LinkedIn
  • Be the first to comment.

Keep Reading

Posts you might find interesting based on what you just read.