
Three documents say three different things about what an "active customer" is, the agent picks one, and the team's response is to add a fourth document. Basia Kubicka and Henry Cipolla posted the better response today. I have been running a small version of it for two weeks, and there is one piece I would add.
The short version
Basia Kubicka posted Henry Cipolla's three rules for a knowledge harness on October 5: a side agent re-reads old threads for stale rules and attaches the evidence, open questions go to the person who owns the answer, and every definition change is a versioned diff with a name on it. A small agent can have all of that with one file in version control. Mine is a JSON file that defines "relevant" for a reading agent across 43 sources. This week it read 136 posts and kept 45, and its last change was a one-line diff on September 28 with a date, a name, and the post that proved it. The piece to add is the reject pile: definitions fail in what they exclude, and I found every bug in mine by reading what it dropped. The piece I am missing is the version stamp on each report.
What they posted
Kubicka wrote it up with Cipolla, whose team at Pedestal has, per her post, deployed agents at Fortune 500 supply chain, manufacturing, and retail companies and now runs them at 97%+ accuracy. I read her post in full. I have not read Pedestal's own write-up, so that figure is her report of their result.
The diagnosis is the part to keep. When an agent gives a wrong answer, most teams add more documents and build a bigger retrieval layer. But the agent was rarely missing the fact. It had three versions of it and no way to tell which one was current.
Three rules follow.
It reads last quarter, not last week. A separate agent re-reads old threads looking for gaps, two teams using one word differently, and rules the business has moved past. Every finding arrives with the chats that proved it.
It asks the person who owns the answer. Open questions go to the human who knows, in Slack or Teams. The reply is the evidence and their name travels with it.
The gate shows a diff, and it is version controlled. A change sits as a pending row under the one it would replace. Approve, reject, or send back. Approval bumps the version, the old copy stays readable, and every saved report points at the version it used.
All three end the same way, she writes: a person's name on a definition.
The one-file version
My reading agent scans 43 sources every week and decides which posts are worth my attention. The whole decision hangs on one word, "relevant," and that word is defined in one place: a JSON file with five clusters, each carrying a keyword list and a set of tags.
This week it read 136 posts and kept 45. It dropped 91.
The rule I set on day one, in I Automated My Reading. The Filter Was Wrong Three Times., was to fix the vocabulary and never the code. When the filter gets a post wrong, the correction goes into the definition file. No special case for a source, no private exception only I understand.
Here is what one change looks like. On September 28 the filter scored a Marily Nika flashcard, on when it is safe to ship hallucinations, at zero. That is ship or no-ship by consequence, detectability, and recovery, which is about as close to my subject as a post gets. The word "hallucination" was not in the file.
The fix was one line. One tag added to one cluster. The commit carries a date and a name, and the note in that day's digest names the post that exposed the gap and says what else the change touched.
Hold that against the three rules. The finding arrived with its evidence. There is a name on the change. It is a diff, the old version is readable, and it can be rolled back. That is a knowledge harness for one definition, and the tooling is git.
It is small. Cipolla's team does this across an enterprise, with agents hunting for the conflicts and routing questions to owners. I am one person with one file. The shape is the same, and the shape is what a team can start on this week.
What I would add: the reject pile
Rule one sends an agent back through old threads to find conflicts. That looks at what is there.
A definition also fails in what it leaves out, and nothing in a thread will tell you. The Nika post was never in a thread. It was in the discard pile with a score of zero.
All three bugs from the filter's first day were the same kind. An enterprise-agents case study dropped while a demand-generation post was kept. A piece on a coding agent scored zero because the vocabulary had concepts and no product names. A company flagged as a gap because it appeared in 45 posts as a line in a list. Every version ran correctly and produced a clean, plausible keep list. I found each one by reading the rejects.
So the filter prints a drop count per source on every run, and a flag lists every rejected title with its score. I learned the same thing from the other side in Every Check Was Green. Four Sources Were Dead.: the confident, well-formed output is where the stale input hides.
For an agent that answers from company knowledge, the reject pile is the questions it declined, the documents retrieval skipped, and the records a definition of "active" excluded. Sample ten a week. That is where the next definition change is.
The drop
- One file per definition, in version control. "Active customer," "resolved," "relevant." If the agent relies on the word, the word has a file.
- No change without a name and the example that proved it. "The system decided" is not an owner, as Kubicka puts it.
- Print the drop count on every run, and keep the rejects one flag away. The daily output stays clean. The discard pile stays reachable.
- Stamp each report with the definition version it ran on. So a change flags drift in old numbers and does not silently move them.
I do not do the fourth. My digests are dated and the file's history is dated, and when I need to know which vocabulary produced which digest I line them up by hand. That works at one change a week. It would not work at ten, and it is the first thing I would fix before handing this to a team.
This sits in Enterprise AI Agents, where the claim is that what matters is whether a deployed agent still does production work after ninety days. An agent whose definitions drift without a name on them is one that quietly stops.
Pick one thing this week. Find the word your agent gets wrong most often, put its definition in a file, and write your name next to it.
Related answer: What is a knowledge harness for AI agents?
Sources: Basia Kubicka with Henry Cipolla (Pedestal), "Your agents don't need a bigger knowledge base. They need a knowledge harness," LinkedIn, October 5, 2026. I Automated My Reading. The Filter Was Wrong Three Times., falkster.com, September 20, 2026. Filter figures are from this site's monitoring script on 2026-10-05: 43 sources, 136 posts in the 2026-09-28 to 2026-10-05 window, 45 retained. The definition change is commit 9ebed52 of 2026-09-28.
Also on Medium
Full archive →AI Agents and the Future of Work: A Pixar-Inspired Journey
What product managers can learn about AI agents from how Pixar runs a film team.
Many AI Agents Are Actually Workflows or Automations in Disguise
How to tell agents from workflows from cron jobs, and why it matters for what you ship.
Frequently asked
What is a knowledge harness for AI agents?+
The term is from Henry Cipolla's team at Pedestal, as described by Basia Kubicka on 2026-10-05. It is the set of rules around what an agent treats as true: a side agent re-reads past threads for gaps and stale rules and attaches the evidence, open questions go to the named person who owns the answer, and every change to a definition is a pending diff that gets approved, versioned, and kept readable after it is replaced.
Why do agents give wrong answers when the knowledge base is large?+
Kubicka's summary of Cipolla's experience is that the agent was rarely missing the fact. It had three versions of it and no way to tell which was current. More documents add versions. A named owner deciding which definition is the real one removes them.
Can you build a knowledge harness without buying a platform?+
For a small agent, yes. Put each definition the agent relies on in one file under version control, require a name and a proving example on every change, print how much the definition excluded on every run, and stamp each output with the definition version. My reading agent's definition of relevant is one JSON file with five keyword clusters, and its changes are commits.
What is the reject pile and why does it matter?+
It is everything the definition excluded on a run. A keep list is selected to look right, so errors in a definition show up in what was thrown away. All three bugs in the first day of my relevance filter were found by reading rejected titles, and the 2026-09-28 fix came from one rejected post that scored zero.
What should a definition change record?+
Who made it, when, the diff, and the example that proved it was needed. On 2026-09-28 my filter's change was one tag added to one cluster, committed with a date and a name, with a note naming the Marily Nika post that had been dropped.
What is the step most small teams skip?+
Stamping each report with the definition version it ran on. I skip it too: my digests are dated and the definition file's history is dated, and matching them is manual. Cipolla's third rule covers it, so that a definition change flags drift in old reports.

Comments (0)
Sign in with LinkedIn to leave a comment.
Sign in with LinkedIn