
Six researchers at UC Berkeley published a study on September 16 called HarnessTax. Melissa Z. Pan, Shuo Yang, Negar Arabzadeh, Wei-Lin Chiang, Ion Stoica, and Matei Zaharia ran seven models through three coding harnesses, Claude Code, Codex CLI, and Pi, on 30 tasks each from SWE-bench Lite and Terminal-Bench 2.0. Claude Fable 5 solved 97.8 percent of attempts in Claude Code, 96.7 percent in Codex, and 96.7 percent in Pi. Claude Code cost about twice what Pi did per attempt, $1.33 against $0.67. Their sentence: "The same model can achieve similar success rates at up to 5x costs."
Tomasz Tunguz read it the next day and did the margin math. GPT-5.6 Sol costs 71 percent less on Pi than on Claude Code for the same result. None of the 42 harness comparisons showed a statistically significant quality difference. Cost spread, depending on harness, ran from 1.1x to 5.1x. His worked example, an illustration rather than a case: two startups bidding the same $250,000 contract, one at 38 percent gross margin and one at 75, decided by the harness alone.
It got 231 points and 97 comments on Hacker News, nearly all about coding agents. Nobody said the quiet part. The harness is what a forward deployed engineer changes. Tools, skills, context, the loop. That layer sets what every answer costs and barely moves quality. So the FDE who adds three tools and a 2,000-word skill for the customer's vocabulary on a Tuesday just made a margin decision for that account. No pull request, no review, no price. It was "just configuration."
This is part 4 of The FDE Transition. Part 3 was what an FDE owes product. This part is the contract with engineering: what an FDE may change on site, what lives in the harness versus code, and what engineering owes back.
The short version
Three layers, three rules. The customer layer, prompts, rules, thresholds, connector settings, and the customer's eval rows, belongs to the FDE, changed the same day, no review, on two conditions: it's versioned, and it diffs against the shipped default in one command. The harness, the tool set, the loop, context assembly, guardrails, and model choice, belongs to engineering, and an FDE changes it only by pull request, with the customer's eval rows and the cost per case attached, because a harness change is a cost change for every customer on it. The core belongs to engineering, no forks, ever. Engineering owes four things back: a configuration that diffs, a sandbox where a cost number means something, an eval run on every customer's rows on every harness change, and cost per resolved case per customer on the dashboard. Nabeel Qureshi's Palantir had a clean version of this line, FDEs overfit and product engineers generalize, and it held because the two groups wrote different artifacts. In an agent product the line runs through one file, which is why it has to be written down.
Why the line moved
Palantir's split worked because the artifacts were different. Per Reflections on Palantir, the FDE built "Asana, but for building planes" at Airbus, and product development engineers productized it. Qureshi's description of the code is the honest one: FDE code "gets the job done fast, which usually means, politely, technical debt and hacky workarounds," and PD code "scales cleanly, works for multiple use cases, and doesn't break." Two kinds of artifact, and the deployment-to-product loop between them.
In an agent product there's one artifact. A prompt. A skill file. A tool description. One Hacker News commenter measured Claude Code's opening instructions to the model at 80 kilobytes, most of it tool descriptions. That's harness. The tool an FDE adds for the customer's ticketing system and the skill that explains their approval hierarchy are harness too, on every call that customer makes. Tunguz's line is the one to underline: "Harnesses coalesce common workflows into repeatable patterns: deterministic code or skills." Code or skills. That "or" is the whole contract. A skill an FDE writes on site is a harness change, and nobody has said which changes are the FDE's to make and which need a second pair of eyes and a price.
I've run product on top of deployment teams, and the failure always had the same shape. The person on site makes the change that unblocks the customer, and nobody upstream knows it exists until it shows up as a bill. The fix was never "stop letting them change things." It was a written line.
Layer one: what the FDE owns outright
The customer layer. Prompts and vocabulary, business rules, thresholds, routing, connector settings, and the customer's eval rows. The FDE changes it the same day, with no review, because that's the reason to have an FDE. Part 2 made the point: if a prompt change is a ticket and a release train, you have a queue with a travel budget.
Two conditions. It's versioned. And it diffs against the shipped default in one command, because that diff is the delta in the handback packet, and a delta the FDE has to remember instead of generate doesn't get sent.
Decagon's Jesse Zhang argues that needing FDEs at all is a sign of a weak product, and claims a customer he shares with Sierra built three workflows in a year with Sierra's embedded engineers and seven in a month after switching. I can't check the customer. The claim is about layer one: how much of the product lives there, and whether the customer's own people can drive it. Decagon's bet is that the layer can be pushed out until the FDE isn't needed. Perspective AI's survey, where 64 percent of FDEs described their main work as deploying with significant customization, says nobody's layer one is big enough yet.
Layer two: the harness, by pull request, with a price
Tool set, agent loop, context assembly, guardrails, model choice. Engineering owns it. An FDE can change it, and often should, because the FDE knows the customer's ticketing API times out under load. But the change goes in as a pull request, and the PR carries two things it usually doesn't: the customer's eval rows that motivated it, and the cost per case before and after, measured in the sandbox.
HarnessTax has the detail that should haunt every FDE who reaches for a new tool. Pi reached the Pareto frontier on both benchmarks with four tools: read, write, edit, and bash. Every tool description is tokens on every call, and the FDE's instinct at a customer is to add. The PR is where that instinct meets a number: this adds a tool, it moves resolution on this customer's rows by this much, it costs this much more per case. Engineering decides inside the same two-week window product gets for the handback: generalize it for everyone, keep it as a per-customer skill, or decline. If the window closes without an answer, the FDE ships it as a skill in layer one anyway, and now it's a harness change nobody priced.
Leo Mehr grew Ramp's FDE team from two engineers to sixteen in eighteen months, inside engineering, working on the core product. His Ramp post has the rule I'd put at the top of the PR template: always be scoping. Who's driving the urgency, who actually uses it, will other prospects need it? In his AI Engineer talk he says the team is heading toward spending most of its time operating harnesses and maintaining evals, and he warns what an unscoped request produces once automated, a "token maxing slop cannon." If the FDE is becoming a harness operator, and I think that's right, the PR-with-a-price rule isn't a constraint on the job. It's the job.
Layer three: the core, and the word fork
The FDE never forks the core. Not for a demo, not for a deadline, not because the customer's CTO is in the room. A fork is a confession that layer one is too small, and the fix is a layer-one or layer-two PR that moves the boundary for everyone. The handbook chapter on what has to change in the product has the test, and I'll restate one line only: if any of the last three things an FDE built needed a core PR, the boundary's in the wrong place.
One exception: an FDE embedded with a product squad for a week to move that boundary. That's product engineering with a visitor badge, not deployment.
What engineering owes
Four things, and a clock.
A configuration that diffs. Per customer, versioned, one command to produce the delta against the shipped default. Substrate-first engineering, applied to the seat.
A sandbox close enough to the customer that a cost per case measured there means something. Neej Gore, Zeta's chief data officer, drew the line at VentureBeat between the sandbox, where an engineer installs a missing component and feeds it back so it ships again, and the mud, where the same component gets hand-built at every customer. A contributed piece, so a vendor's view, and the distinction holds anyway.
An eval run on every customer's rows on every harness change. A harness change touches every customer's cost, and per HarnessTax sometimes their quality: an alternative harness had the highest success rate in nine of twelve model comparisons. The rows come from the handback packet. If part 3 doesn't happen, engineering is changing the harness blind. The eval is the spec covers the general case.
Cost per resolved case, per customer, on the dashboard next to resolution rate. Gore wants productization lag on it too, the time from a field discovery to a tested capability the next customer can use. The economics chapter explains why the margin depends on it.
And the clock. A review window for FDE pull requests, two weeks at most, owned by a named engineer the way the handback is owned by a named product person. Engineering that lets FDE PRs sit gets skills in layer one it never priced.
What this means in the seat, and at the top
If you're the FDE: you've always known which tool the customer actually needs. Now you can price it. A PR with the rows and the cost delta attached is the most persuasive document in the building, and it gets you into the room where the harness gets generalized.
If you run engineering: the harness is a product with a profit and loss per customer, and the FDE is its most active contributor. Tunguz's closing line is the brief: "Effective harnesses are not cocktails of exotic ingredients: a deep customer understanding, a collection of relevant evals & a factory for automating hill climbing." The FDE is the customer understanding and the evals. Build them the factory, and write down what they may touch.
One thing to try this week: pick one live customer and diff their configuration against the shipped default. For every line that touches a tool, the loop, or the context the model sees, ask whether engineering knows it's there and what it costs per case. Count the lines nobody priced. That's the size of the missing contract.
Sources: Melissa Z. Pan, Shuo Yang, Negar Arabzadeh, Wei-Lin Chiang, Ion Stoica, and Matei Zaharia, "HarnessTax: How Much Does the Harness Matter for Coding Agents?", UC Berkeley via Arena (September 16, 2026, updated September 18) · Tomasz Tunguz, "The Harness Margin Opportunity," tomtunguz.com (September 17, 2026) · Hacker News discussion of HarnessTax (September 2026) · Nabeel Qureshi, Reflections on Palantir · Leo Mehr, "Forward Deployed Engineering," Ramp Builders Blog (August 5, 2025, updated February 12, 2026) · Leo Mehr, "How Forward Deployed Engineering is done at Ramp," AI Engineer talk (video) · Neej Gore, "Forward-deployed engineering is how enterprise AI learns," VentureBeat (September 2, 2026) · PYMNTS, "Enterprise AI's Hottest Job Just Found Its Biggest Skeptic" (August 14, 2026), with Jesse Zhang quoted · Perspective AI, "State of Forward Deployed Engineering 2026: Survey of 1,500 FDEs" (May 18, 2026)
Part of the running argument on AI Product Management: once building is cheap, the scarce work is deciding what gets built for everyone, and the codebase contract is where engineering and the seat agree on who decides.
Related answer: What can a forward deployed engineer change in the codebase?
Frequently asked
What can a forward deployed engineer change in the codebase?+
Three layers, three rules. The customer layer (prompts, business rules, thresholds, routing, connector settings, and the customer's eval rows) belongs to the FDE: changed the same day, no review, on two conditions, that it is versioned and that it diffs against the shipped default in one command. The harness (tool set, agent loop, context assembly, guardrails, model choice) belongs to engineering: an FDE changes it only by pull request, and the PR carries the customer's eval rows and the cost per case measured in the sandbox. The core belongs to engineering only, and the FDE never forks it.
Why does the HarnessTax study matter for FDEs?+
Because the harness is the layer an FDE touches on site. HarnessTax, published September 16, 2026 by Melissa Z. Pan, Shuo Yang, Negar Arabzadeh, Wei-Lin Chiang, Ion Stoica, and Matei Zaharia, ran seven models through Claude Code, Codex CLI, and Pi on SWE-bench Lite and Terminal-Bench 2.0. Claude Fable 5 solved 97.8 percent of attempts in Claude Code and 96.7 percent in both Codex and Pi, while Claude Code cost about twice Pi per attempt ($1.33 versus $0.67). Tomasz Tunguz's read: 71 percent less for GPT-5.6 Sol on Pi than on Claude Code, no statistically significant quality difference across 42 comparisons, and a 1.1x to 5.1x cost spread. Every tool and skill an FDE adds is a harness change, so it is a cost change for that customer, and the contract has to say who approves it.
What does engineering owe a forward deployed engineer?+
Four things. A per-customer configuration that is versioned and diffs against the shipped default in one command, so the delta in the handback packet is generated rather than remembered. A sandbox close enough to the customer that a cost per case measured there means something. An eval run on every customer's rows on every harness change, because a harness change touches every customer's cost and sometimes their quality. And cost per resolved case, per customer, on the dashboard next to resolution rate. Plus a review window: an FDE pull request that sits for weeks becomes a skill in layer one that nobody priced.
Should an FDE ever fork the core product for a customer?+
No. Not for a demo, not for a deadline. A fork is a confession that the customer layer is too small, and the fix is a layer-one or layer-two pull request that moves the boundary for everyone. The one exception is an FDE embedded with a product squad for a week to move that boundary, which is product engineering with a visitor badge, not deployment. The test from the handbook still holds: take the last three things an FDE built for a customer, and if any needed a core pull request, the boundary is in the wrong place.
What did Palantir's split between FDEs and product engineers look like?+
Per Nabeel Qureshi's Reflections on Palantir, forward deployed engineers went on site three or four days a week and solved the customer's problem without worrying about overfitting, and product development engineers took whatever the FDEs built and generalized it to sell elsewhere. FDE code got the job done fast, which he politely calls technical debt and hacky workarounds, and PD code scaled cleanly. That line was easy to keep when the two groups wrote different artifacts. In an agent product the line runs through one prompt, one skill, one tool description, which is why it has to be written down.
What is The FDE Transition series?+
A twice-weekly series on falkster.com for forward deployed engineers and the leaders building around them. Part 1 is the thesis that the FDE is the Product Builder arriving from services. Part 2 argues the FDE shortage is a product problem. Part 3 is the handback to product. Part 4 is the codebase contract with engineering. The parts that follow cover who owns the promise with sales, month two with customer success, the FDE scorecard, career, comp, buying from an FDE company, and when not to hire FDEs at all.

Comments (0)
Sign in with LinkedIn to leave a comment.
Sign in with LinkedIn