Blog · The Agentic CTO — part 6 of 6

Your sceptics are right

The two people blocking your AI rollout are both right — and that's the most useful thing you'll learn this quarter.

Two people in your company are quietly deciding whether your AI rollout happens: the engineer who looks at agent output and sees slop, and the teammate doing silent arithmetic about their own job. The case I want to make is uncomfortable for the hype crowd:

Both of them are right.

And it matters more than the tooling, because none of the machinery in this series — the agent-run SDLC, the learning loops, the reaction ladder, the blast-radius security model, the open repo — survives contact with an org that doesn't want it.

There are two kinds of resistor, and they are not the same person. The first is the greybeard — the engineer who's watched three hype cycles and cleaned up after all of them, and who looks at AI output and sees slop. The second is the person doing quiet arithmetic about their own job. One fear is about quality. One is about survival. Neither is answered by an all-hands slide that says velocity.

Two resistors, both rational, answered differently
The greybeard
“The code is slop.”

Right: raw models made experienced devs 19% slower; 2.6% of them highly trust the output.

→ the harness, then one overnight migration they referee
The worried
“This is how I get replaced.”

Right: some companies did freeze hiring and brag — before reversing.

→ reduce hiring, not people — and context as the moat
One fear is about quality, one about survival. A velocity slide answers neither; each needs its own designed answer.

The greybeards are right about the code

Start by conceding the point, because the data does. A controlled METR study found experienced developers were 19% slower with AI assistance — while believing they were 20% faster. A forty-point gap between vibes and reality. And in Sonar's 2026 developer survey, 96% of developers said they don't fully trust AI-generated code; among experienced developers, "highly trust" polls at 2.6%.

Your greybeard has read these numbers, or lived them. When they say the code is bad, they are not behind the curve. They are describing what a raw model produces when you point it at a serious codebase and hope.

Here's what changed, and it isn't the models: nobody serious ships raw models any more. The discipline now has a name — harness engineering — and the industry's framing is blunt: agent = model + harness, and the harness is the part you own. The result that should make every sceptic sit up: in current benchmark work, mid-tier open models inside a well-built harness match or beat frontier models run naked. Quality is not a property of the model. It's a property of the harness the model runs in.

You've been reading a harness manual for five parts

I didn't have the word when I started this series, but that's what parts 1 through 4 are:

agent = model + harness — and the harness is parts 1–4
Verification loopsmandatory bug-hunt · QA agent drives to a terminal statepart 1
Guardrailsrules enforced at the tool layer, never as prosepart 1
Context & memory~30 versioned skill files, written by the learning loopspart 2
Observabilitymonitor the monitoring · a silent channel is a designed-for failurepart 3
Blast radiusfree-fire staging · read-not-write production · break-glasspart 4
The five layers the harness-engineering literature names, and where this series already built each one. The model is the only part you rent.

The mandatory bug-hunt that fails the build on unanswered findings. The QA agent that opens a real browser and drives every change to a terminal state — in CI, and runnable locally before you push. The lint, the tests, the guardrails enforced at the tool layer because prose instructions don't hold. The learning loops that turn every session's scars into versioned docs. The blast-radius model that makes a bad day recoverable. Point your greybeard at the gates, not at the model. The honest pitch is never "the AI writes good code". It's "nothing reaches you that hasn't survived the harness — and you get to help design the harness."

That last clause matters most. The greybeard's pattern-matching for how code fails is the single most valuable input a harness can get. You're not asking them to lower their standards. You're asking them to encode their standards into something that enforces them at 3am without their involvement — which, framed properly, is the most senior work in the building.

Then run the conversion play

Arguments convert nobody. Overnight results do. Pick the piece of tech debt your team has dreaded longest — the library upgrade three quarters deep in the backlog, the migration nobody volunteers for — and give it to a long-running agent with one instruction: drive it to completion. Review at breakfast.

This play now has receipts at every scale:

The overnight play, receipted at scale
4,500developer-years saved on Java upgradesAmazon
79%of generated changes shipped uneditedAmazon
6 wksfor a migration estimated at 1.5 yearsAirbnb
74%of migration changes written by the LLMGoogle
Tech debt is the perfect conversion instrument: work the sceptic knows cold, resents personally, and can judge from the diff alone.

Amazon ran it on Java upgrades: tasks that took 50 developer-days each now take hours — 4,500 developer-years and $260M saved, with 79% of generated changes shipping without edits. Airbnb migrated 3,500 test files from Enzyme to React Testing Library in six weeks against a 1.5-year estimate — 97% automated, and their headline lesson was pure harness thinking: retry loops beat clever prompting. Google's migration tooling generated ~74% of the changes across 39 internal migrations and halved total migration time.

Why this task, specifically? Because the greybeard is the referee. It's work they know cold, in code they can judge blind, solving a problem they personally resent. You're not asking for faith in a demo — you're handing them a reviewed PR that deletes their least favourite chore. Nobody argues with a green diff on a migration they've been avoiding since March.

The worried are right about the org

Now the harder conversation, and the same rule applies: concede the true part first. The fear isn't irrational — some companies did freeze hiring and brag about it. But watch how that movie actually ended. Klarna stopped hiring, publicly credited its AI, then quietly started rehiring humans because the AI-only version was measurably worse. Duolingo's "AI-first" memo triggered a public revolt — and the CEO walked it back within weeks: "I do not see AI as replacing what our employees do; we are in fact continuing to hire." The companies that framed AI as replacement paid for it in talent, trust and headlines, then reversed.

So here's the version I give founders of fast-moving startups, and it has the advantage of being both kind and true. We reduce hiring, not people. You're a startup: the plan was always exponential. If everyone here gets dramatically more leveraged, we don't need fewer of you — we need the growth we were promising investors anyway, built by the people who already hold the context. And context is the moat: the model can write the code, but it cannot know why the pricing page is shaped like that, which customer broke the last migration, or what the founder actually meant in that Slack thread. Every person in the room is sitting on the one asset that doesn't ship with the weights. Congratulations — you're about to be shockingly senior in a company growing underneath you.

Be honest about the boundary of that promise: it's the startup version. It works because the denominator is growing. A flat company making the same pledge is writing a cheque its P&L has to cash — which is exactly why the Klarna-style reversals happened at companies trying to shrink their way to leverage.

The ratio moves. It doesn't reach zero.

For the shape of the transition, I keep pointing people at autonomous vehicles. Waymo's arc ran: multiple staff per car in testing → one safety driver in every seat → and today, roughly 70 remote operators supporting a fleet of ~3,000 vehicles — about one human per forty cars.

The ratio moves. It never hits zero.Testingseveral humans : 1 carLaunch1 safety driver : 1 carToday≈1 operator : 40 cars×5humans move from operating every unitto owning the exceptions
Waymo today: ~70 remote operators supporting ~3,000 vehicles. Plan careers on the escalation layer, not on one-human-per-step.

Two things are simultaneously true in that arc, and your team needs to hear both. The ratio moved brutally — 1:1 became 1:40, and anyone who planned a career on "every car needs a driver" planned wrong. And: the number never reached zero. The humans didn't vanish; they moved up a level, from operating the vehicle to being the escalation layer for the fleet. That is exactly the two-gates world this series has described — Part 3's whole argument was that being woken should be the last step of the machine's process, not the first step of yours. The role that disappears is doing every step. The role that grows is owning the exceptions — and exceptions are precisely where context beats weights.

Run the two plays in order. Give the greybeards the harness and the overnight migration, because respect plus evidence converts the people hype insults. Give the worried the honest maths — reduced hiring, growing denominator, context as the moat — because the companies that skipped that conversation are the cautionary tales now.

The rollout that fails is almost never the one with the worst tooling. It's the one where the two most rational people in the room were never answered.

That's the series. Agents run the SDLC; the org self-improves; it reacts while you sleep; blast radius makes trust unnecessary; the gates make the repo everyone's; and the people — sceptics included, sceptics especially — come along because you built them answers instead of a slide deck. The machinery took me a year. The consent is the part I'd start earlier.


This completes the series: Part 1 — Agents run your SDLC · Part 2 — Your agentic org self-improves · Part 3 — Your agents react · Part 4 — Securing the agentic org · Part 5 — Make everyone a developer · or start at the hub: The Agentic CTO.

Bring your sceptics along.

The rollout that dies isn't the one with bad tooling — it's the one where the two most rational people in the room were never answered. Designing those answers is coachable work.

Talk to me

← All articles