Guardrails are the licence to go fast
Every team that ships a thousand agent-written changes a week has more process than the team that ships ten. Boris Cherny's adoption ladder names why: each step up requires a control you have to build before you can climb.
Boris Cherny, who created Claude Code, published a maturity model on 16/07/2026 describing the steps he keeps seeing as organisations adopt AI. It is four days old as I write this and it has already reframed an argument I have been having for months.
His five steps, with his rough agent counts:
| Step | Name | Agents | What the human is |
|---|---|---|---|
| 0 | Gated | 0 | Locked out of capable models |
| 1 | Assisted | 1 per person | A pair |
| 2 | Parallel | ~10 per person | An orchestrator |
| 3 | Supervised autonomy | ~100 | A manager of managers |
| 4 | AI-native | 1,000+ | An intent-setter |
Cherny says Anthropic operates at step 3 and that he personally reached step 4. His starting observation will be familiar to anyone running an engineering organisation right now. One person is getting ten times the output. The rest of the company has not moved.
The part I want to pull out is that the role names matter more than the agent counts, and each transition has a prerequisite you cannot skip.
What each step actually requires
Going from assisted to parallel means one person stops reading keystrokes and starts reading final diffs. That only works if something else checks the code, because the human attention that used to do it has been spent. So step 2 requires an automated verification loop you actually trust. Not a linter. Tests, security scanning, and a review pass that runs on every change.
Going from parallel to supervised autonomy means agents start work that other agents finish, and the work crosses team boundaries. Now an agent is touching code its originating human does not own. So step 3 requires a registry: something that knows what every tool and agent is, who owns it, what data it touches, and how much autonomy it is allowed. Without that, the first cross-team agent action that goes wrong ends the programme.
Going from supervised autonomy to AI-native means pipelines start themselves from signals and schedules. Nobody is watching at the moment of action. So step 4 requires everything below it, proven, plus an audit trail that reconstructs what happened without anyone having been present.
Read that sequence and the conclusion is uncomfortable for a certain kind of engineering culture. The teams shipping the most agent-written code have the most process, not the least.
The evidence points the same way
The published accounts are consistent on this, and it is worth being specific because the reverse claim gets made a lot.
Stripe merges over 1,300 PRs a week with zero human-written code in them. Every single one is human-reviewed. They also isolated the devboxes from production and the internet, and they run a three-million-test suite. That is more control than most teams apply to human PRs.
Shopify's CEO made reflexive AI use a performance-review expectation in April 2025. In the same period, senior-engineer PR review stayed mandatory, and all model traffic was routed through one proxy with per-user cost alerts.
StrongDM went the furthest, removing human code review entirely. They did not replace it with nothing. They replaced it with thousands of end-to-end scenarios per hour run against behavioural clones of Okta, Jira, Slack, and Google Workspace. The claim is not "we trust the agent". The claim is "we built a validation suite strong enough that reading the diff adds nothing".
McKinsey's QuantumBlack architecture puts the human in exactly one place, at the pull request, reviewing a complete feature. One review point, deliberately chosen, with a deterministic orchestration layer controlling everything around it.
None of these is a story about removing controls. They are all stories about moving controls from human attention, which does not scale, into machinery, which does.
Risk tiers, or how to stop arguing per tool
Here is the practical problem with "every change passes the same gates". Applied literally it means a Slack bot that posts a daily digest gets the same treatment as a service that writes to the billing database. That is how you end up with a process everyone routes around.
The fix is to derive strictness from the thing itself rather than negotiating it. Classify each tool automatically from properties you can observe: what data class it touches, what it can write to, who can reach it, and how much autonomy the agent has. That classification sets which checks are required.
Loading diagram...
Two properties make this work, and both are easy to get wrong.
The builder cannot pick their own tier. The moment self-selection is possible, everything becomes tier 1. The tier has to fall out of what the tool touches, not what its author thinks it deserves.
The low tier must be genuinely fast. A read-only dashboard over data the requester already has access to should go from description to running in a day, fully automated, no committee. If it does not, people build it themselves instead and you are back to the shadow inventory. This is the number I would put on a wall: time from idea to running tool on the lowest tier. If that number is worse than doing it by hand, the platform has failed regardless of how good its gates are.
Strictness gets spent only where the data warrants it. That is the whole trick, and it is also what makes it reasonable to let a BI analyst build things. Without tiers, handing non-engineers a build pipeline is reckless. With them, a read-only dashboard is a read-only dashboard no matter who described it.
The audit side effect
There is a benefit here that nobody plans for and everybody appreciates later.
If every change goes through the same recorded path, the evidence for an audit is generated as a side effect of building. Who approved what, what checks passed, when it deployed, what data it touches, who has access. That stops being a quarterly scramble through Jira and becomes a query.
I am wary of overselling this, because "audit-ready" is a phrase that has been ruined by vendors. The concrete version is narrower and more useful: the answer to "what runs, who owns it, what data does it touch, and who approved it" either exists in one place or it does not. In most organisations it does not exist anywhere, and assembling it takes weeks of interviews.
Where this leaves the argument
The framing I want to retire is governance as the tax you pay for speed. On the published evidence it is the other way round. Controls are what make it defensible to grant speed, and each one you build lets you climb a step you could not otherwise climb safely.
Cherny's ladder makes the sequencing explicit, and sequencing is the part organisations get wrong. Skipping the verification loop to reach step 2 faster does not get you to step 2. It gets you an incident, and then a retreat back to step 1 with a new policy that makes step 2 harder to reach than it was before you started.
The one thing the ladder does not tell you is where you actually are. Not where your best engineer is. Where the organisation is, including the people who have never opened a terminal and the apps nobody wrote down. That is next month, and it is the last post in this series before I start writing about the pipeline itself.
dkod rebuilds ungoverned AI-built internal apps under governance, from approved templates. Discovery starts with dkod-signals.
Tell us what you are seeing in your own org, and we will answer.
support@dkod.ai