There is no playbook for running AI agents in a company
Three engineers at StrongDM ship production security software without reading the code. Somewhere else in your org, a marketing analyst is pasting a database password into a chat window. Same year, same technology, no shared rulebook.
Two things happened in February 2026, three weeks apart.
StrongDM's AI team published a manifesto describing what they call a Software Factory. Three engineers. Two rules: no human writes code, and no human reviews code. Humans write specifications, curate test scenarios, and read scores. The agents do the rest. Their AI Context Store is roughly 32,000 lines of Rust, Go, and TypeScript that nobody on the team has read line by line, and it ships as part of a security product.
Around the same time, Stripe documented Minions, a fleet of unattended coding agents that merge over 1,300 pull requests a week with zero human-written code in them. Every one of those PRs is still reviewed by a person. The agents share a single internal MCP server, Toolshed, holding close to 500 curated tools, and they run on pre-warmed EC2 devboxes that spin up in about ten seconds and cannot reach production or the open internet.
Read those two accounts back to back and it is easy to conclude the industry has figured this out.
It has not. What those companies have is a system they built themselves, from scratch, over a year or more, with dedicated platform engineers. What most enterprises have is a few hundred people improvising.
The two failure shapes
I keep seeing the same split inside organisations, and the two halves fail in opposite directions.
The developers move fast with nothing holding them. They have Claude Code or Cursor or Codex on their laptop. They have their own CLAUDE.md, their own hooks, their own prompt library. It works, genuinely well, and it works differently for every single one of them. Nothing they learned is written down anywhere the next person can find it. When the developer leaves, the setup leaves.
The quality problem here is not that agents write bad code. It is that nobody is checking in a consistent way. Veracode tested more than 100 models on 80 coding tasks for its 2025 GenAI Code Security Report and found security flaws in about 45% of the generated code, with Java above 70%. Security pass rates stayed flat through 2025 even as the models got better at the functional part, because nothing in the training objective rewards a safe answer over a working one. That is a fixable problem. You fix it with a gate. Most teams do not have one, because gates are somebody's job and nobody owns it.
Everybody else moves at zero. Product, BI, data, finance, support. These teams do not write code and never will. Their tool requests sit in a queue behind roadmap work, and the queue is measured in quarters. They have watched the engineers get ten times faster and their own wait has not moved at all. So a few of them start building anyway, in a chat window, and now there is a Python script on a laptop that reads production data and emails a report to twelve people. No git. No auth. Credentials in a plain text file.
Both halves are behaving rationally. Neither is doing anything wrong. There is simply no rulebook.
Why nobody can hand you the rulebook
Every published account of an organisation doing this well describes a bespoke internal platform with a name.
Shopify open-sourced Roast in April 2025, workflows declared in YAML that interleave deterministic steps with LLM steps, with every session persisted so it can be replayed. A workflow grades tests, and failing grades trigger a second workflow that regenerates them until they pass. All AI traffic routes through one central LLM proxy run by a small team, with per-user cost alerts. Senior-engineer review of PRs stayed mandatory.
Spotify has run Fleet Management since 2021, executing code changes across thousands of repositories and auto-merging when tests pass. In February 2025 they replaced the hand-written transformation scripts with Honk, a background agent that takes migration instructions as plain English, with an LLM judging the resulting diff. By November 2025 Honk had merged over 1,500 AI-generated PRs. The surrounding machinery, targeting and PRs and review and automerge, did not change at all.
Airbnb migrated about 3,500 test files from Enzyme to React Testing Library in six weeks against a 1.5-year manual estimate. The mechanism is not a smarter prompt. It is a state machine of validate-then-fix steps with retry loops of 50 to 100 error-enriched re-prompts. 97% automated.
Block's Goose went the other way and pushed agents to non-engineers. Roughly two-thirds of Block's 10,000-plus employees use it weekly, across sales, design, and support, according to Forbes in November 2025. The mechanism that made that possible is the recipe: a shareable YAML file bundling instructions, required extensions, and parameters. One engineer's working setup becomes something a salesperson can run.
Look past the names and the same four parts show up every time.
- A hardened starting point, so quality does not depend on who or what wrote the code.
- Deterministic checks that every change passes, the same way, every time.
- One controlled path to the models, with logging and spend caps.
- Something that feeds failures back into the starting point instead of into somebody's memory.
None of these companies sells you that. They published how they built it, which is generous, and then you still have to build it.
The part nobody publishes
Here is what I find most striking about the published accounts. Every one of them starts from a known inventory. StrongDM has three markdown files. Stripe has one Toolshed. Shopify has one proxy. They can say "every change passes these gates" because they know what every change is.
Nobody publishes the step before that, which is finding out what already exists. And that is the step almost every enterprise is actually stuck on. You cannot put a gate in front of a tool you do not know about. You cannot apply a template to an app that lives in someone's home directory. The dashboard that runs on a laptop under a desk is not in any registry, and no amount of policy fixes it, because policy only reaches things you can name.
That is the gap this blog is going to be about. Not whether agents can write code. That question is settled, loudly, by companies shipping security products and payment integrations with them. The open question is how an organisation of a few hundred people gets from scattered private setups to one path that everyone uses, without stopping anybody's work while it happens.
Next month I want to look at the ungoverned half of that: the internal apps nobody wrote down, why they exist, and what they actually look like when you go and count them.
dkod rebuilds ungoverned AI-built internal apps under governance, from approved templates. Discovery starts with dkod-signals.
Tell us what you are seeing in your own org, and we will answer.
support@dkod.ai