June 11, 2026

What the companies that got this right actually built

StrongDM, Stripe, Shopify, Uber, Wix and McKinsey have all published how their agent platforms work. Read them side by side and the same four components show up every time, under different names.

Enough companies have now published the internals of their agent platforms that you can stop guessing. I went through the primary sources, mostly engineering blogs and a few earnings calls, and lined them up.

The headline numbers are the least useful part. What matters is the machinery underneath, because that is the part you would have to build.

The accounts

StrongDM, Software Factory. Published February 2026. Three engineers, founded as a team in July 2025, under two rules: no human writes code, no human reviews code. The core is Attractor, a non-interactive coding agent whose pipeline is a directed graph of phases defined in Graphviz DOT, built from a natural-language specification of roughly 5,700 lines. Human review is replaced by validation: agents run thousands of end-to-end scenarios per hour against behavioural clones of Okta, Jira, Slack, and Google Workspace. By publication the factory had produced around 32,000 lines of production code. The Attractor spec is open source under Apache 2.0.

Stripe, Minions and Toolshed. Published February 2026. Toolshed is one centralised MCP server with close to 500 curated tools, docs search, tickets, CI status, code search, Slack, Drive, Git, shared by every agent system in the company. Minions are unattended agents running on pre-warmed EC2 devboxes that provision in about ten seconds, isolated from production and the internet. They read a ticket, write code, run it against Stripe's three-million-test suite, and submit a PR. Over 1,300 such PRs merge weekly with zero human-written code, all human-reviewed. Around 8,500 of Stripe's roughly 10,000 employees use LLM tools daily, and a pan-European payment integration that had taken two months shipped in two weeks.

Shopify, Roast. Open-sourced April 2025, days after CEO Tobi Lütke's memo made reflexive AI use a baseline expectation tied to performance reviews. Roast declares workflows in YAML that interleave deterministic steps with LLM steps, persisting every session so it can be replayed. A workflow grades tests; failing grades trigger a second workflow that regenerates them until they pass. All AI traffic routes through one central LLM proxy operated by about six engineers, with per-user cost alerts. Senior-engineer PR review stayed mandatory.

Uber, uReview and friends. DragonCrawl has been in production since October 2023, running mobile end-to-end tests with a deliberately small 110-million-parameter model at over 99% stability across 85 cities. AutoCover is a multi-agent test generator, scaffolder plus generator plus executor plus validator, that raised platform coverage by 10% and saved roughly 21,000 developer hours. uReview is a four-stage prompt-chained review pipeline with confidence filtering and a feedback loop; it now analyses over 90% of Uber's roughly 65,000 weekly diffs, sustains a 75%-plus usefulness rating, and about 65% of its comments get addressed in the same changeset.

Wix, Octocode Orchestrator. Detailed on the Wix Engineering blog in 2026. It monitors user complaints in real time, uses AI to decide whether a behaviour is a real bug or by design, investigates against production data through Trino for databases and Grafana for live logs plus org-wide code intelligence, then has an agent generate the fix and review the PR, all before a traditional ticket would reach R&D. Three layers: a Planner that unpacks intent and queries memory of prior learnings, a Researcher orchestrating specialised sub-agents in parallel with focused MCP toolsets, and a vector-database Memory layer that embeds every completed research cycle. Wix reports bug resolution moving from weeks to hours.

McKinsey QuantumBlack. Published its agentic software-development architecture in April 2026: a deterministic orchestration layer controls the workflow while bounded, specialised agents handle requirements, architecture, and coding. Project context, specs, and accumulated assumptions live in a .sdlc/ folder inside the repository. Humans enter the loop exactly once, at the pull request, reviewing the complete feature rather than each phase. Their diagnosis of the alternative is worth quoting in substance: different developers prompting the same model get different results, quality depends on individual skill rather than process, and the reasoning behind a decision ends up scattered across chat windows where an auditor cannot find it. McKinsey's earlier Seizing the agentic AI advantage (June 2025) describes a large bank modernising around 400 legacy applications with human-supervised agent squads, cutting engineering time by roughly 60%.

Three more for scale rather than mechanism. Sundar Pichai said at Cloud Next in April 2026 that about 75% of new Google code is AI-generated and engineer-approved, up from roughly 25% in late 2024. Meta created a dedicated Applied AI engineering organisation in March 2026 under Maher Saba, with the stated goal of agents doing the bulk of the work to build, test, and ship products while people monitor them. Zeta Global's CEO told the Q1 2026 earnings call that 75% of new code is auto-generated.

The four components

Strip the branding and the same four things appear in every account. Not three, not seven. Four.

1. One hardened starting point

Nobody starts a tool from an empty file. Shopify has YAML workflow definitions. QuantumBlack has the .sdlc/ folder. StrongDM has a 5,700-line specification. Block has recipes.

The reason is not convenience. Agent output quality tracks the quality of the thing it starts from far more than it tracks the model. If logging, secrets handling, access control, telemetry, and tests are already inside the starting point, the agent inherits them and cannot forget them. If they are not, you are relying on a prompt to remember five things, every time, forever.

The compounding property matters more than the quality one. Fix the template once and every tool built from it inherits the fix. Spotify's Fleet Management exists specifically to make that propagation automatic across thousands of repositories.

2. Deterministic checks on every change

This is the component people most want to skip, and it is the one that decides whether the whole thing is safe.

Stripe runs a three-million-test suite. Uber's uReview covers over 90% of 65,000 weekly diffs. Shopify's grading workflow re-runs generation until tests pass. StrongDM replaced human review entirely with thousands of scenario runs an hour against behavioural clones of its dependencies.

Every one of these is deterministic in the sense that matters: the same change gets the same checks, and no human decides per change which checks apply. That property, not the checks themselves, is what makes the volume safe. When 1,300 PRs a week arrive, "we review carefully" stops being a policy and starts being a wish.

3. One controlled path to the models

Shopify routes all AI traffic through a central proxy run by six engineers, with per-user cost alerts. Stripe centralises tools in Toolshed. Uber deliberately runs a 110-million-parameter model where a small one suffices.

Two things fall out of a single path. Cost becomes visible before it becomes a problem, which is the mundane reason. The interesting reason is that you can swap the engine. If model choice is configuration in one place rather than scattered across fifty codebases, a new model version is a config change and a re-run of the eval suite, not a migration project. Given that capability has been roughly doubling every six months, that difference compounds fast.

4. Something that turns failures into system changes

This is the component that separates a platform from a collection of scripts, and it is the one most often missing.

Wix's Memory layer embeds every completed research cycle so the next investigation starts from it. QuantumBlack stores accumulated assumptions in the repository. Uber's uReview has an explicit feedback loop. Shopify persists every session for replay.

The test is simple. When an agent produces something wrong, where does the lesson go? If it goes into a Slack thread and a person's memory, you have not built a system. If it goes into the template, the policy, or the eval suite, then the next run starts from it and the fix reaches everything, including tools built by people who never heard about the incident.

What is missing from all of them

Every account above governs a known set of things. Stripe knows its repositories. Shopify knows its proxy sees all traffic. StrongDM has three markdown files.

None of them documents how you get to a known set, because none of them had to. These are platform teams that built their systems from the inside out, in companies where the code already lived in one place.

That is not the position most enterprises are in. The starting condition is scattered: private agent setups on every developer's laptop, and improvised apps on non-developers' machines that were never in any repository. You cannot put deterministic gates in front of changes you cannot see, and you cannot apply a template to an app that has no home.

So the four components are right, and they are also step two. Step one is finding out what you have.

Next month I want to argue the least popular version of this. The gates are not the price you pay for speed. They are the only reason anyone will let you have it.


dkod rebuilds ungoverned AI-built internal apps under governance, from approved templates. Discovery starts with dkod-signals.

Questions about this post?

Tell us what you are seeing in your own org, and we will answer.

support@dkod.ai