September 14, 2026

How the harness builds an app

The harness is not one clever agent. It is a few small agents, each with one job, joined by hand-offs the organisation writes down, with a person's signature in front of anything that gets written.

The harness is the part of dkod that takes an app somebody built at their desk with an AI tool and rebuilds it as something the organisation is allowed to run.

This post is about its shape, and the seven decisions behind that shape. No implementation detail. The decisions are the part that matters.

1. The rules belong to the organisation

Every organisation has rules. Which identity provider signs people in. Which checks must pass. Which registry images come from. How a service reaches a cluster. Most of this lives in people's heads.

The harness does not own those rules. The organisation does. They live as plain files in a repository in its own GitHub organisation. One manifest for the policy. Templates that generate applications. Rules written in prose. Examples of code the team likes. The prompts for the agents.

dkod writes that repository once, at the start. After that it only reads it. When the harness wants to change something there, it opens a pull request and a person decides.

A rule in a file can be reviewed, diffed and reverted. A rule inside a vendor's prompt cannot.

2. The floor does not move

Before any agent runs, the harness renders the application's shell from a template. No model is involved. The container, the chart, the pipeline, the health endpoints, the sign-in wiring. Same template and same answers give the same files every time, and every file is checked against the rules by code before a person sees it.

The agents write on top of that floor. Take them away and the harness still produces a correct, boring, governed skeleton. A model having a bad afternoon cannot take the floor with it.

3. Small agents, one job each

The temptation is to build one agent that does everything. We went the other way. Anthropic's guide to building effective agents argues for simple, composable steps over one big autonomous loop. That matched what we saw.

A build runs a few agents. Each one has one role and a short, fixed list of tools.

  • The proposer reads the rules and the discovered app and suggests answers to the questions the template asks. It can read. It cannot write a file.
  • The implementer writes the application inside the rendered shell, one whole file at a time, by the organisation's rules and examples. It can write files. It cannot see the scan report about the risky original.
  • The reviewer reads what the implementer wrote, against the rules, and returns findings. Some findings block.

An organisation can shorten a tool list for a node in its own repository. It cannot lengthen it. An agent that cannot reach a tool cannot misuse it. And a short list is one a security team will actually read.

4. Hand-offs are written down

A build runs top to bottom like this. The organisation declares the hand-offs between the agent nodes in the same manifest that holds the policy, which is what turns the harness from a script into something it can shape.

Loading diagram...

Some hand-offs carry a condition. "Only if the reviewer ran and found nothing blocking." "Only if a person approved." A node that never ran is not a pass.

One rule is fixed. Every path into "deliver" goes through a gate. A manifest with a shortcut around the gate fails the scan, with the line number. The organisation can shape the flow. It cannot remove the signature.

5. The signature belongs to a responsibility

Most tools let "an admin" approve. That is wrong in a specific way. The admin of a dashboard is rarely the person accountable for what ships.

So a gate names a responsibility. Security. Source control. Platform. Only a current holder of that responsibility can approve. The approval is recorded against the exact version of the rules the build used. If the rules change before delivery, the approval stops counting and the build waits. Nothing gets written on a stale signature.

6. Spend is capped where it happens

Each agent node can carry its own budget, in dollars per run and in turns. The organisation sets the ceiling. A node can come in under it, never over. The cap is enforced while the run is happening. When a node hits it, it stops, keeps what it wrote, and says so.

Unglamorous. Also the difference between a pilot and a programme. Finance does not approve an agent system that can surprise them.

7. Every run leaves evidence, and the harness learns by pull request

Each run is stamped: which version of the rules, which node, which model, how many turns, what it cost. Each approval is stamped with who and when. Each review is stamped with what it found. The pull request the harness opens carries the same record in plain text.

"What ran, under which rules, who approved it, what did it cost" has one answer, produced as a side effect of building.

And when a build asks a question whose answer belongs in the policy, "which registry", "which GitOps repository", the harness does not keep the answer to itself. It writes it back into the organisation's manifest as a pull request, on the line the manifest left for it. The second build asks fewer questions than the first.

Why this shape

None of this is new. Small parts with one job. Declared interfaces. A human signature on consequential actions. Budgets. An audit trail. Changes by pull request. Good platform teams have run infrastructure this way for a decade.

What is new is applying it to agents, while most agent tooling goes the other way: one big loop, broad permissions, and "trust me" at the end. Our bet is that organisations will not accept "trust me" for software that touches their data. They will accept a system where the rules are theirs, the flow is written down, a person they chose signs off, the spend is capped, and the evidence is already there when someone asks.

That is the harness. The next posts get specific about what it produces and what it refuses to produce.


dkod rebuilds ungoverned AI-built internal apps under governance, from approved templates. Discovery starts with dkod-signals.

Questions about this post?

Tell us what you are seeing in your own org, and we will answer.

support@dkod.ai