August 12, 2026

You cannot govern what you cannot see

Every governance plan starts at step two. Registry, templates, gates, all of it assumes you know what exists. Here is what we built to answer the step-one question, and what it deliberately refuses to do.

Five posts into this series, the pattern in the published accounts is clear enough to state plainly. Companies that run agents at scale have a hardened starting point, deterministic checks on every change, one controlled path to the models, and a loop that turns failures into system changes.

Every one of those assumes you know what you are governing.

Stripe can say every change passes its test suite because every change is in a repository it controls. Shopify can say all model traffic goes through one proxy because the proxy is the only route out. These are true statements about a known set of things.

Ask an enterprise of a few hundred people what it is governing and you get a repository count. That is the tools engineering knows about, and it is not the tools that exist. The rest are on laptops: a Flask app reading a production replica, a scheduled script that emails a PDF to twelve people, a Streamlit dashboard behind no authentication at all. Nobody wrote them down because nobody asked, and nobody asked because there was no way to ask that produced a real answer.

That is the step-one problem. It is unglamorous and nobody publishes about it, and it blocks everything else.

Why asking does not work

The obvious approach is a survey. Send a form, ask people to register their AI-built tools, follow up with the teams that do not respond.

I have watched this fail in a specific way. The people you most need to hear from are the ones who built something under time pressure using a tool they were not sure was approved. A registration form asks them to volunteer that. Some will. The response rate will not be the number you can plan against, and the tools you miss are systematically the riskiest ones, because the correlation between "I would rather not mention this" and "this touches data it should not" is not zero.

The second approach is network scanning, looking for unexpected services. That finds hosted things. It does not find the script running from a cron entry on a laptop, and it does not tell you what built it.

What we built instead

Coding agents leave traces on the machines they run on. Claude Code, Codex, Cursor and the rest keep session directories, project configuration, and history. Those traces describe what was built on that device, and you can read them without reading the work itself.

dkod-signals is one static Rust binary. IT pushes it to every macOS, Windows, and Linux device through the MDM they already run, Jamf or Intune or Kandji. It executes once at low priority, walks the agent traces and the projects they point at, scores each app it finds for risk, and writes one JSON report.

What the report contains is counts and signals. How many AI-built apps on this device. Which agents produced them. Which ones look like they hold credentials in plain files. Which ones appear to reach a database. Which ones have no version control. Which have been touched recently and which are dead. Nothing else.

What it deliberately does not do, and this is the part I care most about being exact on:

Dashboard upload is opt-in. Every scan produces a metrics-only report and mirrors the same bytes into a required S3-compatible bucket in your own cloud account, using credentials you configure. The scan refuses to start without the bucket and its credentials. With the upload flag on, an additional copy goes to your organisation's dashboard. See the raw-report mirror documentation for the current setup.

Correction, 5 September 2026: the original no-upload statement and the 26 August update described earlier scanner behavior. The current scanner requires the bucket mirror on every scan. Dashboard upload remains opt-in; the earlier no-network and identical air-gapped-operation claims no longer apply.

No telemetry, licence checks or version pings. The only network destinations are your required report bucket and, when enabled, the dashboard you configure. The scanner must be able to reach the bucket.

It never executes what it finds. It reads file metadata and structure. It does not run the app, import it, or evaluate any of its code. An inventory tool that executes untrusted code found on employee laptops would be a worse problem than the one it is solving.

It does not collect content. Not source code, not prompts, not the text of agent conversations, not customer data. The report is metrics. That is a design constraint, not a setting, and it is the only version of this I would be willing to push to a whole fleet of employee machines.

The v1 binary is complete and available now. The aggregation layer that rolls per-device reports into an org-wide view is a separate piece, and since this post first ran it has shipped: the managed dashboard at dashboard.dkod.app, with self-hosting available on request. When it was not shipped, this paragraph said so — I would rather date a correction than imply a dashboard existed before it did.

What the report is good for, and what it is not

Being honest about the limits of a device inventory matters, because this is exactly the kind of tool that gets oversold.

It gives you the shape of the problem. How many AI-built apps exist, on how many machines, in which departments, carrying which risk signals. For most organisations that is the first real number they have ever had, and it usually lands somewhere between "more than we thought" and "considerably more than we thought".

It does not give you certainty. An app whose traces were cleaned up, or one built entirely in a web chat window and pasted into place, leaves less to find. Risk scores are signals derived from structure, not verdicts. Something flagged as holding a credential needs a person to look.

So the output is a place to look, ranked. Not a compliance certificate. Treated as the second thing it is useful. Treated as the first thing it is misleading, and I would rather lose a sale than have someone quote a scan count as coverage.

Then what

An inventory is only worth producing if something happens next, and the honest answer is that the tools split three ways.

Some should be switched off. Plenty of what a scan finds was built for one meeting and never opened again, and it is still holding a live credential. There is nothing to migrate. Turn it off, rotate the credential, and that risk is gone the same day. This bucket is usually larger than anyone expects and it is the cheapest win in the exercise.

Some should be rebuilt. A scheduled report twelve people depend on is a real requirement with a bad implementation. It wants to be the same report, running from a template that already contains logging, secrets wired to the real secret store, access control behind the company identity provider, telemetry, and tests. Not rewritten by hand. Rebuilt from an approved starting point, through the same checks as everything else, deployed into the customer's own cloud.

That is the governed build pipeline, and it is in early access. The first slice is deliberately narrow: a scheduled-digest template, a read-only dashboard template, and an internal-web-app template that is internal-only with no payment handling. Three shapes, chosen because they cover most of what the improvised apps actually are.

Some cannot be rebuilt safely, and get flagged with the reason. An app that writes to a production database is outside the envelope. So is anything handling payments. The right output there is a named blocker, "writes to production DB", attached to the app, so a human decides. An automated pipeline that quietly attempts the hard cases is how you cause the incident that ends the programme.

Having an edge to the envelope, and saying where it is, is the difference between a pipeline and a promise.

The part I am least sure about

The narrow template set is the bet, and it is a real bet.

If most improvised internal apps are genuinely digests, dashboards, and small internal web apps, three templates cover the majority and the approach works. If the distribution is much longer-tailed than that, the flagged-and-blocked pile stays big enough that the pipeline does not change anyone's life, and the templates need to multiply faster than we can harden them.

The device inventory is what settles that argument, which is the other reason discovery comes first. It tells you what the tools actually are before anyone commits to what the pipeline should build.

If you want the number for your own organisation, dkod-signals is available today, and the pipeline is in early access.


dkod rebuilds ungoverned AI-built internal apps under governance, from approved templates.

Questions about this post?

Tell us what you are seeing in your own org, and we will answer.

support@dkod.ai