April 9, 2026

The internal apps nobody wrote down

A Flask app on a laptop that reads the production database and emails a PDF every Monday. No git, no auth, credentials in a plain file. Multiply it by the number of people in your company who got tired of waiting.

Ask a CTO how many internal tools their company runs and you get a number. Ask how that number was produced and it is almost always the same answer: someone counted the repositories in the GitHub organisation.

That count is the tools engineering knows about. It is not the tools that exist.

What the uncounted ones look like

The shape repeats with unnerving consistency. Someone who is not a developer needed a thing, could not wait for it, and asked an agent. What came back works. What came back also looks like this:

  • A single Python file, often Flask or a Streamlit app, sitting in a folder on a laptop or a spare VM.
  • A connection string to a production read replica, and sometimes not a replica.
  • An API token in a .env file next to the script, or worse, pasted directly into the source.
  • No git. There is one copy, and it is the copy that is running.
  • No authentication, because it is "internal", which in practice means anyone who can reach the host.
  • A cron entry nobody else knows about.
  • Twelve people who now depend on the Monday email it sends.

That last point is the one that matters. These tools are not abandoned experiments. They get used. They get used by people who make decisions based on the numbers in them, and nobody has ever checked whether the numbers are right.

Why they exist

It is tempting to frame this as a discipline problem. It is not. It is a queue problem.

The internal-tool request goes into the engineering backlog behind revenue work, which is the correct prioritisation, and it stays there. The person who needed the tool watched engineering get dramatically faster over the past two years and watched their own wait time stay exactly the same. Then they discovered that describing the tool to a chat window produces a working version in an afternoon.

Nobody in that story behaves badly. The analyst solved their own problem with the tools available. Engineering prioritised correctly. The gap between those two correct decisions is where the app lives.

The surveys bear this out, in the sense that the reasons people give are mundane. When asked why they use AI tools their IT department has not approved, the most common answers are that they prefer working without waiting, and that IT does not offer what they need. Malice does not appear on the list.

The risk is not "AI writes bad code"

There is a real quality issue with generated code and it is worth being precise about it.

The Veracode number I quoted in March is about 45% of generated code carrying a security flaw. Stanford researchers found something that bothers me more. Developers working with an AI assistant wrote less secure code than the control group, and were more confident it was secure. The confidence gap is the interesting half, because a scanner catches the flaw and nothing catches the confidence.

But a 45% flaw rate is a solved problem when you have a scanner in CI. Every company on the published list runs one. The flaw rate only becomes a breach when the code never goes near CI, because the code never goes near a repository.

So rank the risks by what actually bites:

Credentials in plain files. A production password in a .env on a laptop is not a code-quality problem. It is a credential you cannot rotate, because you do not know where it is. If that laptop is lost or the folder gets synced to a personal cloud drive, nobody can tell you what was exposed.

No authentication in front of data. An unauthenticated dashboard reading customer records is a data-protection incident waiting for someone to notice. It does not need an attacker. It needs one wrong network rule.

Nobody can patch it. There is one copy and it is the running copy. A dependency CVE lands, and the tool is not in any inventory, so nobody patches it. It just keeps running.

No lineage on the numbers. People make decisions on the Monday email. Nobody has checked the query. This is the risk that never appears on a security review and does the most quiet damage.

It dies with its author. The person leaves. The tool keeps running until it breaks, and then twelve people file a ticket about a system nobody has ever heard of.

Notice that none of these are fixed by a better model. They are all fixed by the app existing in a place with rules.

Policy does not reach these

The standard response is a policy: AI-built tools must be registered, reviewed, and approved before use.

The tools this policy is aimed at were built by people who did not know the policy existed, to solve a problem the sanctioned path could not solve in time. A policy is a request for voluntary disclosure of something that felt like initiative when they did it. Some people will comply. The compliance rate will not be the number you need.

Worse, a strict policy pushes the behaviour further underground. If registering means a three-week review that ends in "no", the rational move is not to register.

The two things that actually change the outcome are unglamorous.

Make the sanctioned path faster than the improvised one. If describing your tool and getting a governed, hosted, authenticated version takes a day, and doing it yourself takes an afternoon plus ongoing maintenance forever, most people pick the governed path. If the sanctioned path takes three weeks, you have written a policy nobody can follow.

Find the existing ones without asking. This is the part that gets skipped. Coding agents leave traces on the machines they run on. Session directories, configuration files, project histories. You can read those traces on a device and produce an inventory of what was built there without anyone filling in a form, and without reading the contents of the work.

That second part is what dkod-signals does. One static binary, pushed via Jamf or Intune or Kandji, runs once at low priority on each device, inventories the AI-built apps it finds, scores each one for risk, and writes a single metrics-only JSON report. It never runs what it finds. Every scan mirrors its metrics-only report into a required S3-compatible bucket in your own cloud account; an additional upload to your organisation's dashboard is opt-in. Correction, 5 September 2026: the earlier claim that nothing uploads by default described a previous scanner version. The bucket mirror is now required on every scan. Counts and risk signals, not source code and not content.

I want to be careful about what that gives you, because it is easy to oversell. A device inventory gives you the shape of the problem: how many tools, on how many machines, touching what classes of data, in which departments. It does not give you a fixed estimate. It tells you where to look.

That is still a change in kind. Every company on the published list of successes governs a known set of things. Getting to a known set is the first move, and most organisations have not made it.

What comes after counting

Counting is only useful if something happens next, and the sorting is where the judgement lives. Some of these apps should simply be switched off, credential rotated, done in an afternoon. Some are real requirements with bad implementations and deserve to be rebuilt properly. Some should not be touched by any automated rebuild, and the right output for those is a named reason rather than a silent skip. I will come back to how that sorting works once there is something to sort.

Next month I want to take the other half of this seriously. The people who caused none of it, who still cannot ship anything, and who are quietly the largest bottleneck in most engineering organisations.


dkod rebuilds ungoverned AI-built internal apps under governance, from approved templates. Discovery starts with dkod-signals.

Questions about this post?

Tell us what you are seeing in your own org, and we will answer.

support@dkod.ai