Solution · Autonomous software delivery

The Dark Software Factory

Specs in, pull requests out, people only at the gates. This is what makes step 4 affordable: one developer can prototype with a client in weeks and keep the result running, and a run that did not do what it claimed cannot report success.

Live orchestrator v0.2.97, read from the running containers on 4 September 2026. Four production projects, one Compose stack.

What goes dark, and what does not

A software factory with the lights off changes who does what before it changes the software. When machines write, review, test, deploy and document the code, the person who used to do those jobs moves to the gates, and the gates become the job. So the question to ask of any ‘autonomous’ pipeline is where its gates sit, what a gate can see when it opens, and whether the system tells the truth about what happened on the far side. A pipeline that cannot is a dark room with an optimistic status light on the door.

This one hands a spec, a markdown document describing one unit of work, to an orchestrator built on DBOS and Postgres. It dispatches Claude Code agents with one role each: plan, write, review, audit security, test, deploy and document, and an eighth, off the main line, that turns a plan into specs. Every step is checkpointed, so a restart resumes at the last completed step rather than paying for the same model call twice. Four projects run on it. One is the quoting engine at eBatt.ai.

The hard part was never calling a model seven times. It was making the sequence durable through restarts, redeploys and expired credentials, and making it honest, which is harder: each failure in the evidence below looked like success at first, and each now has a named check in front of it.

I first wrote it as terminal panes and a polling loop, and it lost state every time the laptop slept. The rebuild did not remove the person; it gave the person a shorter list of better questions. The automation is allowed to be fast. It is not allowed to be sure.

The pipeline, with the gates drawn as gates

Eight agents, two lanes and three locks. The security stage cannot be named as a successor by any agent, so nothing can reach it and nothing can skip it. The deployer’s claim of success is checked against the remote before the docs agent is allowed to run. Click any node for what it does and what it is not allowed to do.

Loading diagram...

One attempt, told honestly

The beats below are lifted from the orchestrator’s own dispatch log for 28 July 2026, with the project anonymised. The first twenty seconds are what a good run looks like. The last seven are what a false signal looked like on the same morning, and why the release that followed added an attempt number to every idempotency key.

Live demo

Lights off. Gates on.

One real attempt from the orchestrator's own log: eight agents, one human gate, one check the agent cannot fake, and the honest coda where a cancelled worker kept running. Timestamps are real; the project is anonymised.

  • DBOS
  • Postgres
  • Claude Code
  • Docker Compose
  • GitHub
  • Cloudflare
  • ntfy
  • Honcho
Dark Software Factory — specs in, pull requests out · loops every 34 s

agent.trace

Live tool calls

streaming
0.0s / 34s

The evidence, including the failures

424 of 533
merged eBatt.ai pull requests from factory branches, since 26 April 2026
141
merged pull requests on the orchestrator itself, the last on 3 September
3,011
structured handoffs across 255 specifications in the handoff store
444
of those handoffs addressed to a human rather than the next agent
13
named halt reasons, derived from the code, not from a grep
15 to 25 min
wall-clock for a full run of the pipeline

The numbers above are counts I can point at, not estimates, and none of them is a promise about the next run. The orchestrator carries 39 documented API operations, nine scheduled monitors and a test suite that stood at 1,768 passing on 30 August 2026. Of the handoffs in the store, 2,522 ended in success, 54 failed, 63 were blocked and 72 were partial. Those last three numbers are the ones a brochure would leave out.

The failures are the part worth reading. The workflow engine’s status row says SUCCESS for a run that halted, because a halt is a normal return, so only the on-disk marker counts; four green rehearsal runs in a row reported writing a file that never existed, which is why rehearsals now prove the shape of a prompt and nothing about an agent. A deployer wrote ‘success’ while the code never reached the remote, so the orchestrator now reads the remote’s tree itself. On 28 July an attempt died in 86 seconds without launching a model, because an idempotency key carried no attempt number, and on 30 July a telemetry lane went silent for 53.67 hours without the dead-man switch noticing; the diagnosis records that the gap ‘would not page today’. Each now has a named check.

What still needs a person

The list is written down in the operating documents rather than implied by a diagram, and it is short enough to print. The watchdog that detects a stalled dispatch alerts but does not restart anything, after a false positive once cancelled a healthy pipeline and threw away six paid-for dispatches. Stranded workflows are detected and reported, not forked back to life. And there is no migration tool: tables create themselves with idempotent statements at boot, which works and cannot express a destructive change, so the first destructive change will need a real tool and a person holding it.

Deploy approval

The one place a person is required by default. A durable wait of up to twenty-four hours that survives an orchestrator restart, answered with a one-tap notification that carries a low-privilege token rather than the API key. A destructive database migration forces the gate on regardless of any appetite setting.

Reading the spec

In programme work an agent authors each specification from the plan and the facts harvested from specs already done. Nothing advances until a person has read the staged spec and approved it; a gate pending for a week costs nothing and cannot be stranded by a redeploy.

Plan revision

Every change to a programme's plan is operator-initiated, and there is no way to auto-apply one by construction. The gate is hard on.

Persistent disagreement

A reviewer may send work back to the builder twice. The third disagreement halts the run for a human read, not a fourth attempt.

Want a factory of your own?

The runtime is Apache 2.0 and self-hostable. If you want one for your own team, I can help you decide where its gates sit and set it up, starting from the failure record above.