The argument on this site's front page is that a single reviewing its own work walks the same path twice and keeps its blind spots, while isolated agents on different paths cover each other and a person arbitrates what's left. When we first wrote it down, it was a thesis. It's now how we build software day to day. This post describes the working version, as concepts, including what it costs.

The short version: agents are disposable; context is not. Everything else follows from taking that seriously.

The problem with one long session

The natural way to use a coding agent is one long conversation: explain the project, ask for the feature, correct it, ask for the next. It works, until it doesn't. The context fills with half-relevant history. Early assumptions harden into facts. The agent optimizes for what it can see, which by hour three is mostly its own earlier output. It marks its own homework and gives itself good grades.

That's the failure this site keeps describing, at the scale of a whole project. The fix isn't a better session. It's making sure no single session carries anything it would be costly to lose.

The pipeline A work item flows through prep, debate, sign-off, implement, checks, review, and land. Each stage is a short-lived agent in its own sandbox, except sign-off, which is a person. Beneath the stages runs the durable context they read from and write to: the item's description and journal, the plan and its decisions, and runbooks and decision records. item FILED prep QUESTIONS debate 2 MODELS sign off build OWN CLONE checks GATES review OTHER MODEL land findings: another round, up to a budget DURABLE CONTEXT item + journal plan + decisions runbooks decision records
Every box above the line is short-lived. Everything worth keeping is written to the line.

1. Ephemeral agents

Every run starts clean: its own , its own working copy of the code, a fresh home directory with no credentials in it. When the run ends, all of that is thrown away. Nothing an agent “remembers” survives, and that's the point. An agent that can't carry state can't carry drift, stale assumptions, or a private theory of what the task is really about.

Isolation is also what makes the rest safe. An agent in its own clone can be as wrong as it likes without touching anyone else's work, and its output reaches the shared code only through the checks below.

2. Durable context

If agents forget, something else must remember. Each piece of work lives as an item with a description, an append-only journal of what happened to it, and (for anything non-trivial) a plan. When an agent is launched, its instructions are assembled from those, plus standing rules for the project. The agent doesn't need to remember the task; it's told the task, every time, from the record.

Runs write back too. An agent finishing a job declares what it left undone or discovered along the way, and those declarations become new items instead of evaporating with the session. Beyond the items, operating procedures live as runbooks marked unproven until someone has run them exactly as written, and architectural decisions live as records that bind until superseded.

3. Planning agents

Before anyone writes code, a planning agent does a prep pass on the item. Its job isn't to decide. It's to find the decisions: read the code the item touches and write the plan's open questions, each with the options and what hangs on them. A plan whose questions are all written down is a plan whose disagreements are visible, and that's the input the next stage needs.

4. Debating agents

The open questions are then argued, not accepted. In a debate walk, two seats on different models take each question: one presents it, the other argues, and together they reach a decision that's recorded in the plan with its reasons. Different models matter. Two copies of one model share their blind spots; the arithmetic on our concepts page only holds when the paths really are different.

Not every question should be debated. A question of taste (what a page is called, how a feature should feel) has no right answer for two models to converge on. Those are walked with a person.

5. A person signs off

A decided plan waits for sign-off. A person reads the decisions, overrules any they disagree with, and only then does the item become buildable. It's a small act and it's deliberately not automated: it's the point where someone who can be asked “why?” accepts responsibility for the answer.

The same principle covers anything that can't be undone. Agents can propose deleting work, canceling it or rewriting history; they can't do it without an explicit confirmation from a person.

6. The implementer never reviews itself

The build runs in its own clone, against the signed-off plan. Before anyone judges it, mechanical gates run: the project's tests, and a check that the commit messages describe the change and that the docs were considered. Then a reviewer (a different agent, on a different model, starting from a clean context) rules on the . It can accept the work, fix small things, or send it back with findings for another round.

Rounds have a budget. A change that fails review twice doesn't get a third, fourth and fifth attempt from an agent determined to make the reviewer happy; it stops and comes back to a person with the reviewer's findings attached. A loop that isn't converging is information, not something to retry.

7. Measure, and be honest about whose failure it was

Every attempt is recorded: which model, which runtime, what the gates said, what the reviewer said. That record is what later decides which model gets which kind of work. One lesson came early and was humbling: when a run fails, check whether the agent failed or our machinery did. In our first pilot, every failure was ours: a test that assumed a file only present on one machine, a gate without the project's environment. Scored naively, the record would have taught us to distrust a model for our own bugs. So failures caused by the are classified separately and kept out of the model's record.

What it costs

This is slower per item than a good engineer with one agent in one session. It spends more : a prep pass, a debate, a build and a review, where one session would have done all four at once. It needs a person paying attention at sign-off and when a loop stalls. And it has a failure mode of its own: the process can become the product, and you can spend a week improving the pipeline instead of the thing it builds.

What it buys is work we can trust without re-reading all of it. Each stage has one job, a clean context and someone else checking it. When something goes wrong, the record says where. And none of it depends on any agent being brilliant, or remembering anything, which is fortunate, because none of them do.

The thesis, in practice

Isolated contexts, different paths, no self-review, durable intent, and a person as arbiter. Each of those started as a line on a diagram. Each is now a stage in a pipeline we use every day. The deep dives are in other posts: why isolation has to be structural, why agents write code for nobody when nothing asks who it's for, and why reasoning is worth keeping after the session ends.