The Implementer Can't Close Its Own Work
What you'll learn
- Why an agent's own “done” is a stand-in for finished work, and how to stop it closing anything
- How to give a judge agent its own checkout so it cannot edit what it judges
- How to write a verdict format that fails closed when nothing could check the work
- How to run a bounded rework loop with any headless coding agent that merges only on an accepted verdict
- How to write findings that make the next round converge, each with the test that would have caught it
Introduction
A coding that finishes a task tells you it's done. The tests are green, the commit is in, the summary is confident. That report is the of finished work: the visible signal we use in place of checking. The is that the work does what you asked for, the way your plan said to do it.
In our own runs, those two came apart on a single day. Two unattended agent runs satisfied every gate our pipeline had and closed as done. One had shipped the approach its plan had explicitly rejected. The other had softened a stated test requirement until it passed. Every gate asked whether the run had reported: a commit, a resolution, a summary. None asked whether anything had checked the work. Both defects were found later, by a person reading the .
The fix we settled on is simple to state. Finished work goes to a review stage, and only a separate judge's verdict can close it. This tutorial builds a small version of that with git, a shell and a coding agent run non-interactively. Examples use Claude Code and opencode; any agent with a non-interactive mode works (Codex CLI, Gemini CLI and aider all have one), and the shape matters more than the tool. By the end you will have a judge agent, a it reviews from, a verdict file format, and a short loop script that relaunches the implementer on findings and merges only on acceptance.
Step 1: Make “done” something only a judge can say
The first rule is a rule about state, not about prompts. Work that finishes moves to review. It does not move to done. The only thing that moves it on is a verdict from something other than the agent that wrote it.
There are two gates here, and you need both:
- The close gate. Nothing merges without an accepting verdict on the exact commit being merged.
- The launch gate. If no reviewer is configured, the implementer doesn't start. A close-side rule on its own leaves the original hole: work that says nothing about verification still runs, and someone eventually closes it on the agent's word. Silence should block the launch, not leak an unverified close.
We had a judge long before we had this rule. It sat in a separate workflow you could route work into, and it had been run eleven times. Both defective runs took the direct path straight around it. A guarantee that depends on launching work the right way is only as strong as your launch discipline, which is the argument of the previous tutorial applied to verification.
Step 2: Write the judge
The judge is a plain markdown prompt file. Create
prompts/judge.md:
Decide whether this work does what the task and its plan asked for.
Exercise it: a plan decision implemented as its rejected alternative
is a defect however green the suite is.
Verify by running things, not by reading the implementer's summary:
- `git status` and `git diff BASE..HEAD` in your checkout
- the test suite
- the feature itself, the way a user would reach it
Write every finding for the implementer who gets the next pass:
name the file and lines, the pattern to copy, and what NOT to redo.
Every finding must name the test that would have caught it, so the
fix and its test ship together.
If you could not check something, list it under "could_not_check".
Never accept on evidence you could not gather.
Your final message is the verdict JSON and nothing else.
The loop passes this file to the agent along with the commit to
review. In Claude Code it can also live in
.claude/agents/judge.md as a , with YAML
frontmatter naming its tools.
The first three lines of that goal are adapted from the reviewer goal we run in practice. Notice what the file does not do: it doesn't ask for a different model. You can point the judge at one, but independence doesn't come from there. The implementer's goal is “make it work”. The judge's goal is “does this match the intent”. The same model, given a different goal and a fresh context, is a real review. In our setup nothing compares the reviewer's model to the implementer's; the close gate is what carries the guarantee, not who sits in the chair. A second model family is a reasonable extra when you want one, and we've measured where that helps, but it isn't the foundation.
The other line that matters is “verify by running
things”. A judge that only reads the diff is checking the same
stand-in the implementer produced. In one of our runs, a judge
allowed to fix things ran the test suite and found three failures
where the implementer had reported two. In another, an implementer
reported four fixes as done; the judge ran git status
and found every one of those edits still unstaged. That was the
fifth false completion claim an executing judge had caught for us.
Step 3: Give the judge its own checkout
Asking the judge not to edit is a request. A tool list that looks read-only is barely more: any agent with a shell can write anything. You want structure: the judge reviews a pinned commit in a tree of its own, the loop throws that tree away afterward, and what gets merged is the pinned commit, by its hash. Whatever the judge does in its tree never reaches the branch.
# the implementer works on its branch in one worktree
git worktree add ../wt-impl feature/csv-export
# after a round, pin what is being judged
SHA=$(git -C ../wt-impl rev-parse HEAD)
# the judge gets a detached worktree at exactly that commit
git worktree add --detach ../wt-judge "$SHA"
# after the verdict, discard anything the judge changed
git worktree remove --force ../wt-judge
Two details here come from things that went wrong for us. First,
review a pinned ref, not a live view of a moving tree.
We once gave a judge a read-only view of the live checkout; during
a long review the checkout's HEAD moved, and the judge ran
the gate against a tree missing the commits under review, and
accepted. Second, merge the hash the judge saw
(git merge --ff-only "$SHA"), never “whatever is
on the branch now”. If the implementer commits again after
the verdict, that commit hasn't been judged.
A worktree still shares the repository's branches, so for a wall a shell can't reach around, run the judge in a that sees only its own checkout; “Run the agent in a box” shows how. Your agent's allow/deny list can add a layer on top. In Claude Code, for example, the judge's call could be:
claude -p "$prompt" --agent judge \
--allowedTools "Read" "Grep" "Glob" "Bash(git status:*)" "Bash(git diff:*)" "Bash(git log:*)" "Bash(npm test:*)" \
--disallowedTools "Edit" "Write"
The separate tree is the part that makes it safe; a deny list only stops most attempts before they start.
Step 4: A verdict format that fails closed
The judge's final message is a JSON object. Keep it small:
{
"verdict": "rework",
"summary": "Export works for ASCII input; the plan's encoding decision is not implemented.",
"checked": [
"git status: clean",
"ran npm test: all passing",
"exported a file with non-ASCII names and opened it"
],
"could_not_check": [],
"findings": [
{
"what": "Writes Latin-1, but the plan decided UTF-8 with a BOM and rejected Latin-1",
"where": "src/export/csv.ts:41-58",
"fix": "Use the writer pattern in src/export/json.ts:12-30; do not touch the column mapping, it is verified",
"covering_test": "test/export/csv.test.ts: 'writes UTF-8 with BOM for non-ASCII names'"
}
]
}
The three verdicts mean different things, and the difference should be mechanical:
accepted: the work does what was asked. The loop may merge the pinned commit.rework: real problems, listed as findings. The work goes back to the implementer.fixed: the judge corrected small problems itself and accepts the result. Only legal for a judge you've deliberately given write access on a branch of its own. Our judge here is report-only and its tree is discarded, so afixedfrom it means something is broken; nothing it changed could reach the branch.
“Fail closed” means: wherever the verdict can't be read,
nothing merges. A crashed judge, an empty reply, prose where JSON
should be, a verdict outside the judge's powers, a rework with no
findings, a finding with no covering test: all of these are
“could not check”, and “could not check”
must never read the same as “checked and fine”. The same
goes for an accepted that lists anything under
could_not_check. That isn't a rejection, but it isn't
an acceptance either; it goes to you.
This isn't paranoia about the implementer. Judges drift from the contract too. On one day in our runs, two different judges reviewing the same piece of work broke the verdict format in two different ways: one invented its own scoring dimensions, the next wrote prose into numeric fields. Both verdicts were refused. The substance underneath was a clean acceptance, and refusing them was still right, because the alternative is a parser that guesses.
Step 5: The loop
Here is the whole thing as one script. It runs from your main checkout, takes a branch and a task file, and gives the implementer at most four rounds.
#!/usr/bin/env bash
# review-loop.sh <branch> <task-file>
set -uo pipefail
# The only line that names a specific agent. Swap in yours.
run_agent() { claude -p "$1" --permission-mode acceptEdits --allowedTools "Bash(git:*)" "Bash(npm test:*)"; } # Claude Code
# run_agent() { opencode run "$1"; } # opencode: any provider, add -m provider/model
BRANCH=$1; TASK=$2
MAX_ROUNDS=${MAX_ROUNDS:-4}
WORK=../wt-impl; JUDGE=../wt-judge; OUT=.review
mkdir -p "$OUT"
# launch gate: no reviewer, no run
[ -s prompts/judge.md ] || { echo "no judge configured; not launching" >&2; exit 1; }
GOAL=$(cat prompts/judge.md)
[ -d "$WORK" ] || git worktree add "$WORK" "$BRANCH"
BASE=$(git merge-base main "$BRANCH")
notes=$(cat "$TASK")
round=1; voided=0; rejudge=0
while [ "$round" -le "$MAX_ROUNDS" ]; do
echo "== round $round/$MAX_ROUNDS ($voided voided)"
# implementer: works and commits on its branch (skipped when re-judging)
if [ "$rejudge" -eq 0 ]; then
(cd "$WORK" && run_agent "$notes
Commit your work on this branch before you finish. Uncommitted edits are not delivered.") \
> "$OUT/impl-$round.txt"
fi
rejudge=0
# pin what is judged; give the judge its own tree at that commit
SHA=$(git -C "$WORK" rev-parse HEAD)
DIRTY=$(git -C "$WORK" status --porcelain)
git worktree remove --force "$JUDGE" 2>/dev/null
git worktree add --detach "$JUDGE" "$SHA"
V="$OUT/verdict-$round.json"
(cd "$JUDGE" && run_agent "$GOAL
Review commit $SHA against base $BASE.
Task: $(cat "$TASK")
Files the implementer left uncommitted (NOT part of this commit): ${DIRTY:-none}") > "$V"
# the judge's tree is disposable: whatever it changed is dropped here
git worktree remove --force "$JUDGE"
# fail closed: anything unreadable voids the round, uncharged
if ! jq -e '(.verdict == "accepted" and (.could_not_check | length) == 0)
or (.verdict == "rework" and (.findings | length) > 0
and all(.findings[]; (.covering_test // "") != ""))' "$V" >/dev/null 2>&1
then
voided=$((voided + 1)); rejudge=1
echo "verdict unreadable, out of powers or incomplete: round voided" >&2
[ "$voided" -lt 3 ] || { echo "judge keeps failing; stopping. See $V" >&2; exit 1; }
continue
fi
if [ "$(jq -r .verdict "$V")" = accepted ]; then
git merge --ff-only "$SHA" && echo "accepted and merged $SHA"
exit $?
fi
# rework: the judge's findings ARE the next round's instructions
notes="Address these findings. They are the reason this work is not merged.
Fix each one and add the covering test it names. Do not redo verified work.
$(jq -r '.findings[] | "- \(.what) [\(.where)]\n fix: \(.fix)\n covering test: \(.covering_test)"' "$V")"
round=$((round + 1))
done
echo "round budget spent; over to you. Last verdict: $V" >&2
exit 1
The verdict check reads the agent's standard output, so it expects
only the final message there; claude -p prints exactly
that. If your agent prints more, the round is voided rather than
misread. A few other things in there are deliberate.
Only real verdicts spend rounds. When the judge
crashes or returns something malformed, the machinery broke, not
the work. That round is voided and the counter doesn't move, so the
next pass re-runs the judge on the same commit rather than asking
the implementer for another attempt. We learned this one by
spending it: before we voided machinery failures, a hiccup used up
a round, and two of our live cycles spent their whole budget on a
single real verdict. The separate voided cap keeps a
judge that fails every time from looping forever.
The budget is a safety net, not a sizing tool. Our default is two rounds, calibrated for frontier implementers; for cheaper local implementers we run four, because a weaker implementer paired with a strong judge converges by iteration. That holds only while the findings stay within the implementer's reach. One of our runs used four rounds out of four and diverged: the first two built the core honestly, the third scored badly, and the fourth spent eight minutes, committed nothing, and reported four fixes done with the edits unstaged. The plan behind it carried nine architectural decisions, which was too much for the smaller local model we'd given it. The budget turned that into a bounded, evidenced stop for a person rather than an endless burn. Size the plan to the implementer, or the implementer to the plan.
The findings go back raw. Nobody rewrites them between rounds. If the judge writes them well, they are already the instructions, which brings us to the part that decides whether the loop converges at all.
Step 6: Findings that converge
A findings block that only condemns cannot converge. “The export handling is wrong” tells the next implementer that it failed, which it already suspects. It doesn't say what right looks like. The verdict that set our standard named the exact working pattern to copy (file, lines, mechanism) for the one decision that was missing. Findings like that usually need one more round.
So the judge writes for the next implementer, not for you. Each finding says where, what, the pattern to follow, and what not to touch because it's already verified. That last part matters more than it looks: without it, a retry tends to “improve” things that were fine and introduce new findings.
The second half is the covering-test ratchet. Every finding names the test that would have caught it, and the retry is told to write that test along with the fix. A defect caught once then stays caught mechanically, by the suite, without relying on any future judge noticing it again. Once we made this a standing part of the reviewer's goal, every findings-bearing verdict in the cycle we watched named its covering tests, and one round's crash regression arrived with its newly failing tests attached. The loop above enforces it: a rework verdict with a finding that lacks a covering test is treated as incomplete.
When it works, it looks like this. In the run that first carried our loop to an automatic close, the judge's findings went from six, to one, to one, to one, and then the verdict was accepted and the work merged without anyone touching it.
What the judge still can't see
An executing judge verifies what it runs. Code that nothing runs passes every round. When we hand-finished the diverging run above, we found that its earlier, judge-passed commits carried two latent crashes: one imported a name from a module that doesn't export it, the other imported a module that doesn't exist. Four judge rounds never hit either, because no test reached those paths.
That's an argument for the ratchet rather than against the judge. Each round where the judge asks “which test would have caught this?” grows the set of paths something actually runs. It's also a reason to read a diverging run's self-reports, including any list of “remaining work” it files, with the same suspicion as its green claims. In that run, two of the four follow-up items it filed claimed work was missing that had in fact been committed.
Then there's cost. With a judging, our review rounds took about 17 to 21 minutes and cost about $2 to $2.70 each. Once the test suite was fast, the judge's own diff-reading and verification became the floor. Judging every round is the expensive part of the loop, and it's still cheaper than finding the rejected alternative in production. A cheap mechanical check before the judge (does the named failing test now pass?) can save judge rounds on obviously unfinished work; we haven't yet measured whether thinning the intermediate judgments pays, so we don't recommend it yet.
Try it yourself
Set up the loop on a small repository you don't mind breaking.
- Add
prompts/judge.mdandreview-loop.shfrom this page and setrun_agentto your agent. With Claude Code, change the test command in its allow list to your project's. - Write a task file with one decision in it that has a plausible wrong answer, and state which alternative you rejected. For example: “Store timestamps as UTC ISO strings, not epoch integers.”
- Before running the loop, test the judge. On a branch, commit the rejected alternative yourself, with passing tests. Run only the judge half against that commit. If it accepts, your judge is checking the suite, not the intent; tighten its goal until it catches the planted defect.
- Now run the loop for real. Read each
verdict-N.json. Does every finding tell the next implementer what to do, or only what went wrong? - Break the machinery on purpose: make the judge print a line of prose before its JSON. Confirm the round is voided, the round counter doesn't move, and nothing merges.
Delete prompts/judge.md and run the script
once more. It should refuse to launch the implementer at all.
An implementer's green self-report is the extension of finished work. The verdict of something with a different goal, run on a tree it can't touch, is the closest a loop gets to the intension.
For the wider pipeline this loop sits inside, from plan to land, see How We Build With Agents Now.
Key Takeaways
- Finished work goes to review; only a separate judge's verdict closes it. No reviewer configured means the work doesn't launch.
- The judge reviews a pinned commit in its own worktree, and you merge that hash; anything it changes there is discarded. It can't change what it judges, by structure rather than by request.
- Independence comes from a different goal and a fresh context, not from a different model.
- Fail closed: a crash, a malformed verdict, or an acceptance with unchecked items never merges. Machinery failures void the round instead of spending it.
- Findings are written for the next implementer, and each names the test that would have caught it. Findings that only condemn don't converge.
- An executing judge verifies what it runs. Code no test reaches passes every round, which is why the ratchet matters.
Next Steps
- Reviews That Can't Quietly Drop Findings — the next tutorial: making sure every finding a review raises is answered, not lost between rounds
- Plans That Record Decisions, Not Guesses — a judge can only catch a rejected alternative if the plan wrote it down
- Multi-Agent Systems Drift. Here's Where to Put the Walls. — why isolation between agents beats asking them to stay independent
- Is a Debate Better Than One Strong Model? — what we measured about a second model's view
- Concepts — interactive visualizations of blind spots in multi-agent systems