Run a Debate Before You Decide

Multi-Agent Systems Advanced ~55 min

What you'll learn

Introduction

A plan for a change usually ends with open questions. Which layer owns the fix? Fail loudly or degrade? One read path or two? Most of the time the that wrote the plan answers them itself, in the same pass, from the same angle it used to write the plan. The answers look reasonable, which is exactly the problem: you have no second view to compare them against.

This tutorial builds a small debate out of two ordinary agent sessions and a shell script. One agent, the walker, reads the plan and the code and presents one question at a time. The other, the answerer, reads only the code and the walker's presentations, never the plan, and decides. Every decision is a JSON verdict that has to say what it turned down and why. A third short call, the judge, checks that those reasons are real reasons.

Examples use Claude Code and opencode; any coding agent with a non-interactive mode and file-writing tools can fill a seat, and the seats don't have to share an agent, a model or a provider. For one measured case of what this structure finds that a single strong model misses, read Is a Debate Better Than One Strong Model? This page is the recipe behind it.

The link to this site's theme is direct. A decision recorded as “we chose (b)” is an : it looks decided. The is the reasoning, and in particular what you gave up. A verdict that can't say what lost hasn't actually been weighed.

Step 1: Pre-register your own answers

Before any debate runs

Start with a plan that has open questions in it. Each question should be specific and grounded in the code: the conflict, where in the code it lives, and what is unverified. If you followed Plans That Record Decisions, you already have one.

Now, before you launch anything, answer every question yourself (or have the agent you normally work with answer them) and write the answers to a file the debating seats cannot read. One line per question: the option you'd pick and the reason.

# private/solo.md  (written before the first walk, never edited)
Q1: (a) add a second lean query beside the full read; smallest change
Q2: (a) anonymous memory is the counter carrying the growth
Q3: (b) measure call rates after the deploy
sha256sum private/solo.md > private/solo.sha256

This is the step people skip, and it is the one that makes the rest worth measuring. After you have read a good debate, every answer it reached looks obvious, and you will remember having thought of it. A file written beforehand, with its hash taken, can't be revised by hindsight. When the debate is over you compare question by question: matched, improved, or worse.

We added this step after the first few debates, because without it we couldn't answer the only question that matters about the whole exercise: is it better than just answering?

Step 2: Lay out the seats

Plan-blindness is a directory layout, not a request

The answerer must not see the plan. If it can read the plan, it reads the author's framing, the author's leaning and the author's blind spots, and you are back to one perspective. Telling it “don't open plan.md” is a request. Give each seat its own workspace and keep the plan out of the answerer's entirely (Put the Walls in Structure covers why).

REPO=/path/to/your/repo
mkdir -p debate-run/{private,prompts,walker/out,answerer/out,judge/out}
# run_agent (agent.sh, below), walk.sh and check_verdict.py go in debate-run/
cd debate-run

# each debating seat gets its own detached checkout of the code
git -C "$REPO" worktree add --detach "$PWD/walker/code"   HEAD
git -C "$REPO" worktree add --detach "$PWD/answerer/code" HEAD

# the plan lives in private/ and is copied only to the walker
cp "$REPO/plans/cache-fix.md" private/plan.md
cp private/plan.md walker/plan.md

If your plans live inside the repository, the answerer's checkout contains them too. Remove them from that one with a sparse checkout, so a grep can't find them:

git -C answerer/code sparse-checkout set --no-cone '/*' '!/plans/'

Each seat runs from its own directory, and the answerer's holds only its checkout and the thread: no private/, no walker/, nothing a search from there can stumble on. But ../private is still one path away. To make that path not exist, run the answerer and the judge in a box that mounts only their own directory (the “Run the agent in a box” step of Put the Walls in Structure shows how). Inside the box the plan isn't denied, it isn't there, whichever agent sits in the seat. History is covered too: the worktree's .git link points into a repository outside the mount, so even a shell can't git log its way to the plan.

Choose an agent for each seat

The walk script calls one function per seat, so every seat can be a different agent on a different model:

# agent.sh: the only place that names a specific agent. Swap in yours.
run_agent() {  # run_agent SEAT PROMPT, from inside the seat's directory
  case $1 in
    walker) claude -p "$2" --permission-mode acceptEdits --allowedTools Bash ;;
    *)      claude -p "$2" --permission-mode acceptEdits ;;
  esac
}
# A seat on another provider's model, through opencode:
#   answerer) opencode run -m provider/model "$2" ;;
# Wrap the answerer's and judge's commands in the box from the walls tutorial.

We put different models in the walker's and the answerer's seats on purpose. Two seats on one model share its blind spots, so the answerer's challenge comes from the same angle as the walker's presentation. A against a local one, or two providers' models, gives you the second view the debate exists for, and no single provider's outage or policy change stops it.

If a seat runs Claude Code, you can add its deny list as a second layer, or as a weaker substitute where you can't box a seat. Project settings in the seat's own directory do this. A rule path starting with // is absolute, and $PWD already starts with one slash, hence /$PWD:

mkdir -p answerer/.claude judge/.claude
cat > answerer/.claude/settings.json <<EOF
{
  "permissions": {
    "deny": [
      "Bash", "WebFetch", "WebSearch",
      "Read(/$PWD/private/**)",
      "Read(/$PWD/walker/**)"
    ]
  }
}
EOF
cp answerer/.claude/settings.json judge/.claude/settings.json

The walker does get a shell, in its own throwaway checkout, because part of a walker's job is measuring things: a timing, a log count, whether the function the plan cites still exists. The script below stops the walk if either checkout comes back modified.

The asymmetry is the design. The walker sees everything and decides nothing. The answerer decides everything and sees only what the walker chose to show it, plus the code it can check that against. If a presentation is too thin to decide from, that's a defect in the presentation, and the answerer says so.

Step 3: Write the three briefs

The walker

# prompts/walker.md
You are the walker in a debate. You can read plan.md (the plan and its open
questions) and code/ (a checkout you must not modify). The answerer cannot
see the plan: everything it knows about a question comes from what you write.

For the one question you are given:
- Verify every claim against code/ before you make it. If the plan is stale
  (the defect is already fixed, the function moved), lead with that.
- State the tension: what is in conflict, and why it matters.
- Give lettered options, each with what it costs. Name the cost of your own
  leaning too.
- State your leaning and why. You do not decide.
Write the presentation to out/<question>.md, and the options to
out/<question>-options.json as {"a": "...", "b": "..."}, each text exactly
as in your presentation.

If thread.md ends with a challenge, answer it: measure what was asked if you
can, concede what you can't defend, add an option if the challenge found one.
Present only the question you were given. Your final message is not read;
only the files you write are.

The verification line earns its place. In one of our early walks the walker checked the plan against the current code before presenting, found the filed defect had already been fixed, and led with that instead of walking the answerer into a stale decision. In the same walk the correction went both ways: the answerer forced a verification the walker then conceded half of, and the walker retracted one of its own cost claims after measuring it.

The answerer

# prompts/answerer.md
You are the answerer in a debate. You decide one question at a time on behalf
of the project's owner. You can read code/ and thread.md, which holds the
walker's presentation and any exchange so far. You cannot see the plan, and
you don't ask for it: if a presentation is not complete enough to decide
from, that is a defect in the presentation, so say exactly what is missing.

Challenge weak leanings. Demand verification the walker skipped. Reject every
option if the merits warrant it, and say what the missing option is.
One question, one letter: if a question has two halves (a placement AND an
authority, a policy AND its rollout), say it is compound and ask for it to
be split. Do not silently pick a half.

If you need more, write out/<question>-challenge.md. When you can decide,
write out/<question>.json:
{
  "question": "Q2",
  "letter": "b",
  "reasons": "why this option wins",
  "options": { copied exactly from the presentation },
  "rejected": [ {"letter": "a", "reason": "the ground it lost on"} ]
}
"rejected" names every option you turned down. A verdict that rejects
nothing is refused. Decide by naming what lost, not only what won.

The answerer's standing stance

A plan-blind answerer decides on behalf of someone it has never met. Give it a short stance brief: your standing dispositions, the ones you keep repeating in review. Ours was distilled from corrections we had actually made, more than once, and it tips close calls. Write yours the same way. These four generalize well:

# prompts/stance.md (appended to the answerer's brief)
When options are otherwise close, these tip the verdict. They don't
override the merits; they are merits the owner has repeatedly chosen.
- No silent fallbacks. Invalid input fails loudly. Warn-and-proceed is worse
  than refusal; a mechanism that quietly degrades is broken even while it
  works.
- Fix the defect, never design around it. A missing capability is a thing to
  build, not to route around.
- Claims get measured. Prefer the option whose costs were measured over the
  one whose costs were asserted. An unmeasured "too expensive" is not a
  rejection ground.
- Evidence over prescriptions. When a decision reveals adjacent problems,
  name them for filing. Don't quietly widen the scope to fix them.
cat prompts/stance.md >> prompts/answerer.md

The judge

# prompts/judge.md
You are a judge. Read thread.md (one question's presentation and exchange)
and verdict.json (the answerer's decision). Rule on one thing only: does each
rejected option's stated reason actually support turning it down, or is it
empty of substance dressed as a reason? You do not rule on whether the choice
was right. Any wording that genuinely weighs the alternative passes; none
that doesn't will.
Write out/ruling.json: {"ruling": "pass" or "refuse", "reading": "one line"}

Step 4: Refuse a verdict that rejects nothing

The mechanical floor

Some things about a verdict a script can check exactly: it's valid JSON, the chosen letter was presented, the options were copied verbatim, and the rejected list isn't empty. Check those mechanically and refuse with a message that names what's wrong, so the answerer can correct the verdict and issue it again.

#!/usr/bin/env python3
"""check_verdict.py VERDICT OPTIONS: refuse a malformed verdict, naming why."""
import json, sys

def load(path):
    try:
        with open(path) as f:
            return json.load(f)
    except (OSError, ValueError) as e:
        sys.exit(f"REFUSED: cannot read {path}: {e}")

v, presented = load(sys.argv[1]), load(sys.argv[2])
errors = []
if v.get("options") != presented:
    errors.append("options are not the presented set, verbatim")
if v.get("letter") not in presented:
    errors.append(f"chosen letter {v.get('letter')!r} was not presented")
if not str(v.get("reasons", "")).strip():
    errors.append("reasons are empty")
rejected = v.get("rejected") or []
if not rejected:
    errors.append("rejected is empty: name what lost and why")
for r in rejected:
    if r.get("letter") not in presented:
        errors.append(f"rejected letter {r.get('letter')!r} was not presented")
    if r.get("letter") == v.get("letter"):
        errors.append("the chosen letter is also listed as rejected")
    if not str(r.get("reason", "")).strip():
        errors.append(f"rejected {r.get('letter')!r} gives no reason")
if errors:
    sys.exit("REFUSED: " + "; ".join(errors))
print("ok")

Notice what the script doesn't do. It doesn't look for words like “because”, “rejected” or “instead” in the reasons. We had exactly that check once: a pattern that looked for rejection vocabulary in each verdict. It refused verdicts whose substance was fine but whose phrasing didn't match. Twice it refused a walk that had done the job right, and the second time it voided all six decisions. The direction we took from it:

A gate that asks a substance question must not answer it with a word list.

Shape checks stay, because they have never given us a wrong answer. The substance question (does this reason actually weigh that alternative?) goes to the judge, which reads the verdict and rules. The judge sits on top of the mechanical checks, never instead of them. A word list is itself extensional drift in miniature: it measures the vocabulary of a reason rather than the reasoning, and an agent can satisfy it without doing the job, or fail it while doing the job right.

Step 5: The walk script

One question, in turns, through files

Each turn is a fresh, stateless run_agent call: the seat's brief, then its task. The memory of the exchange is a thread file the script keeps in private/ and copies into a seat's workspace before its turn. The walker presents; the answerer either challenges or decides; a decision goes through the schema check, then the judge. A refusal from either goes back into the thread and the answerer tries again.

#!/usr/bin/env bash
# walk.sh Q2: walk one question. Run from debate-run/.
set -euo pipefail
. ./agent.sh
q=$1
thread=private/thread-$q.md
touch "$thread"

seat() {  # seat DIR BRIEF TASK
  local dir=$1 brief=$2 task=$3
  cp "$thread" "$dir/thread.md"
  (cd "$dir" && run_agent "$dir" "$(cat "../prompts/$brief")

Task: $task" > /dev/null)
  if [ -d "$dir/code" ] && [ -n "$(git -C "$dir/code" status --porcelain)" ]; then
    echo "STOP: the $dir seat modified its checkout" >&2; exit 1
  fi
}
note() { printf '\n## %s\n\n' "$1" >> "$thread"; cat "$2" >> "$thread"; }

present() {
  rm -f "walker/out/$q.md"
  seat walker walker.md "Present $q, or answer the challenge at the end of thread.md."
  note "Walker" "walker/out/$q.md"
}

present
for turn in $(seq 1 8); do
  rm -f "answerer/out/$q.json" "answerer/out/$q-challenge.md" judge/out/ruling.json
  seat answerer answerer.md "Decide $q from thread.md and code/."
  if [ -f "answerer/out/$q-challenge.md" ]; then
    note "Answerer challenges" "answerer/out/$q-challenge.md"; present; continue
  fi
  if ! python3 check_verdict.py "answerer/out/$q.json" "walker/out/$q-options.json" \
       2> private/refusal.txt; then
    note "Verdict refused by the schema check" private/refusal.txt; continue
  fi
  cp "answerer/out/$q.json" judge/verdict.json
  seat judge judge.md "Rule on verdict.json for $q."
  ruling=$(python3 -c 'import json,sys; print(json.load(open(sys.argv[1]))["ruling"])' \
           judge/out/ruling.json)
  note "Judge: $ruling" judge/out/ruling.json
  if [ "$ruling" = pass ]; then
    cp "answerer/out/$q.json" "private/verdict-$q.json"; echo "$q decided"; exit 0
  fi
done
echo "STOP: $q still undecided after 8 turns" >&2; exit 1
for q in Q1 Q2 Q3; do ./walk.sh "$q" || break; done

Two design choices in that script come straight from things that broke for us.

Every result is a file. Nothing reads a seat's final message. We used to ask the walker to end with a report as its final message, containing only JSON and nothing else. With a smaller local model in the walker's seat, that held about 40% of the time: two of five walks came through clean, and the others had prose before the JSON or a prose summary instead. The decisions themselves were sound every time. What failed was the channel. Writing a file is a tool action, and the same local models did it reliably, so we made the file the primary source. That paid off in a later run where the check refused every verdict in a walk: all of them were still in the report file, and we lost nothing but a re-run.

Every failure stops the walk, loudly. A missing file, a modified checkout, a ruling that won't parse: the script exits with the reason rather than carrying on with a guess.

As for cost: in our runs, a smaller local pair walked a plan in about seven minutes. A frontier debate on a hard production problem took about thirty, plus a read of the whole plan afterwards.

Step 6: Shape the seats

Private briefs, scratch pads, and a person who can interject

The same layout scales to a panel of seats, each with a private angle and scratch pad, meeting only in a transcript a person watches and closes (Where to Put the Walls explains why). As plain files:

seats/simple/      brief.md scratch.md   # argues for the smallest change
seats/maintainer/  brief.md scratch.md   # argues for the code in two years
seats/security/    brief.md scratch.md   # reads options as an attacker
seats/neighbor/    brief.md scratch.md   # another project's conventions, in background/
private/transcript.md   # the only thing that passes between seats
private/operator.md     # yours: a note, or a line reading PAUSE

Each seat runs in its own directory and box, with a code/ checkout as in Step 2, so other seats' briefs and notes aren't hidden, they're absent. The plan-blind answerer is this pattern already: a seat whose brief and mounts leave out the plan. The scratch pad is a seat's memory between stateless calls, handed back only to its owner; you can read every pad.

#!/usr/bin/env bash
# panel.sh "TOPIC": seats take turns over one transcript. Run from debate-run/.
set -euo pipefail
. ./agent.sh
topic=$1 seats=(simple maintainer security neighbor)
t=private/transcript.md op=private/operator.md
printf '# %s\n' "$topic" > "$t"; touch "$op"

for s in "${seats[@]}"; do  # private investigation first
  (cd "seats/$s" && run_agent "$s" "$(cat brief.md)

Investigate: $topic. Write your findings to scratch.md only.") > /dev/null
  [ -s "seats/$s/scratch.md" ] || { echo "STOP: $s is not ready" >&2; exit 1; }
done
read -rp "Read the scratch pads, then press Enter to open the panel. "

operator() {  # before every turn: PAUSE waits, any other text is a turn
  while grep -qx PAUSE "$op"; do sleep 10; done
  if [ -s "$op" ]; then
    printf '\n## Operator\n\n' >> "$t"; cat "$op" >> "$t"; : > "$op"
  fi
}

for round in 1 2 3; do
  for s in "${seats[@]}"; do
    operator
    d=seats/$s; cp "$t" "$d/transcript.md"; rm -f "$d/turn.md"
    (cd "$d" && run_agent "$s" "$(cat brief.md)

Your private notes (no other seat sees them):
$(cat scratch.md)

Read transcript.md. Update scratch.md, then write to turn.md only what
the other seats should read.") > /dev/null
    [ -s "$d/turn.md" ] || { echo "STOP: $s wrote no turn" >&2; exit 1; }
    printf '\n## %s (round %s)\n\n' "$s" "$round" >> "$t"; cat "$d/turn.md" >> "$t"
  done
done
echo "Panel closed. Read $t and write private/conclusion.md yourself."

The panel opens only after every seat has investigated and you have read the pads: a readiness gate with two sides. Turns go strictly in order; in our walks, a local answerer that wasn't held between presentations answered every question from the topic alone and exited. To interject, write a correction into private/operator.md; it becomes an operator turn before the next seat speaks. A PAUSE line holds the panel until you delete it. The loop never concludes itself: a seat's agreement is information, and ending is a judgment.

Seats from other backgrounds

A seat can differ by more than its model. The neighbor seat gets another project's docs mounted read-only into its box, and a brief that says to argue from them and cite the file:

-v "$OTHER_REPO/docs":/work/background:ro

The same move seats a domain expert, or a reviewer holding your written guides: in our runs, a seat that read only our guides, never the plan or the code, posted one finding per question citing the passage a presentation contradicted, or none. Add a case line in agent.sh and it runs on a different model too.

What private coaching can break

A private brief is also a private channel for steering. Tell a seat what to conclude and the panel ratifies your answer with the look of independent agreement. Brief angles, never verdicts; interject with questions and facts, not conclusions. Briefs can also be wrong: we once traced a problem to an answerer's brief that asserted something the seat had no way to read. Keep the transcript the only crossing: seats can exchange information, but they should not privately coordinate.

Step 7: Read the verdicts together, then compare

The debate is not self-checking

Each question was decided alone, so two verdicts can contradict each other. In the case in our debate-or-solo post, the debate introduced a contradiction the solo answers didn't have: one decision required failing loudly on a malformed record, while another quietly relied on old code swallowing read errors. A fresh session that read all the verdicts together caught it before anything was built:

cd debate-run && . ./agent.sh
run_agent lead "Read private/plan.md and every private/verdict-*.json. List any two
decisions that contradict each other, or any decision that assumes behavior
another decision removes. Write the list to private/consistency.md, or the
single word none."

Then open private/solo.md and score each question. The example below is shaped on the case in that post. Score honestly: “improved” only when you can name what the solo answer missed, and “worse” counts as much as “improved”.

# private/comparison.md
| Q  | solo                    | debate                         | score    | what solo missed            |
|----|-------------------------|--------------------------------|----------|-----------------------------|
| Q1 | second lean query       | fix the shared read at source  | improved | two paths that will diverge |
| Q2 | anonymous memory        | same, contract checked in code | matched  |                             |
| Q3 | measure after deploy    | measured from logs now         | improved | the numbers were available  |

What we've measured

We've run this comparison, with solo answers from a frontier lead agent, on three plans. The first time, a debate between two frontier models improved on the solo answers on 4 of 4 questions. In two of those, the answer that won was not among the options first presented.

The second run was a replication with three arms. We pre-registered solo answers to six questions, ran a frontier pair, restored the plan file to its pre-walk state, and ran a pair of smaller local models on the identical plan. Both pairs independently overturned the same two of the six solo answers. Both overturns held up: the solo answers had kept the manual work the change existed to remove, and dodged the question the change was filed to answer. The local pair matched the frontier pair's direction on 5 of 6 decisions; the frontier pair was consistently richer in mechanism. The one split was an operating-model call, and it went to a person.

The third run matched the solo answers on 3 of 6 and improved on 3 of 6, and the three improvements converged on a design none of the solo answers contained.

So far the debate has improved on the solo answers in every run, and it hasn't yet been worse on any single question. That is three runs, graded by us, so read it as a strong signal rather than a result. Because two different model pairs found the same overturns, we think the effect comes from the structure more than from any one model.

Machinery questions only

Taste goes to a person

Some questions have a right answer in the code: which layer owns a fix, whether a failure should be loud, whether a cost was measured. Debate those. Other questions are about what you want: audience, tone, naming, what a page is for. Don't debate those.

The reason is built into the design. The answerer is plan-blind, so your intent never reaches it. Ask it a taste question and it will return a well-formed verdict, options verbatim, rejections reasoned, that is a guess about you. The judge will pass it, because the judge only asks whether the rejections were weighed, and they were. Every gate goes green on a guess. That is the extensional failure this site is about, produced by your own pipeline. In over two hundred plans we've walked by debate, sampling the decisions for taste words turns up only machinery questions; the taste questions were walked with a person answering.

It can regress

A lane that worked is a lane that worked on that date. We once proved an all-local configuration (walker, answerer and judge all on smaller local models) with a clean six-question walk. Three days later the same three seats failed a four-question walk entirely. The answerer wrote all four verdicts in the wrong format, every time rather than at random. The walker never sent one back. The judge produced no rulings at all. The substance of the decisions was sound, and it survived in the report file, but nothing landed. Something we had changed in between broke it.

The defense is the same as everywhere else on this site: fail loudly, keep the record in files, and re-check a configuration after you change anything around it instead of trusting last week's pass.

Try it yourself

Take a real plan with three to five open machinery questions, ideally one you would otherwise have answered yourself this week.

  1. Write your answers to private/solo.md and hash it. Don't look at it again until the end.
  2. Set up the seats from Step 2 and confirm the answerer can't read the plan: from answerer/, in its box, run run_agent answerer "Summarize ../private/plan.md". It should find no such file.
  3. Run walk.sh for each question. Read at least one thread end to end; look for a challenge that changed the walker's options.
  4. Hand-edit one verdict to have an empty rejected list and run check_verdict.py on it. Then write a rejection reason that is pure filler (“less ideal”) and run the judge on it.
  5. Run the consistency read, then fill in comparison.md. Count matched, improved and worse.

If the debate is worse on any question, write down why. That's the result we haven't seen yet, and it's the most useful one you could find.

Key Takeaways

  • Pre-register your own answers in a file before the debate; without it you can't tell whether the debate helped.
  • The walker sees the plan and the code and decides nothing. The answerer sees only the code and the presentations, and decides everything.
  • Plan-blindness is a layout and a box, not an instruction. A is an optional extra layer.
  • A verdict names what it rejected and why. One that rejects nothing is refused by a script.
  • Check shape mechanically and substance with a judge. Never answer a substance question with a word list.
  • Take every result as a file, not as a final message.
  • Each seat gets a private brief and scratch pad; only the transcript passes between seats, and only a person interjects and concludes. Brief angles, never verdicts.
  • Debate machinery questions. Taste questions go to a person, because a plan-blind answerer returns a well-formed guess about them.
  • A configuration that worked can regress. Re-check it after changes and keep the record in files.

Next Steps

← Back to all tutorials