Our concepts page argues that two on different paths catch what one agent misses. It's easy to believe and hard to check, because you rarely get to see both answers to the same question. Yesterday we did, on a real bug, and the difference was bigger than we expected.
The problem
Our agent-orchestration service, a Python web service, started climbing from 576 MB to 2 GB of memory within eleven minutes of a restart. Its event loop was stalling for up to five and a half seconds at a time, and simple requests took thirty seconds or more to answer. Heap attribution pointed at one place: a status panel that read the service's whole history of work attempts and decoded every stored JSON column, about 108 MB per read, to compute two spending totals.
As usual, a planning agent read the code and wrote the fix's plan with five open questions: is this a leak or churn, which memory counter carries the growth, is the history read in scope, how often is it called, and what was holding 108 MB when the measurement window closed.
The experiment
Both arms read the same prepared plan.
- Solo. Claude Opus 5.5, the agent driving the work, wrote its own answers to all five questions before any debate ran, into a file the debating agents couldn't read.
- Debate. A debate walk: Claude Opus 5.5 presented each question and MiniMax M3.1 Flash Preview argued it. The debate raised two more questions of its own and decided all seven. Afterwards a separate agent (Claude Sonnet 5.5) read the whole plan for contradictions. Claude Opus 5.5 implemented it and DeepSeek v4 Flash reviewed the result.
Then we compared the two sets of answers, question by question.
Where they agreed
On the diagnosis, both arms said the same thing: this was churn, not a leak. Large temporary decodes pushed the heap's high-water mark up, and the allocator didn't give the memory back. Nothing was retained; it was the size of each read, repeated. Both named the same memory counter. Both said to fix the read.
What the debate found that the solo agent missed
1. The real multiplier
The status build had recently been moved off the event loop into worker threads. That was the fix for the stalls, and the solo agent treated it as harmless. The debate measured it. The same decode had run serially on the event loop for about 47 hours before the move. Memory growth per tick went from a median of 4.9 MB before the move to 113.8 MB after it, about 23 times more. The change that cured the stalls had multiplied the memory, because now several threads were each decoding 108 MB at the same time.
That changes what the fix has to do. Making each read smaller helps. Making sure only one read runs at a time is what bounds the peak.
2. One read path, not two
The solo fix added a second, lean query just for the spending totals, next to the existing full read, with a test to keep the two in agreement. That works, but it leaves two ways of reading the same records, and they'll eventually disagree. The debate fixed the shared read at its source instead: select only the scalar columns and the handful of JSON keys the callers actually use, with SQLite's JSON operators, so all six callers get cheaper at once.
# before: every caller paid for every column
rows = [decode_all_json_columns(r) for r in db.execute(
"SELECT * FROM attempts")]
# after: read only what is used, inside SQL
rows = db.execute("""
SELECT id, round_id, model, runtime, cost_usd, started_at,
evidence -> '$.meeting_id' AS meeting_id,
evidence -> '$.judge_verdict' AS judge_verdict
FROM attempts""")
That read about 29 times less text overall, and about 1,500 times less from the one JSON column that made up 99% of the bytes.
3. A decision the solo agent would have made by accident
The old decoder, when it met a malformed JSON value, logged a warning and carried on. Reading the value inside SQL behaves differently, and the debate noticed. A rewrite would silently change what happens to a corrupt record, so that behavior had to be chosen on purpose. The decision was to fail loudly: a malformed record raises an error naming the table, row and column, instead of quietly reading as empty.
4. A fix that doesn't erode
The debate raised a question the solo agent never asked. The history grows by a few hundred records a day. A leaner read is cheaper today, but its cost still grows with the table, so the peak would creep back up. At about five times today's row count, the solo fix would have been back where we started.
The debate added a bound that doesn't grow: a single-flight wrapper around the read. Concurrent callers share one computation instead of each running their own. With one important twist: a caller who arrives while a computation is already running doesn't get that computation's result, which started before it asked. It waits for the next one, which every late arrival shares. So no caller ever gets a stale figure, and at most one read runs at a time.
class FreshSingleFlight:
"""One run per key at a time; late callers share the NEXT run."""
def run(self, key, fn):
with self._cond:
state = self._keys.setdefault(key, KeyState())
if state.running is None:
batch = state.running = Batch() # nobody running: go
else:
batch = state.pending = state.pending or Batch()
while not batch.done:
if state.running is None and state.pending is batch:
state.running, state.pending = batch, None
break # run it for everyone waiting
self._cond.wait()
else:
return batch.result() # someone ran it for us
try:
batch.value = fn()
except BaseException as exc:
batch.error = exc # shared with every waiter
with self._cond:
batch.done, state.running = True, None
if state.pending is None:
del self._keys[key]
self._cond.notify_all()
return batch.result()
Side by side
| Question | Solo | Debate |
|---|---|---|
| Leak or churn? | Churn | Churn, plus a measured control showing threading multiplied it about 23× |
| Which counter? | Anonymous memory | Same, with the counter's contract checked in code |
| The fix | A second, lean query beside the old one | Fix the shared read at its source; all callers benefit |
| Call rates | Measure after the deploy | Measured from logs: about 590 calls an hour, up to 4 at once |
| What held 108 MB? | Reads in flight (inferred) | Reads in flight (reproduced: 111 MB in a replay) |
| Raised by the debate | — | Single-flight, so the bound doesn't grow with the data |
| Raised by the debate | — | A finer memory instrument for the growth still unexplained (split into its own task) |
What it cost
About thirty minutes of debate time, then the whole-plan read. The debate also introduced a contradiction the solo answers didn't have: one decision required failing loudly on a malformed record, while another quietly relied on the old code swallowing read errors. The whole-plan read caught it, and it was settled before anything was built. So the debate is better, but it isn't self-checking. The step after it earned its place too.
The honest limits
- One case. This is one problem, one run. It's an anecdote with numbers attached, not a study.
- Same model on both sides. The solo agent and the debate's presenter were the same model. This compares a debate with one agent answering alone. It doesn't compare models.
- One question was partly primed. The solo agent had written its diagnosis in the task's notes before the experiment was proposed, and the plan quoted it. The fix questions, and the two the debate raised, are the clean comparison.
- Effort, not only perspective. The debating agents ran measurements the solo agent chose not to: log analysis, a replay, the before-and-after split. Part of the gain is work, not just a second point of view. Then again, getting that work done is part of what a debate is for.
So, is it better?
In this case, clearly. Both arms would have fixed that night's symptom. Only the debate's version names the real cause, keeps one code path, makes the error behavior a decision rather than an accident, and holds as the data grows. The service went back into production the same night; whether the memory stays flat under a normal day's load is the check we're watching now.
The broader lesson is the one this site started with. A strong model answering alone takes one path through the problem and answers well along it. The things it missed weren't hard. They were off its path: a side effect of an earlier fix, a behavior change in error handling, a question about next month rather than tonight. A second perspective found them because it was looking from somewhere else.