Our concepts page argues that two on different paths catch what one agent misses. It's easy to believe and hard to check, because you rarely get to see both answers to the same question. Yesterday we did, on a real bug, and the difference was bigger than we expected.

The problem

Our agent-orchestration service, a Python web service, started climbing from 576 MB to 2 GB of memory within eleven minutes of a restart. Its event loop was stalling for up to five and a half seconds at a time, and simple requests took thirty seconds or more to answer. Heap attribution pointed at one place: a status panel that read the service's whole history of work attempts and decoded every stored JSON column, about 108 MB per read, to compute two spending totals.

As usual, a planning agent read the code and wrote the fix's plan with five open questions: is this a leak or churn, which memory counter carries the growth, is the history read in scope, how often is it called, and what was holding 108 MB when the measurement window closed.

The experiment

Both arms read the same prepared plan.

Then we compared the two sets of answers, question by question.

Where they agreed

On the diagnosis, both arms said the same thing: this was churn, not a leak. Large temporary decodes pushed the heap's high-water mark up, and the allocator didn't give the memory back. Nothing was retained; it was the size of each read, repeated. Both named the same memory counter. Both said to fix the read.

What the debate found that the solo agent missed

1. The real multiplier

The status build had recently been moved off the event loop into worker threads. That was the fix for the stalls, and the solo agent treated it as harmless. The debate measured it. The same decode had run serially on the event loop for about 47 hours before the move. Memory growth per tick went from a median of 4.9 MB before the move to 113.8 MB after it, about 23 times more. The change that cured the stalls had multiplied the memory, because now several threads were each decoding 108 MB at the same time.

That changes what the fix has to do. Making each read smaller helps. Making sure only one read runs at a time is what bounds the peak.

2. One read path, not two

The solo fix added a second, lean query just for the spending totals, next to the existing full read, with a test to keep the two in agreement. That works, but it leaves two ways of reading the same records, and they'll eventually disagree. The debate fixed the shared read at its source instead: select only the scalar columns and the handful of JSON keys the callers actually use, with SQLite's JSON operators, so all six callers get cheaper at once.

# before: every caller paid for every column
rows = [decode_all_json_columns(r) for r in db.execute(
    "SELECT * FROM attempts")]

# after: read only what is used, inside SQL
rows = db.execute("""
    SELECT id, round_id, model, runtime, cost_usd, started_at,
           evidence -> '$.meeting_id'     AS meeting_id,
           evidence -> '$.judge_verdict'  AS judge_verdict
    FROM attempts""")

That read about 29 times less text overall, and about 1,500 times less from the one JSON column that made up 99% of the bytes.

3. A decision the solo agent would have made by accident

The old decoder, when it met a malformed JSON value, logged a warning and carried on. Reading the value inside SQL behaves differently, and the debate noticed. A rewrite would silently change what happens to a corrupt record, so that behavior had to be chosen on purpose. The decision was to fail loudly: a malformed record raises an error naming the table, row and column, instead of quietly reading as empty.

4. A fix that doesn't erode

The debate raised a question the solo agent never asked. The history grows by a few hundred records a day. A leaner read is cheaper today, but its cost still grows with the table, so the peak would creep back up. At about five times today's row count, the solo fix would have been back where we started.

The debate added a bound that doesn't grow: a single-flight wrapper around the read. Concurrent callers share one computation instead of each running their own. With one important twist: a caller who arrives while a computation is already running doesn't get that computation's result, which started before it asked. It waits for the next one, which every late arrival shares. So no caller ever gets a stale figure, and at most one read runs at a time.

class FreshSingleFlight:
    """One run per key at a time; late callers share the NEXT run."""

    def run(self, key, fn):
        with self._cond:
            state = self._keys.setdefault(key, KeyState())
            if state.running is None:
                batch = state.running = Batch()  # nobody running: go
            else:
                batch = state.pending = state.pending or Batch()
                while not batch.done:
                    if state.running is None and state.pending is batch:
                        state.running, state.pending = batch, None
                        break  # run it for everyone waiting
                    self._cond.wait()
                else:
                    return batch.result()  # someone ran it for us
        try:
            batch.value = fn()
        except BaseException as exc:
            batch.error = exc  # shared with every waiter
        with self._cond:
            batch.done, state.running = True, None
            if state.pending is None:
                del self._keys[key]
            self._cond.notify_all()
        return batch.result()

Side by side

QuestionSoloDebate
Leak or churn?ChurnChurn, plus a measured control showing threading multiplied it about 23×
Which counter?Anonymous memorySame, with the counter's contract checked in code
The fixA second, lean query beside the old oneFix the shared read at its source; all callers benefit
Call ratesMeasure after the deployMeasured from logs: about 590 calls an hour, up to 4 at once
What held 108 MB?Reads in flight (inferred)Reads in flight (reproduced: 111 MB in a replay)
Raised by the debate—Single-flight, so the bound doesn't grow with the data
Raised by the debate—A finer memory instrument for the growth still unexplained (split into its own task)

What it cost

About thirty minutes of debate time, then the whole-plan read. The debate also introduced a contradiction the solo answers didn't have: one decision required failing loudly on a malformed record, while another quietly relied on the old code swallowing read errors. The whole-plan read caught it, and it was settled before anything was built. So the debate is better, but it isn't self-checking. The step after it earned its place too.

The honest limits

So, is it better?

In this case, clearly. Both arms would have fixed that night's symptom. Only the debate's version names the real cause, keeps one code path, makes the error behavior a decision rather than an accident, and holds as the data grows. The service went back into production the same night; whether the memory stays flat under a normal day's load is the check we're watching now.

The broader lesson is the one this site started with. A strong model answering alone takes one path through the problem and answers well along it. The things it missed weren't hard. They were off its path: a side effect of an earlier fix, a behavior change in error handling, a question about next month rather than tonight. A second perspective found them because it was looking from somewhere else.