Someone asked us a pointed question about multi-model systems: if you put a model with guardrails in a room with a model without them, can one steer the other?

Yes, it can. Put two models in a shared conversation and each one's output becomes the other's input. Framing travels, and so do assumptions and confidence. It doesn't have to be adversarial, either. Two perfectly well-behaved will happily talk each other into the wrong answer. A code review where both reviewers approve a bug, because the first said it looked fine and the second read that first. A research summary where every agent echoes the framing of whichever one spoke first.

The usual fix is to write better prompts: be skeptical, disagree respectfully, think independently. It doesn't work, and the reason is structural. Once an agent can read another agent's reasoning, that reasoning is part of its context. You can ask it to ignore what it has read. You can't make it unread it.

The answer we arrived at, and built into our deliberation platform, Deliberative AI, fits in one sentence:

Agents can exchange information, but they cannot privately coordinate.

The rest of this post is the seven boundaries that make that sentence true, each paired with the failure it prevents.

The lifecycle

In brief: an orchestration starts with at least two agents. Each is isolated; they can't talk to each other until they're brought into a meeting. Before that, a person (or another agent) coaches each one privately, and each writes its thinking to its own scratch pad. When an agent is ready, it's marked ready. The meeting starts, the agents share what they learned, and they try to reach consensus. Someone watches the meeting and can pause it. When it's done, a person concludes it, and the agents go back to private mode, where they can no longer talk to each other and are coached individually on the work that follows.

The deliberation lifecycle Agents A and B work privately in separate lanes, each with its own scratch pad. A two-sided readiness gate admits them to a shared meeting, which can be paused. A human concludes the meeting, and the agents return to separate private lanes. PRIVATE Agent A own scratch pad Agent B own scratch pad coached 1:1 READY ✓ ✓ both sides MEETING A ↔ B seek consensus ❚❚ pause CONCLUDE H person only PRIVATE Agent A does its task Agent B does its task no channel between them
The only shared space is the meeting, and only people open or close it. Everything else is one agent and its coach.

Seven boundaries

1. Private coaching

Each agent gets its brief one to one, in its own channel. The failure it prevents is anchoring on the room: when every agent hears the same framing at the same time, you get several copies of one perspective and call it diversity. Coaching each agent separately lets you give them different angles on purpose (one looks at cost, another at failure modes), and neither sees the other's brief.

2. Context isolation

Outside a meeting, agents have no channel to each other at all. Not a restricted one, none. This prevents backchannel convergence: agents that can read each other's work in progress drift toward one another before any deliberation starts, and the meeting becomes a formality that ratifies an agreement nobody examined.

3. Isolated scratch pads

Each agent thinks in writing, in a scratch pad only it and its coach can read. This prevents premature commitment. An agent that must state positions in public defends them. An agent that can work privately can change its mind cheaply, and the person coaching it sees the reasoning, not just the conclusion.

4. Readiness gates

A meeting can't start until both sides agree it should: the agent signals it has finished investigating, and the marks it ready. Either side alone isn't enough. This prevents the rushed meeting, where one agent has done the work, the other hasn't, and the prepared one sets the terms because it's the only one with anything to say.

5. Meeting pause

Whoever is watching a meeting can stop it, for everyone or for a single agent, add a message, and resume. This prevents runaway consensus, the conversation that locks onto an early idea and keeps building on it. The intervention doesn't have to come from a person. The pause is a generic intervention point, and an oversight agent can use it too. That matters for the original question: an agent with different constraints can be the one that steps in.

6. Conclusion by a person

Agents can signal that they believe they have reached consensus. They can't end the meeting; a person concludes it. This prevents implicit authority: meeting output treated as a decision just because the agents stopped talking. The consensus signal is information, and concluding is a judgment.

7. Private coaching after the meeting

Once a meeting concludes, the agents go back to isolation, and the follow-up work is assigned one to one. This prevents the meeting leaking into execution. The meeting exists to align, and doing the work happens afterwards. An agent executing its part shouldn't be renegotiating it with its peers mid-task, and here it can't.

Why prompts can't do this

Every one of these could be written as an instruction. “Don't be swayed by the other agent.” “Don't end the discussion until the user agrees.” “Keep your working private.” The trouble is that an instruction is one more piece of context, competing with everything else in the window, including the other agent's persuasive paragraph. A boundary doesn't compete. If there is no channel, nothing crosses it, however persuasive it is.

This is the same distinction that runs through everything on this site. A prompt asks for the intended behavior. The architecture makes the unintended behavior impossible, which is a much stronger promise.

What this does not solve

Walls prevent process failures: premature consensus, leaked framing, conversations that decide by stopping. They don't fix capability. Two agents that share the same wrong belief will agree on it, and no amount of isolation will surface it. Isolated reasoning helps most when the agents really do differ (different models, different briefs, different evidence), and less when they are two copies of the same thing.

It is also not free. Seven boundaries mean more steps, more waiting and a person in the loop. For a quick, low-stakes task that's overkill; one good agent will do. The architecture earns its cost on work that is high-stakes, long-running or adversarial, where a confident wrong answer is expensive.

Boundaries first

If you're designing a multi-agent system, treat coordination boundaries as architecture, not as safety features added once the agents have started drifting. Decide early where agents can see each other, who opens and closes the shared space, and what stays private. Then build it so the answers are enforced, not requested.

Our short version: context isolation, private coaching, and a human in the loop.