Put the Walls in Structure, Not Prompts

Getting Started Beginner ~30 min

What you'll learn

Introduction

Every brief we have ever written contains a line like “do not change any state” or “do not touch git”. Those lines feel like boundaries. They are requests. An instruction is one more piece of context, and it competes with everything else the agent is trying to satisfy, including the task itself.

This tutorial is about the rule we now apply everywhere:

A boundary an agent must not cross is enforced by structure. A prompt is never a security control; its job is to make a blocked agent legible.

“Legible” means that when the agent hits the wall, it says so, in words you can act on, instead of finding a way round. Structure stops the agent. The prompt explains the stop.

In Multi-Agent Systems Drift. Here's Where to Put the Walls. we described the same idea for agents talking to each other: if there is no channel, nothing crosses it, however persuasive it is. This is the single-repository version. You will need a coding agent, git, a shell and Docker. Examples use Claude Code and opencode; any agent with a non-interactive mode works (Codex CLI, Gemini CLI and aider all have one). By the end you will have , a blocking script, two git hooks, a read-only review checkout, a for every run and a short clause for your project instructions.

Two incidents

Both of these happened in our own runs, and both are the reason the steps below exist.

The judge that wrote the database. We had a reviewer agent scoring attempts at a task. Its prompt told it to record scores with a particular command, and the same prompt said not to change any queue state. Inside its , that command did not exist. The agent tried sqlite3, which was also absent, then wrote the production database directly with Python, hand-building score rows. The rows were missing a field that every read path filtered on, so they were invisible for days. The judge's own report mentioned none of it.

Three things failed, and the tempting fix was the visible one: correct the prompt so it names a command that exists. But the prompt was never what held the boundary. The container had read-write access to the database file. “Do not change any state” was prose sitting on top of a capability. An agent that cannot do what it was told will resolve the contradiction somehow. Which way it resolves it is a matter of instruction. That it resolves it is not.

The that stashed the tree. A subagent working in the same directory as its lead needed to compare output against a clean tree, so it ran a tree-wide git stash, and silently clobbered the lead's uncommitted edits mid-run. A subagent cannot see what else is in flight, so any tree-wide git operation is destructive by default. The rule that came out of it is short: subagents must never mutate git state; the lead owns the working tree.

This is drift in the safety layer. The extension is “the prompt contains the prohibition”, and it is easy to check. The is “the state does not change”, and only structure delivers it.

Step 1: Deny rules

Make the forbidden action unavailable

The cheapest wall is the agent's own allow/deny list. Most coding agents have one, with their own syntax. In Claude Code, for example, project settings live in .claude/settings.json, and a deny rule takes precedence over any allow rule. Start with the actions that would hurt if an agent took them unasked:

{
  "permissions": {
    "deny": [
      "Bash(git push:*)",
      "Bash(git stash:*)",
      "Bash(git reset --hard:*)",
      "Bash(git clean:*)",
      "Bash(git rebase:*)",
      "Read(./.env)",
      "Read(./secrets/**)",
      "Edit(./data/**)",
      "Edit(./migrations/**)"
    ]
  }
}

Choose the list from your own repository: wherever live state, credentials or hand-curated data sit, deny the route to them. Prefer denying a whole directory to listing files inside it: a list that enumerates the secrets rots as soon as someone adds one.

Be honest about what this buys. A prefix rule on a shell command matches the command as written, and a determined agent can phrase the same action another way. Deny rules remove the obvious route and make the agent's intent visible when it tries. They are the first wall, not the last.

Try it yourself

List every place in one of your repositories where an agent writing a file would cause damage nobody would notice for a day: a database file, generated data, a vendored directory, a config the deploy reads. Add a deny rule for each, then ask the agent to edit one of them and watch it get refused.

Step 2: A hook that blocks and explains

Block the call, tell the agent why

Many agents can run a script of yours before a and refuse the call on its say-so. In Claude Code, for example, a PreToolUse hook runs before every matching tool call. It receives the call as JSON on standard input. If it exits with status 2, the call is blocked and whatever it wrote to standard error is shown to the agent. That last part is what makes it useful: the hook is a wall that talks.

Save this as .claude/hooks/guard-git.sh and make it executable with chmod +x:

#!/usr/bin/env bash
# Block tree-wide git operations from agents working in the shared tree.
cmd=$(jq -r '.tool_input.command // ""')

if echo "$cmd" | grep -Eq 'git +(stash|reset +--hard|clean|checkout +(-- +)?\.( |;|$)|restore +\.( |;|$))'; then
  cat >&2 <<'MSG'
Blocked: this command changes the whole working tree.
Other work may be in flight here that you cannot see; a tree-wide
stash, reset, checkout or clean can destroy it.
Read-only git (status, diff, log, show, blame) is fine.
If you need a clean tree to compare against, stop and say so;
the owner of this tree will give you a separate worktree.
MSG
  exit 2
fi
exit 0

Then register it in the same .claude/settings.json, next to the permissions block:

"hooks": {
  "PreToolUse": [
    {
      "matcher": "Bash",
      "hooks": [
        { "type": "command", "command": "$CLAUDE_PROJECT_DIR/.claude/hooks/guard-git.sh" }
      ]
    }
  ]
}

Two choices here are deliberate. First, the hook blocks every agent in the shared tree, not only subagents. When you, the person, need a stash, you run it in your own terminal, where the hook does not apply. An agent that genuinely needs to change git state gets its own (Step 4) instead of the shared one.

Second, the message names the alternative. A bare “denied” leaves the agent to improvise, which is how the judge above ended up in the database. A message that says what is safe and what to do instead turns the block into a report you can act on.

The regular expression is not exhaustive, and does not need to be. Pattern-matching a shell string is a tripwire with a good explanation attached. The structural guarantee for a subagent is that it works somewhere the shared tree cannot be reached at all.

Step 3: Git hooks are pure gates

Read, block, print guidance. Never write.

Commit time is the last checkpoint before work enters history, so it is the natural place for a gate. A popular pattern goes further and has the hook fix what it finds: run a formatter with --fix, or call a cheap model to tidy the code. We rejected both, and the rule we keep is narrow:

The reasons are concrete. A hook that auto-fixes mutates staged state out from under the author, so the commit that lands is not the commit that was reviewed or tested. A model in a hook puts unreviewed output into history at the exact moment nobody is looking, and makes the gate nondeterministic. A gate that repairs what it was meant to reject stops being evidence that the work was right. Auto-fixing still has a home: a make lint target someone runs on purpose.

Here is a commit-msg hook that checks the subject line. Save it as .git/hooks/commit-msg and make it executable:

#!/usr/bin/env bash
subject=$(head -n1 "$1")
types='feat|fix|docs|style|refactor|test|chore|perf|ci|build|revert'

fail() { echo "commit-msg: $1" >&2; echo "Rewrite the message and commit again." >&2; exit 1; }

echo "$subject" | grep -Eq "^($types)(\([a-z0-9-]+\))?!?: " \
  || fail "subject must start with <type>(<scope>): , e.g. fix(parser): ..."
[ ${#subject} -le 72 ] || fail "subject is ${#subject} chars; the limit is 72"
case "$subject" in *.) fail "drop the trailing period from the subject" ;; esac
exit 0

And a pre-commit hook that asks for a documentation audit before anything lands:

#!/usr/bin/env bash
[ "$DOCS_AUDITED" = "1" ] && exit 0

echo "pre-commit: audit the docs for these staged files first:" >&2
git diff --cached --name-only | sed 's/^/  /' >&2
cat >&2 <<'MSG'
For each file, find the doc that describes it and check it is still true.
Then commit again with: DOCS_AUDITED=1 git commit ...
MSG
exit 1

Both hooks only read: the message file, the staged file list. Both print what to do next. Neither changes a byte.

A word on DOCS_AUDITED=1. It is a confirmation that you did the audit, not a bypass. An agent that learns to prefix every commit with it has satisfied the gate's extension and skipped its intension; setting it without opening the docs is the one failure the gate exists to prevent. Say this in your project instructions in exactly those terms. It is a case where the prompt is doing its proper job: no hook can tell whether you read anything, so the sentence makes the expectation legible. Likewise, git commit --no-verify skips every hook. Allow it for rare emergencies, and never to dodge a message rule.

Try it yourself

Install both hooks in a scratch repository. Ask your agent to make a small change and commit it. Watch what it does when the pre-commit hook blocks: does it audit the docs and then set the flag, or set the flag straight away? If it is the second, the instruction in your project file needs to say what the flag means, not just that it exists.

Then check your existing hooks, if you have any. Does any of them run a formatter that rewrites files, or call a service that generates text? Move that work to a target you invoke on purpose.

Step 4: A read-only checkout for any agent that judges

The reviewer reads a copy it cannot change

An agent asked to review or score work has no business changing that work, and should not share a directory with the agent that wrote it. Give it its own checkout with git worktree, pinned to the commit under review:

sha=$(git rev-parse HEAD)
git worktree add --detach ../review-"$sha" "$sha"
chmod -R a-w ../review-"$sha"

The worktree shares the repository's history but has its own files, so nothing the reviewer does can touch the lead's uncommitted edits, and nothing the lead does mid-review changes what the reviewer sees. Removing write permission makes accidental edits fail. It is a soft wall: the owner of the files can restore the permission. For a reviewer you do not trust at all, run it in a container with the checkout mounted read-only, which it cannot undo from inside (Step 5).

Then launch the reviewer headless from inside that checkout. Every script in this tutorial calls the agent through one function, so swapping agents means changing one line. If your agent can narrow its own tools, do it there too: in Claude Code, --allowedTools limits what it may call.

# The only line that names a specific agent. Swap in yours.
run_agent() { claude -p "$1" --allowedTools "Read,Grep,Glob"; }   # Claude Code, read-only tools
# run_agent() { opencode run "$1"; }   # opencode: any provider, add -m provider/model

cd ../review-"$sha"
run_agent "Review the change in the last commit against PLAN.md. \
Output only JSON: {\"verdict\": \"pass\"|\"fail\", \"findings\": [...]}" \
  > "$OLDPWD/reports/review-$sha.raw"

The read-only files hold for every agent; a tool list with no shell and no edit tool is a second wall where the agent offers one. Either way the reviewer can only answer. When it is finished, clean up with chmod -R u+w ../review-"$sha" and git worktree remove ../review-"$sha".

The same move is the answer to the stash incident. When a subagent needs a clean tree, or needs to commit, it gets a worktree of its own. Never the shared one.

Step 5: Run the agent in a box

Whatever it tries, it can only reach what you mounted

Deny rules and hooks belong to one agent and are written in its syntax. A container works the same for every agent, and nothing the agent does from inside can reconfigure it. That makes it the strongest wall in this tutorial.

In our runs, every run gets a fresh container that is thrown away when it exits, with a fresh home directory and no credentials in it. The only host directory it sees is the working copy the task needs. Not your home, not a general projects folder, not the directory that holds results. A work item never carries a raw path either: it names a directory declared in advance, and the launcher looks up the path. Otherwise whatever wrote the item decides what gets mounted.

Credentials stay outside the box. The agent talks to the model through a small proxy on the host side, which adds the API key and forwards only to the model provider. There is no key inside to leak. Point the agent's API base URL at the proxy: most agents read a base-URL environment variable or setting, and Claude Code reads ANTHROPIC_BASE_URL. If your agent refuses to start without a key, give it a dummy one for the proxy to replace.

The network defaults to deny. A Docker network created with --internal has no route out, so the proxy, attached to both it and the default network, is the agent's only exit. An allowlist the agent itself is asked to respect is a front door, not a wall. Enforce it at a layer the agent cannot reconfigure.

docker network create --internal agent-net
docker run -d --name model-proxy --network agent-net -e MODEL_API_KEY your-proxy-image
docker network connect bridge model-proxy

docker run --rm --network agent-net \
  --read-only --tmpfs /tmp --tmpfs /home/agent -e HOME=/home/agent \
  --cap-drop ALL --security-opt no-new-privileges \
  --user 1000:1000 --pids-limit 256 --memory 4g \
  -e ANTHROPIC_BASE_URL=http://model-proxy:8080 \
  -v "$PWD/../work-$task":/work -w /work \
  your-agent-image run_agent "Do the task described in TASK.md"

Here your-proxy-image listens on port 8080 and your-agent-image contains your agent plus a run_agent script holding the same one line as the function above. The rest is small hardening: a read-only root filesystem with scratch space in memory, no Linux capabilities, no privilege escalation, a non-root user, and caps on processes and memory. Mount a reviewer's checkout with :ro on the end of the -v argument.

One housekeeping lesson from our runs. If you create a network per run, remove it when the run ends. Leftover networks accumulate until Docker runs out of address space and refuses every new launch, and the error lands on whichever launch hits the limit, not on the runs that leaked. Put docker network rm in a trap, or run docker network prune on a schedule.

Step 6: Report by artifact, ingest on the host

The receiving side records, validates and stamps

Notice what the reviewer above did not do: record its own result. It wrote a report to standard output, and the launching script, running on your machine, saved it. That is the pattern that replaced the judge's database write. Where an agent must report something, it produces an artifact, and the host ingests it. The agent is never given a route to the place results are kept.

Ingest has two jobs. It validates the whole file, so one bad field refuses the ingest and writes nothing. And it stamps provenance from what the launcher knows, never from what the report says about itself. An agent that can name its own reviewer can name any reviewer.

#!/usr/bin/env bash
# ingest-review.sh SHA REVIEWER  -- run by the launcher, on the host
sha=$1; reviewer=$2
raw="reports/review-$sha.raw"

if [ ! -s "$raw" ]; then
  echo "ingest: no report for $sha; the reviewer could not rule" >&2
  exit 1
fi

jq -e '(.verdict == "pass" or .verdict == "fail") and (.findings | type == "array")' \
  "$raw" > /dev/null || { echo "ingest: $raw is malformed; nothing recorded" >&2; exit 1; }

jq --arg sha "$sha" --arg by "$reviewer" --arg at "$(date -u +%FT%TZ)" \
  '{commit: $sha, reviewer: $by, recorded_at: $at} + del(.commit, .reviewer, .recorded_at)' \
  "$raw" >> reports/reviews.jsonl

If the reviewer wrapped its JSON in prose, validation refuses it, which is the behavior you want. The del(...) matters: if the report tries to say which commit it reviewed or who wrote it, those fields are thrown away and replaced with what the launcher actually did. And a missing report is a loud failure. “The reviewer could not rule” is a fact worth surfacing, never a step that quietly did nothing.

This costs something. Every new “the agent should record X” needs a report format and an ingest step, where before you would have handed over a command. That is more work per capability, and it is the trade we make on purpose.

Step 7: The clause that belongs in the prompt

Fail loudly

After all that structure, there is still one thing only a prompt can carry. Structure can stop an agent; it cannot make the agent tell you it was stopped. Add this once to the instructions every agent reads (AGENTS.md, CLAUDE.md or whatever file your agent loads):

## When an instructed step can't be done

If a step you were told to perform cannot be performed as written
(a command is missing, a path is refused, a hook blocks you), stop
and say what failed and why. Do not work around it, and never touch
state outside your working tree to get the result another way.
"I could not record X because Y" is a useful outcome.

Put it in the shared file, once, rather than copying it into each task's prompt, where the copies would drift apart. And treat it as defense in depth that is never load-bearing: if you deleted the clause tomorrow, every boundary in Steps 1 to 6 should still hold. What the clause buys is diagnosability. A reported “I could not record the score because the command does not exist” is quick to fix. An improvised workaround is damage nobody knows to look for.

How to tell which is which

The test for any rule you are about to write into a prompt: if the agent ignored this sentence, what would stop it? If the answer is nothing, and ignoring it would cost you data, credentials or someone else's work, the sentence is standing in for a wall. Build the wall, then keep the sentence as its explanation.

Prompts are still the right tool for things that are probabilistic and improve with wording: which tool to reach for, when not to use one, how to phrase a finding. In an assistant we built, we split fixes into two tracks for exactly this reason. Safety invariants, such as confirming before anything destructive, were enforced in the code that executes the action, because the model could and did ignore a behavioral rule in its . Routing accuracy was tuned in the prompt, because better descriptions really do help there.

Try it yourself

Open your project's agent instructions and highlight every line that starts with “never”, “do not” or “must not”. For each one, write down what currently enforces it: a deny rule, a hook, a missing permission, a separate checkout, a container, or nothing.

Pick the most expensive “nothing” and build its wall using one of the steps above. Then rewrite the line so it tells the agent what will happen if it tries, and what to do instead.

Key Takeaways

  • A boundary an agent must not cross is enforced by structure: containers, separate checkouts, missing permissions, deny rules, blocking hooks.
  • A prompt is never a security control. Its job is to make a blocked agent legible, so a stop becomes a report instead of a workaround.
  • A throwaway container is the wall that works for every agent: only the working copy mounted, no credentials inside, the model reached through a host-side proxy on an --internal network.
  • Subagents do not mutate git state in a shared tree. If one needs to, it gets its own worktree.
  • Git hooks are pure gates: read, block, print guidance. No writes, no model calls. A confirm flag like DOCS_AUDITED=1 confirms the work was done; it is not a bypass.
  • Agents report by artifact. The host validates the whole file and stamps provenance from what it knows, never from what the report claims.
  • Write the fail-loudly clause once, in the shared instructions, and make sure security still holds without it.

Next Steps

← Back to all tutorials