Plans That Record Decisions, Not Guesses

Getting Started Beginner ~25 min

What you'll learn

Introduction

Ask a coding to plan a feature and it will hand you a plan with every question answered. It reads well. It is also, mostly, the agent's guesses about what you want, written in the same confident voice as the facts it checked.

That is drift at the planning stage. The measurable stand-in for "planned" is a document where every question has an answer. What you meant by "planned" is a document recording what you chose, what you turned down, and why, so the person or agent who implements it builds the thing you had in mind. The two look identical from the outside. We once had a planning agent decide four questions in a single turn with nobody answering, and the plan passed every check we had at the time.

This tutorial sets up the planning process we use for any work where your judgment decides the outcome. It has three passes over one file: a prep pass that finds the questions, a walk by a different agent that puts them to you one at a time, and a finalize step that folds your answers into a recommendation. The plan file is the only place a decision lives. Not the chat, not a ticket, not the agent's memory of the session.

You need a coding agent, git and a shell. The examples use Claude Code and opencode; any agent that can read your repository and edit one file will do the same job, on a or a local one. By the end you'll have a plan template, a prep prompt, a walk prompt, and one real plan walked to a decision.

Step 1: The Plan Template

Build

Make a plans/ directory in your repository and save this as plans/TEMPLATE.md:

---
status: draft
---

# <Title>

## Problem

<The idea in your own words: several paragraphs, not a line.
Paste it as you wrote it. Do not let an agent paraphrase it.>

## Open Questions

<!-- The prep pass adds questions here. The walk adds one
decision line under each question it settles with you. -->

## Recommendation

<!-- Empty until every question is decided. -->

## Verification

<!-- Checks that only looking at the running thing can answer,
one per line. Leave this comment if there are none. -->

Four sections, each with one job. Problem holds your words, because the agents that walk and implement the plan should read what you said, not a summary of it. Write it in detail: in our runs the prep agent produces questions in proportion to what it is given, and a one-line idea starves it. Open Questions is the decision record and stays in the plan permanently. Recommendation is written last, from the decisions. Verification is for checks you can't settle by reading code.

The status line has two values you set by hand: draft while planning, final when the plan is ready to hand to an implementer. final does not mean built.

Step 2: The Prep Pass

Build

The prep agent's job is to find the decisions, not to make them. Think of a product manager scoping work: they check the stack for relevant facts ("there's already a table that could hold this") and pass those facts on with the question, without designing the table. Save this as plans/prep-prompt.md:

You are the prepping agent for the plan file named below.
Your job is to find the decisions in it, not to make them.

- Read the Problem section, then read the code it touches. Do not
  guess what a file contains; open it.
- Write questions under "## Open Questions" in this form:
    **Q1 — <one question>**
    Background: <the facts that ground it: code you read, a
    tension you found>
- State facts, not solutions. No options, no leanings, no
  "we could", in any section, in any format. Weighing options
  belongs to a later pass.
- If something is already decided (by me, or in the Problem
  section), record it and say who decided it. That is a fact.
- Mark every claim you did not verify, inline, under the
  question it grounds:
    **Unverified:** <what you did not check, and why it matters>
- One question, one answer. If a question has two "?", or an
  "and" joining two asks, split it into two questions.
- Tag at least one question with [change] inside the bold
  heading: a question whose answer could change something,
  not only confirm a fact about the code as it is.
- Do not say prep is finished. A later pass may add questions.
- You write exactly one file: the plan. No code, no fixes, not
  even a one-liner.
- End every reply with the full list of questions, one line
  each, marking the ones added this turn as new.

Plan: plans/<name>.md

Each rule is there because its absence went wrong for us.

No solutions. If prep lists options, the walk inherits a framing chosen before anyone looked at the whole set of questions, and the cheapest option on the list tends to look like the obvious one. Prep that only states facts leaves the option space open for the pass that has to weigh it.

Unverified, inline. A background paragraph with no mark reads as established fact, and a walk that decides on it is deciding on a guess. In one of our walks, three claims had to be re-measured in the middle of a decision because nothing had said they were shaky. The mark goes under the claim, not in a separate list, because that is where the next agent reads.

At least one [change] question. A plan made only of "does X exist?" questions can be walked to a decision on every one of them without anything changing. The tag isn't a priority rating, and it doesn't license proposing an answer. It guarantees there is something to decide.

One question, one answer. Each question ends with a single letter chosen from a set. "Where does the checkpoint go, and who owns its format?" can't be answered with one letter, so whoever answers it either answers in a shape nothing records or quietly drops half. Splitting it costs nothing at prep time.

Here's what a good prep question looks like, for a plan about webhook deliveries that fail:

**Q3 — What should happen to a webhook delivery that fails? [change]**
Background: send_webhook() in notify/webhooks.py makes one POST and
logs the exception on failure; nothing sends it again. The deliveries
table has no attempt counter.
**Unverified:** whether the job runner retries failed jobs by itself.
If it does, failures are already retried at a layer this question
would duplicate.

A fact, a tension, an honest gap. No "we could add backoff".

To run it, copy the template to a new file, write your Problem section, and start your agent with the prompt:

cp plans/TEMPLATE.md plans/webhook-retries.md
# write the Problem section yourself, then, in Claude Code:
claude "$(sed 's/<name>/webhook-retries/' plans/prep-prompt.md)"
# in opencode or any other agent: open it in the repository and
# give it plans/prep-prompt.md with the plan name filled in

Prep is interactive. Stay in the session, raise topics, say which ones deserve a question, and stop when you can't think of more.

Step 3: Walk With a Fresh Agent

Build

Close the prep session. Start a new one. The walk must be done by an agent that did not write the prep, because an agent walking its own prep defends it. A fresh agent receives the prep's claims as someone else's preliminary research and checks them.

That check earns its place. In one of our walks the fresh agent found four material errors in the prep: a symptom blamed on the wrong layer, a thing reported as missing that existed, file paths from before a refactor, and an incomplete picture of how two parts were linked. Decisions came out differently because of it. Walks also find questions prep missed: on one plan, three of the eleven questions decided were raised during the walk itself.

Save this as plans/walk-prompt.md:

You are the walking agent for the plan file named below. Another
agent prepped it. Treat its background text as preliminary research
to verify, not as fact.

Walk the open questions with me, one at a time:

1. Before presenting a question, re-check the claims it rests on,
   starting with anything marked **Unverified:**. Record what you
   checked in its background as
     Verified during the walk (<date>): ...
   replacing any Unverified mark it clears.
2. Present the question: the concrete tension, then a lettered set
   of real, distinct options. For each, say what would be built
   and what it costs or leaves unsolved. Not one-line labels.
   "Fix the underlying gap first" is a legitimate option.
3. Give your leaning, with its reasoning and an explicit rejection
   of the weaker options. Name what your choice costs: what it is
   worse at, and which rejected option was better at that. Weigh
   maintainability and the long term, not ease.
4. Stop and end your turn. Do not present the next question until
   I have answered this one.
5. When I answer, add one line directly under the question heading:
     **Qn — DECIDED (x): <what was chosen>, because <why>;
     rejected (y) <on what grounds>, (z) <on what grounds>.**
   (x) is the letter I chose. If I answer outside your set, record
   my answer in my words. Never write a DECIDED line for a question
   I have not answered. Your leaning is a recommendation, not a vote.
6. If the discussion raises a new question, add it to the plan.
7. End every reply with: n/m decided; remaining: Qx <one line>, ...

Edit the plan in place: change only the lines you mean to change.
Never rewrite the whole file or rebuild it by splitting on text.
You write exactly one file: the plan. No code.

Plan: plans/<name>.md

Then start a new session the same way, with the walk prompt. In any agent that means a session with no history of the prep; in Claude Code, for example, a new claude process, not --continue or --resume on the prep session.

What a question turn looks like

Q3 — What should happen to a webhook delivery that fails?
Verified during the walk (2026-10-01): the job runner does not retry;
an exception marks the job failed and it stops there. The missing
attempt counter is confirmed.

The tension: our own receivers usually recover within minutes, but one
customer endpoint rate-limits and rejects bursts.

(a) Retry in the job runner, for every job type, with backoff. One
    mechanism for everything, but jobs that are not safe to repeat
    would now repeat too.
(b) Retry inside send_webhook only, with backoff, counting attempts
    on the delivery row. Scoped to the one job known to be safe to
    repeat; a second mechanism if other jobs want retries later.
(c) No automatic retry; add a "resend failed" action. Nothing repeats
    unseen, but every routine blip waits for a person.

Leaning: (b). Retrying is only safe where the request is idempotent,
and webhooks are the only job we know is. Reject (a) for repeating
jobs nobody has checked, (c) for handing routine blips to a person.
What (b) costs: it's a second retry mechanism, and (a) would have
been the single one if every job later proves safe to repeat.

0/4 decided; remaining: Q1 who gets alerted, Q2 attempt limit,
Q3 failed deliveries, Q4 retention of failed payloads

Notice what the agent has not done. It hasn't decided anything. It verified the one claim the prep flagged, laid out the choice, said what its own preference gives up, and stopped.

Two tells that a leaning wasn't really weighed: the downside is hard to write at all, or a run of leanings all reduce to the same sentence ("one source of truth" winning every question). If that happens, ask the agent which axis it has been skipping.

Your side of the walk

Answer in your own words. A letter is fine; so is "b, but cap it at five attempts", or an answer none of the letters covered. Correct the framing when it's wrong. In one of our walks a correction from the person answering reshaped a decision halfway through it, which is the walk working as intended. The agent then writes:

**Q3 — What should happen to a webhook delivery that fails? [change]**
**Q3 — DECIDED (b): retry inside send_webhook with backoff and an attempt count on the delivery row, capped at five, because webhooks are the only job known to be safe to repeat; rejected (a) runner-wide retry, which would repeat jobs nobody has checked, and (c) manual resend only, which leaves routine failures waiting for a person.**

The question heading stays untouched. The decision line carries what was chosen, why, and what was rejected on what grounds. A bare "DECIDED (b)" loses the option set, and the option set is what tells a later reader that runner-wide retry was considered and turned down rather than never thought of.

The rule that matters most is the one in item 5 of the prompt: a DECIDED line records your answer, never the agent's. The letter is the one you picked, from a set the agent presented in an earlier turn. No set, no reply, no marker. This is the rule most likely to be argued away when a walk feels obvious, so treat an agent saying "I already know what you'd pick" as the signal it is about to break it.

Step 4: Finalize in Place

Build

When every question has a DECIDED line, ask the walking agent to finalize:

Every question is decided. Finalize the plan:
- Write the Recommendation section as one composed approach, not a
  list of answers. Cite the question behind each paragraph (Q3).
- Include the load-bearing rejections: any alternative an implementer
  could plausibly build without noticing it was considered and turned
  down.
- Leave Open Questions and every DECIDED line exactly where they are.
- Fill Verification, or leave its comment if there is nothing to check
  on the running thing.
- Set status: final.
Edit in place. Do not rewrite the file.

A finished recommendation reads something like this:

## Recommendation

Failed webhook deliveries are retried inside the sender, with
backoff, up to five attempts, each counted on the delivery row (Q2,
Q3). When the last attempt fails, the account owner is alerted (Q1)
and the payload is kept for thirty days (Q4).

Do not add retries to the job runner. It looks like the simpler
change, and it would repeat jobs that are not safe to repeat (Q3).

The second paragraph is a fence. Without it, an implementer who finds the runner is the obvious place for retries will put them there, and nothing in the code will say that was the option you rejected.

Why "in place" is in every prompt

Our own plans carry their planning rules inside the file, in a block the finalizer deletes. One finalization script did that by splitting the document on a section name and reassembling the pieces. The rules mentioned that section by name, so the split landed inside the rules instead of on the section. Everything after the split point was dropped: the questions and every decision marker under them. The script reported nothing.

Any document that talks about its own structure has this trap. So the rule is: change the lines you mean to change. If you do keep scaffolding in the plan, wrap it in a marker pair that appears exactly once, and cut on the markers alone:

grep -c 'BEGIN: planning-rules' plans/webhook-retries.md   # must print 1
grep -c 'END: planning-rules' plans/webhook-retries.md     # must print 1
sed -i '/<!-- BEGIN: planning-rules -->/,/<!-- END: planning-rules -->/d' \
  plans/webhook-retries.md
git diff --stat plans/webhook-retries.md

Commit the plan before finalizing and look at the after. Losing a section shows up as a large deletion you didn't ask for.

Reading the Record

Because decisions live in one file in one format, you can count them with grep. Anchor the pattern to the start of the line: your prompts quote the marker format as an example, and an unanchored search counts those too.

# questions in the plan
grep -E '^\*\*Q[0-9]+ — ' plans/webhook-retries.md | grep -vc 'DECIDED ('
# questions decided
grep -cE '^\*\*Q[0-9]+ — DECIDED \(' plans/webhook-retries.md

Be clear about what that count is. "4/4 decided" is an extension: it can be true of a plan where an agent wrote every line. The count is useful only because the rule behind it (your letter, your words, rejected options attached) is what gives it meaning. Keep the rule and the count is worth reading. Drop it and you have a well-formatted guess.

Try it yourself

Pick one real change you've been putting off because it has a decision in it. Then:

  1. Save the template and the two prompts from this tutorial in plans/ and commit them.
  2. Copy the template and write the Problem section yourself: three or more paragraphs, including what you don't want.
  3. Run the prep pass. Check the result: does every question ask one thing? Is at least one tagged [change]? Is there any "we could" hiding in a background? Is anything you know to be uncertain missing an Unverified mark?
  4. Close that session. Start a fresh one with the walk prompt and answer each question in your own words.
  5. While walking, note how many prep claims the walker corrected and how many new questions appeared.
  6. Finalize, then read only the Recommendation and the DECIDED lines. Could someone who wasn't there build what you meant from them, and avoid what you rejected?

If the walker corrected nothing and raised nothing, look again at whether it was really verifying, or reading the prep back to you.

When a Person Has to Walk

Not every question needs you. Questions about machinery (where a check lives, which store owns a record) can be argued between two agents, which we cover in Is a Debate Better Than One Strong Model? Questions of taste (who a page is for, what a label should say, the tone of a message) need the person whose taste it is. An agent answering those returns a well-formed verdict that is a guess about you, and a check that only asks "did it reject an alternative on stated grounds?" will pass it. We walk plans made of those questions ourselves, every time.

If an agent ever does answer, the record has to say so. A bare DECIDED line means a person answered. A decision made by an agent carries that fact in the line itself, so the two can never be confused later.

The value of a decision record is that a later reader can tell what was decided and on whose authority.

Key Takeaways

  • A plan with every question answered is the extension of "planned". The is a record of what you chose and what you turned down.
  • Prep finds decisions and doesn't make them: facts and questions, an inline Unverified mark on anything unchecked, one question per answer, at least one question that could change something.
  • Walk with a fresh agent. It re-checks the prep, and in our runs that has changed decisions and surfaced questions the prep missed.
  • One question per turn: the tension, lettered options with their costs, and a leaning that says what it gives up.
  • A DECIDED line holds your letter and the rejected options with their reasons, never the agent's own answer.
  • Finalize into a Recommendation that carries its fences, keep the decision trail, and edit the file in place. Never rebuild it by splitting on text.

Next Steps

← Back to all tutorials