Reviews That Can't Quietly Drop Findings

Concepts in Practice Intermediate ~35 min

What you'll learn

The problem: the judgment word

Ask an to review some code and list its findings, and you have asked it to make two decisions at once: what it sees, and which of those things are worth reporting. The second decision is where things go missing. Reviewer agents are trained to avoid false positives, so they quietly filter anything ambiguous. The cautious reviewer drops what it isn't sure of. The tired one drops whatever sits next to something an earlier category already covered. None of this shows in the output. A dropped concern leaves no trace.

We saw this directly. Two frontend reviews of the same project, months apart: the earlier one surfaced 14 findings, each with a file and line. The later one, run after the review instructions had been refactored, surfaced 7, with several "No findings in this category" lines in places where the code clearly had problems. Same project, half the findings, and nothing in the second report that told you so.

In this site's terms, the is "a review happened": a report exists, it has sections, it reads well. The is that every real concern reached whoever acts on the review. Prompting for findings optimizes the first and loses the second, because the filter that drops observations sits exactly where nobody can see it.

The fix that worked for us was to take the judgment word out. The reviewer doesn't decide what counts as a finding. It anchors everything, and a later step decides what to do. We call it density, not judgment.

Step 1: Four markers, numbered across the report

The marker vocabulary

The report is ordinary exploratory prose, one section per category, with inline markers at the point each one applies:

Each type is numbered independently and sequentially across the whole report: [R1] to [R47], no resets at section boundaries, no skipped numbers. That makes the report countable, which matters in Step 4.

Three rules give the markers their force:

  1. Every paragraph carries at least one [O] or [W]. A paragraph with only [R] anchors is narration. It describes the code without reviewing it, and the next step will drop it, because checklist entries only exist where [O], [W] or [V] markers do.
  2. If the prose makes a claim about the code, it carries a marker. There is no "is this worth marking?" gate. Vague generalizations such as "15 other hooks do the same" can't carry a marker, so the reviewer either enumerates each one with its own [R] or drops the claim. The rule forces enumeration, and that is the point.
  3. Keep [O] and [W] apart. "Strict mode is enabled" filed as [O] sends the implementer to investigate a non-problem. "This component leaks memory" filed as [W] turns a real concern into a don't-touch item. Choosing the type is part of the reviewer's job.

Here is what a section looks like:

## Category 3: Error Handling

[O4] Three handlers catch and log without re-raising, so a failed
write returns 200 to the caller:
- [R11] `api/orders.py:88`
- [R12] `api/orders.py:131`
- [R13] `api/refunds.py:42`

[W5] Every outbound HTTP call goes through [R14] `lib/http.py:fetch`,
which sets a timeout and retries only idempotent methods. Preserve
this when touching the handlers above.

Just as important is what the report leaves out. It has no ### Finding N blocks with File / Issue / Fix fields, no severity tiers and no suggested work order, and the reviewer doesn't prescribe fixes. Each of those is the judgment call coming back in. The report is exploratory: it points the implementer at things, and the implementer decides what each one needs.

Step 2: Density isn't depth

Two rules for what markers can't see

Markers solved the dropping problem and exposed the next one. The first marker-era review of that same frontend produced 259 markers, about 18 times as many anchors as the findings-era report. It also missed three things the older, sparser report had caught: no Web Vitals monitoring, no bundle chunking for three.js, and a raw polling timer that bypassed the project's own visibility-aware polling helper. All three were things the project didn't have. Density makes the reviewer enumerate what it reads. It doesn't make the reviewer look for what isn't there.

The depth check. Before closing a category, the reviewer asks: what are the three to five most commonly overlooked concerns in this domain, and did I investigate each one? For performance that might be bundle analysis, image optimization and render frequency; for error handling, external monitoring, silent catches and error boundaries. Each gets a marker: [W] if the project has it, [O] if it doesn't, and [O] with a one-sentence reason if it doesn't apply. Then the reviewer counts the [O] and [W] markers it wrote for the category. Zero means the category was described without being reviewed.

Confirm an absence before claiming it. The depth check creates a new risk: the reviewer now writes "no X here" markers, and some will be wrong. In one run, a backend reviewer marked "no global exception handler" when one existed and was registered, and the same report cited it elsewhere as a practice to preserve. The frontend reviewer on the same project marked "no skip link", and that existed too. A false absence is the most expensive error a review makes, because it sends the implementer off to build something that already exists. So the rule is to grep for it, open the file that would contain it, or check the config key before asserting it's missing. If the check finds it, it's a [W]. If the check can't settle it, the report says the absence is unverified. After this rule shipped, every re-run on that project found the handler, and false absences on that known error went to zero.

Step 3: Write the reviewer prompt

Hands-on

The reviewer is a plain markdown prompt, so any agent can run it. Create prompts/marker-reviewer.md:

You are a reviewer. Your job is to point the implementing agent at
things it should investigate, not to decide which observations are
worth acting on. Density, not judgment.

## Report header
Start the report with a `## Review Summary` heading, then this block
verbatim:

> **Purpose**: This report is generated by a reviewer agent. It is an
> exploratory report, not a prescriptive plan. Every marker below is a
> reference point the implementing agent must investigate, acknowledge,
> and respond to — there are no optional items, no "nice-to-haves," no
> "fix later" items. The implementing agent decides what action each
> marker warrants, but no marker is closed until it has been
> investigated and a decision recorded.

## Categories (one `## Category N: Title` section each, in this order)
1. Error Handling
2. Input Validation
3. Logging & Monitoring
4. Tests
Write each section to the report file as soon as you finish it.

## Markers (each type numbered 1..N across the whole report)
- [R<n>] every concrete location, one per distinct location.
- [O<n>] every concern to investigate.
- [W<n>] every concrete good practice to preserve.
- [V<n>] every runnable verification step, project-relative paths.

## Rules
- Every paragraph carries at least one [O] or [W]. [R]-only prose is
  narration and will be dropped.
- Any claim about the code carries a marker. No "worth marking?" gate.
- Before any marker that claims something is absent, confirm it with
  a grep or a direct read. Found it: that's a [W]. Can't tell: say the
  absence is unverified.
- Before closing a category, list the 3-5 most commonly overlooked
  concerns in that domain and mark each. Then count the [O] and [W]
  markers in the section. Zero means you described it without
  reviewing it: go back.
- A category with nothing to flag gets one sentence naming what you
  checked: [W] if there's a practice to preserve, [O] with a reason if
  the subject is absent.
- No severity tiers, no Fix fields, no priority order, no prescribed
  remedies.

## Finish
End with `## Verification Commands`: at least lint, type check and
build (or the project's equivalents). Put each [V<n>] in the prose
line above its command, never inside the code block. Then run
`scripts/check-markers.sh reviews/report.md` and fix any failure with
targeted edits. Don't regenerate the report. State the count for all
four types in your reply, including zeros.

Two notes on that file. Nothing in a prompt stops the reviewer from changing the tree, so run it against a separate checkout or a (see Put the Walls in Structure, Not Prompts). In Claude Code, for example, the same prompt can also live in .claude/agents/ as a whose tools list leaves out the editing tools, but Bash and Write can still change files, so the separate checkout is the guarantee. And "write each section as soon as you finish it" earns its place. In one before-and-after pair on the same target, the reviewer that wrote per category produced 61 preserve markers where the one that wrote everything at the end produced 20, and 18 verification steps against 8. Written at the end, early categories had faded from context and got summarized rather than anchored.

Step 4: The marker count, before anyone says "done"

A runnable self-check

The simplest version of the check fits on one line:

for t in R O W V; do grep -oE "\[${t}[0-9]+\]" report.md | sort -u | wc -l; done

Every count must be at least 1, and each type's numbers must run 1, 2, 3 up to N with no gaps. Both halves are there because of failures we saw.

A walk by a smaller local model covered all 16 categories with plenty of [R] anchors and produced zero [O] markers. Its closing section asserted that "all observations carry [O<n>] markers", copied from the rules it never applied. Our structural scored it near clean, since every marker present was valid. After the paragraph rule and the per-category count were added, the next walk by the same model produced healthy [O] density and put the marker-count grep in its own verification section. Countable rules are the ones smaller models actually follow.

The second failure is subtler. A different report had every category in order, 19 [R], 24 [O] and 22 [W], and a Verification Commands section with four well-chosen commands. None of them carried a [V] marker. The report's own self-check certified "all markers numbered sequentially" by listing the ranges it had used. The lint scored it 0.9909 with perfect sequence integrity, because a type that was never written has no gaps to find. Once we added a presence check, the same report scored 0.9307 with a named "no [V] markers" issue. Across 124 reports, 102 (82%) carried [V], so the omission is random, not something a better example would fix. The answer is a check that fires when it happens.

Save this as scripts/check-markers.sh. It counts only prose and skips fenced blocks, because a verification block that quotes [O3] as text would otherwise invent a marker, and a [V] hidden inside a fence anchors nothing downstream:

#!/usr/bin/env bash
# Usage: scripts/check-markers.sh REPORT
# Fails if any marker type is missing or numbered with a gap.
report="$1"; fail=0
prose=$(awk '/^```/ {f = !f; next} !f' "$report")

for t in R O W V; do
  nums=$(grep -oE "\[${t}[0-9]+\]" <<< "$prose" | tr -dc '0-9\n' | sort -n -u)
  count=$(grep -c . <<< "$nums")
  max=$(tail -n 1 <<< "$nums")
  echo "[$t] count: $count"
  if [ "$count" -eq 0 ]; then
    echo "  FAIL: no [$t] markers anywhere"; fail=1
  elif [ "$count" -ne "$max" ]; then
    echo "  FAIL: $count distinct numbers but highest is $max (gap)"; fail=1
  fi
done

# Per-category [O]/[W] count: a zero is a section described, not reviewed.
awk '/^```/ {f = !f; next} f {next}
     /^## / {if (h) print c "\t" h; h = $0; c = 0; next}
     {c += gsub(/\[[OW][0-9]+\]/, "&")}
     END {if (h) print c "\t" h}' "$report"

exit $fail

Make it executable with chmod +x. A zero count for [V] while the section visibly has markers means they are inside the fence. Move each one into the prose line above its command. The script can't tell whether a number was reused for two different observations, so skim for that by eye.

Step 5: Convert the report into a checklist

A separate formatting step

The review is only half the contract. The other half is that nothing in it can be skipped. A separate formatting step copies the report unchanged and appends one entry per [W], [O] and [V] marker, in that order (what not to break comes first). [R] markers get no entries. They are site anchors under an [O], and the parent entry covers the decision. Each entry has three lines:

[O4]
Addressed appropriately: (yes/no)
Code changes? (yes/no)
Address summary:

A marker is closed only when "Addressed appropriately" is yes and the summary is filled in. "Code changes?" can honestly be no: investigated and decided against, or preserved as is. The summary has no length cap, so "fixed 60 of 64 sites, left 4 intentional" fits.

This step involves no judgment, so it doesn't need a model. Save it as scripts/to-checklist.sh:

#!/usr/bin/env bash
# Usage: scripts/to-checklist.sh REPORT > PLAN
report="$1"
cat "$report"
printf '\n---\n\n## Implementation Checklist\n\n'
printf '> One entry per [W], [O], [V] marker. Closed only when\n'
printf '> Addressed appropriately is yes AND the summary is filled in.\n\n'

prose=$(awk '/^```/ {f = !f; next} !f' "$report")
for t in W O V; do
  for n in $(grep -oE "\[${t}[0-9]+\]" <<< "$prose" | tr -dc '0-9\n' | sort -n -u); do
    printf '[%s%s]\nAddressed appropriately: (yes/no)\nCode changes? (yes/no)\nAddress summary:\n\n' "$t" "$n"
  done
done

Two practices stop the implementer from emptying the checklist the way the reviewer used to empty the report. First, decisions you'd give the same answer to every time ("always replace internal hostnames with placeholders") go into the implementer's standing instructions as set rules, each with what to do and what not to do. Otherwise every run punts them back to you as "operator decision" recommendations. Second, on refactor work, add a no-punt clause: don't defer a marker because it's effortful and then call it "risky". A specific, reasoned safety objection is fine. Risk as a blanket excuse is not. Things that genuinely need judgment at run time stay out of the standing rules and come back honestly as Code changes? no, Addressed appropriately? no, which is where you want to see them.

Try it yourself

Pick a target where a naive reviewer would score clean. The target is chosen to be discriminating. A repo that obviously lacks the thing tests nothing, and neither does one that obviously has it. One of ours had a fully pinned lockfile with no hashes that the setup path never installed from, so "is there a lockfile?" passed and the real defect was a level down.

  1. Read the code yourself and write the expected markers into a file before running anything. Note which ones must come back [W], not [O]. A false absence on something the project has is the main risk, and you can't see it in a report you read for the first time afterwards.
  2. Run the reviewer. The examples use Claude Code and opencode; any agent with a non-interactive mode (Codex CLI, Gemini CLI, aider) works:
    # The only line that names a specific agent. Swap in yours.
    run_agent() { claude -p "$1"; }          # Claude Code
    # run_agent() { opencode run "$1"; }     # opencode: any provider, add -m provider/model
    
    run_agent "$(cat prompts/marker-reviewer.md)
    
    Review src/ and write reviews/report.md."
  3. Run scripts/check-markers.sh reviews/report.md. Every type present, no gaps, no category at zero.
  4. Grade the report against your table by hand, then read the rest for anything you missed.
  5. Run scripts/to-checklist.sh reviews/report.md > reviews/plan.md and count the entries against the [W]+[O]+[V] totals.

Writing the table first is what makes a fluent wrong finding visible. In one validation run, two findings that looked correct turned out wrong only when graded against a table written before launch: a missing control reported as "a mitigating factor", and a lockfile described as "gitignored or unused" by a walk that never opened .gitignore. Both sat in a report that passed every structural check.

What the count can't tell you

Everything in Step 4 measures form. A report can be anchored, correctly typed, contiguously numbered, properly sectioned and wrong. In one comparison of two reports on the same project, the one that scored a perfect 1.0 carried an inverted finding: it recommended making a type checker fail when absent, which the build target already did. The lower-scoring report had the defect right, root cause included.

A clean 1.0 is necessary, not sufficient.

That is the drift pattern again, one level up. The marker count is a better stand-in than "a report exists", and it is still a stand-in. It catches a missing type or a dropped section, and it can't catch a confident false claim. Keep the count, and keep spot-checking a few markers against the code.

Key Takeaways

  • Asking for "findings" builds a hidden filter into the reviewer. Anchor every claim with [R], [O], [W] or [V] instead: density, not judgment.
  • Every paragraph needs an [O] or [W]. A category with zero of them was described, not reviewed.
  • Density doesn't find what's missing. Use a per-category depth check for that, and confirm any absence before asserting it.
  • Run a marker count before calling a review done. Check that every type is present, not just that the numbers are contiguous, because an empty type passes a gap check.
  • Turn every [W], [O] and [V] into a checklist entry that closes only on a recorded decision.
  • Validate against a discriminating target with the expected markers written down first. A structural pass is necessary, not sufficient.

Next Steps

← Back to all tutorials