The looked good. That was the problem.
We were reviewing a batch of staged changes in one of our services,
the one that runs multi-agent deliberation sessions. Most of it was
the work we had asked for. Mixed in with it was something we hadn't
asked for: a 16-line function called
find_sessions_by_prefix, a 22-line change to the
session lookup endpoint (try an exact match, fall back to a prefix
match, return 409 if the prefix is ambiguous), and 75
lines of tests covering all of it.
All three pieces were well written. The function was small and named for what it did. The endpoint change handled the ambiguous case properly instead of guessing. The tests covered the exact match, the unique prefix, the ambiguous prefix and the miss. If you were grading it as code, you would pass it.
We took it out anyway, because it had no user.
Where the shape came from
We recognized it at once. The day before, we had added git-style
short-ID matching to a different service, our task dispatcher. That
change had a user. Someone monitoring a task with
curl had typed a shortened task ID, copied from a log
line, and got back an empty 200: no task, no error,
nothing to say the ID was simply too short. The fix closed a real
gap in a real interaction. A person typed something reasonable and
the system answered unhelpfully.
The working on the deliberation service had that change in view. It saw a similar lookup endpoint, noticed the missing prefix match, and ported the feature across. Same pattern, same care, same test discipline.
Who is this for?
So we asked the question the diff couldn't answer: who types a session ID into this endpoint?
- Not people. Session IDs reach humans through list endpoints and task records, which return full IDs. Nobody reads one off a log and types it in.
- Not agents. Agents receive session IDs as full strings in their tool context and pass them back unchanged. An agent does not abbreviate an identifier it is holding.
- Not scripts. Every script that talks to the service gets IDs from the service, in full.
The answer was nobody. The feature was built for an interaction that never happens: someone typing a partial session ID into an API that only ever hands out whole ones.
That is what made it scope creep, not the quality of the code. The agent wasn't solving a problem. It was completing a symmetry: the dispatcher has short-ID matching, so this service should too. That reads like sound engineering until you notice where the reasoning started. It started from what we have (a pattern) instead of from what someone needs (a failure).
Supply-driven: “we have this pattern, so it belongs here too.” Demand-driven: “this person hit this failure, and here is what fixes it.” Only one of them has a user.
Why the quality made it worse
A sloppy version of this change would have been caught in seconds. An untested helper, a hacked-in fallback, a missing error case: reviewers are trained to notice those. The polished version gets waved through, because every signal a reviewer checks is green. The code is clean, the tests pass, the naming is consistent, and the change even mirrors an existing feature, which makes it look deliberate.
Seventy-five lines of passing tests are a strong signal that something is finished and correct. They say nothing about whether it should exist. A test suite can tell you the feature works. It can't tell you that it tests a feature nobody will call.
In the terms this site keeps returning to, every check passed: the code compiles, the tests run, the behavior matches the specification the agent wrote for itself. The check (what is this for?) was never run, because nothing in the pipeline asks it.
Why agents are prone to this
This isn't carelessness, and it isn't a capability gap. It follows from how an agent sees a codebase.
Agents work from code, not from users. An agent with several repositories in context sees functions, endpoints and repeated shapes. Users show up as abstractions at best. So “add the missing function” is a far more salient move than “ask who needs it.”
Consistency is a real virtue, applied at the wrong scale. Within one service, similar operations should behave alike, and duplicated logic should usually be shared. Agents learn that heuristic well. Across services with different users, though, mechanical consistency manufactures features nobody asked for. Consistency is a local virtue and a global vice.
Nothing in the loop is designed to say no. The agent writes the feature, writes the tests, runs the tests, and reports success. Each step confirms the one before. None of them checks whether the whole thing should exist.
The test
The rule that would have caught this is short enough to put in an agent's instructions and in every review checklist. Before a feature is added, name three things:
- The user. A person, an agent or a script. Be specific about which.
- The interaction. What that user actually does that the feature serves.
- The failure. What goes wrong for them today without it.
If any answer is vague, or reaches for “consistency,” “symmetry” or “nice to have,” the feature goes.
Applied to our two cases:
| Dispatcher | Deliberation service | |
|---|---|---|
| User | A person monitoring tasks with curl | — |
| Interaction | Typing a short ID copied from a log | — |
| Failure | An empty 200 instead of the task | — |
| Verdict | Build it | Drop it |
Putting it into practice
- Make the question a field, not a vibe. Any change with agent authorship names its user, interaction and failure in the description. A blank is a finding.
- Tell the agent the difference. Say in its instructions that every new feature must trace to a concrete demand (a filed bug, a documented need, a request), not to a pattern it noticed elsewhere.
- Treat porting as a claim that needs evidence. “Service A has it” is not a reason for service B. A port needs its own user in the destination.
- Look hardest at the clean additions. Unrequested, well-tested, internally consistent changes are exactly where this hides, because their quality is the camouflage.
Code that solves nothing
Agent scope creep isn't a problem of bad code. It's a problem of good code that solves nothing. Reviewing the diff harder won't catch it, since the diff is fine. What catches it is asking a question the diff can't answer.
Who is this for? If you can't name them, you don't need the feature.