Most of what we write about here concerns models: an that satisfies the and misses the point, a reward optimized at the expense of what it was meant to encourage. It's comfortable to treat that as a machine problem. It isn't. The same drift shows up in the writing people do every day, for the same reason: we write for whatever will score us.

Our tutorial on extensional drift defines it as optimizing for something measurable and quietly losing the thing you actually cared about. Here's what that looks like when the thing being optimized is your own prose.

Writing for the scorer

The keyword filter. Job applications are screened by software before a person reads them, so people learn to write CVs for the software: the posting's exact phrases, repeated. The CV that passes the filter describes the job posting more than the applicant. Its (matched keywords) goes up while its (who this person is and what they did) goes down.

The rubric. A performance review graded on “impact,” “collaboration” and “ownership” produces self-assessments with a paragraph under each heading, whether or not anything happened under that heading. The form gets filled, and the story of the year gets chopped to fit the form.

The readability score. Readability formulas count syllables and sentence lengths. Write to raise the score and you get short sentences and short words. That's often good advice, and it stops being good the moment a precise long word is replaced by a vague short one. The formula can't tell “clear” from “simplified until wrong.”

The skim. Content written to be skimmed (bold lead-ins, bulleted takeaways, a summary before the summary) scores well on time-on-page and scroll depth. It also trains readers not to read, which rewards writers for making every text skimmable, which is how an argument turns into a list of assertions.

The newest scorer: “does this sound like AI?”

In our post on jargon we noted that words models overuse (“delve,” the em dash, “it's not just X, it's Y”) have become tells. That has created a new scorer, part software and part the reader's suspicion, and people have started writing for it too.

So writers remove the em dash they've used for twenty years. They swap a precise word for a clumsier one because the precise one is on somebody's list. Some add a deliberate rough edge so the text reads as human. None of this is about saying the thing better. It's about the text passing a test that has nothing to do with whether it's true or clear.

It's the same shape as every other case: a surface feature that once correlated with something (“written by a person”) becomes a target, and the correlation breaks. A text with no em dashes isn't more human. It's just been optimized against a detector.

What gets lost

In each case the loss is the same: the reason for writing is replaced by the reason the writing will be scored. A CV exists to tell someone who you are. A review exists so that someone understands your year. An essay exists to make an argument. Each scorer measures something that usually comes along with that purpose, and each one, once optimized, can be satisfied without it.

The cost is higher than it looks, because writing is also how many of us think. Drafting for a rubric doesn't only change the text; it changes what you notice. If you write every week for a scorer that rewards confident bullet points, you get better at producing confident bullet points and worse at the long, uncertain paragraph in which you find out what you actually believe.

Catching it

The tutorial's three questions work on prose as well as on metrics:

  1. What was this writing for? Not the format, not the audience's checklist, but the reason it exists.
  2. Could it score better while doing that job worse? If yes, you're writing for the scorer.
  3. What would show the gap? Usually a person who needs the text to do its job. Ask them what they took from it.

A few habits help. Draft for the reader first and fit the form second, so the form is a container rather than a mold. Keep the precise word, even if it's long or unfashionable, and explain it if you must. And when you catch yourself changing a sentence to avoid sounding like something, ask whether the new sentence says more or only says it differently.

The same problem, on our side

It's tempting to describe extensional drift as something AI does and we correct. The more honest picture is that it's what any optimizer does under a measure, and people are optimizers too. The models learned their habits from our text, including text written for scorers. If we want them to write for meaning rather than for metrics, it helps to keep doing that ourselves.