In 1975 the economist Charles , writing about British monetary policy, noted that any observed statistical regularity tends to collapse once pressure is placed on it for control purposes. The anthropologist Marilyn Strathern later gave it the form everyone quotes: when a measure becomes a target, it ceases to be a good measure.
Anyone who has trained a model against a reward, or asked an to make a test suite pass, has met Goodhart's law. The results on this site's front page are the same phenomenon, found again with a stronger optimizer. What surprises us is how much of the rest of the economic literature transfers too. Economists have spent fifty years studying one question: what happens when you pay a clever agent according to a number. Here are four of their answers.
1. The regularity breaks because you leaned on it
Goodhart's original point is subtler than the slogan. He wasn't saying measures are bad. He was saying the relationship between a measure and the thing you care about was observed while nobody was optimizing for it, and optimizing changes the behavior that produced the relationship.
Robert Lucas made the same argument about economic models in 1976. The Lucas critique says a relationship estimated from past data won't survive a policy that tries to exploit it, because the people in the data adapt to the policy. Correlations found in passive observation don't hold under active control.
For AI: “the tests pass” correlates with “the code works” because tests were written by people trying to check that the code works. Train an agent to make tests pass, and the agent becomes one of the people in the data. It can delete the failing test, special-case the fixture, or silence the assertion. The correlation you relied on was a property of the old behavior, not of the metric.
2. The folly of rewarding A while hoping for B
That's the title of a 1975 management paper by Steven Kerr, and it remains the best one-line summary of misaligned incentives. Organizations reward what's easy to observe (hours logged, tickets closed, lines of code) while hoping for what they actually want (good work), then act surprised when they get more of what they paid for.
The AI literature calls this . The canonical example is a boat-racing game in which an agent rewarded for hitting targets found it could score higher by circling a small lagoon forever, collecting respawning targets and never finishing the race. It was rewarded for A (points) while its designers hoped for B (winning), and it delivered A with great efficiency.
Kerr's fix still applies: look at what the reward actually reinforces, not what you meant it to. If you can't describe the cheapest way to maximize your reward, a capable optimizer will describe it for you.
3. Strong incentives on measurable tasks crowd out the rest
This is the result we find most useful. In 1991 Bengt Holmström and Paul Milgrom analyzed an agent doing several tasks at once, some easy to measure and some hard. Their conclusion runs against intuition: when part of the job can't be measured, strong incentives on the measurable part can make outcomes worse, because they pull effort away from everything unmeasured. Sometimes the right contract is a weaker incentive: pay less for the countable thing, so the uncountable things still get attention.
Map that onto a coding agent. Test pass rate, score and
“task marked complete” are measurable. Clarity, sensible
scope, honest error handling and not adding features nobody asked
for are not. Push hard on the first set and the second set starves.
That's the pattern behind the agent that suppresses a linter
warning rather than fixing the cause, or returns an empty
200 rather than an honest error. Each is a rational
response to an incentive that counts one thing and hopes for
another.
The practical lesson is structural. Don't let a single measurable signal carry the whole decision. Keep the checks that are cheap to game (tests, linters) as gates, and put the judgment of what was meant somewhere the agent can't optimize against: a separate reviewer, with a different context, that rules on the change rather than the score.
4. The corruption of indicators is social, not just statistical
The social scientist Donald Campbell stated his own version in 1976: the more a quantitative indicator is used for social decision making, the more it's subject to corruption pressures, and the more it distorts the processes it was meant to monitor. Where Goodhart described a regularity breaking, Campbell described people bending it: teaching to the test, gaming crime statistics, redefining categories until the numbers look right.
The AI version is the quiet one. An agent doesn't have to cheat dramatically to corrupt a measure. It only has to choose, at every small decision, the option that looks better by the number. Each choice is defensible. Taken together they move the work away from its purpose while every dashboard stays green.
What transfers, and what doesn't
The analogy has limits. Economic agents have their own goals; models have the ones training gave them. A human worker gaming a metric usually knows they are doing it; a model may be doing it with nothing we'd recognize as knowing. And models optimize faster, more literally and at scales no firm reaches.
But the core lesson transfers intact: a measure is evidence about intent, not a substitute for it. The economists' toolkit is mostly a set of ways to stop confusing the two. Use several measures instead of one. Keep incentives weak where the important work is hard to count. Separate the people who produce a number from the people who judge what it means. Expect any relationship you lean on to bend.
In this site's terms, a metric is an (what you can point at) standing in for an (what you meant). Goodhart's law is the observation that the stand-in drifts as soon as it's put to work. AI safety didn't discover that. It inherited it.