I recently watched the talk “Können wir Entwickler:innen-Produktivität messen?” (Can we measure developer productivity?) by Eberhard Wolff, whose work I have held in high regard for many years. One passage got me thinking in particular: his application of Goodhart's Law to code coverage. This article grew out of it.
Goodhart's Law
In 1975, the British economist Charles Goodhart formulated an observation about monetary policy that has long since gained significance far beyond economics. The phrasing in common use today comes from the anthropologist Marilyn Strathern:
When a measure becomes a target, it ceases to be a good measure.
As long as a metric is merely observed, it can be a useful indicator. As soon as reward or punishment depends on it, however, people begin to optimise the metric itself and no longer what it was actually supposed to measure. The metric thereby loses precisely the informative value for which it was made a target.
Code Coverage as a Target
Code coverage measures the proportion of the source code that is actually executed while the tests run. As an indicator, this is useful: code that is not executed by any test can contain bugs that no test will ever find. Low code coverage is therefore a clear warning sign.
As soon as a team has to reach 80%, 90%, or even 100% code coverage because a requirement, a dashboard, or even a bonus depends on it, the metric gets gamed. This is exactly what Goodhart's Law warns us about. None of this has to happen in bad faith. The team prefers to test the trivial parts of the code because they are easy to cover, while the complex parts, where the bugs are likely to hide, remain untested. Tests call methods without checking their results. Or they merely make sure that no exception is thrown.
In all of these cases, code coverage goes up. The tests still do not find any problems. A perfectly good metric has ceased to be a good metric because it was made a target.
Risky Tests
PHPUnit can detect some of these manipulations. A test that performs no assertion or expectation verifies nothing. It executes code, but it makes no claim about that code's behaviour. This is why PHPUnit marks such a test as risky by default and points this out in the test result.
One detail matters here in the context of Goodhart's Law: risky tests do not contribute to code coverage. PHPUnit discards the code coverage data collected while a risky test was running. A test that merely executes code without verifying anything therefore does not increase your code coverage.
In rare cases, a test without assertions is intentional, for instance when successful execution without an exception is in itself the relevant statement. Such a test can be marked with the #[DoesNotPerformAssertions] attribute. This marking is explicit, lives in the code, and is therefore visible to everyone during code review.
Mutation Testing
A test can contain assertions and still be worthless, for example when it verifies the wrong thing or is too weakly formulated. PHPUnit does not recognise such tests as risky. This is where mutation testing helps: a tool such as Infection makes controlled changes to the code that simulate typical programming mistakes and checks whether the tests notice these changes. Tests that do not “kill” a single mutant obviously test nothing that is relevant to the behaviour of the code.
In my article “Path Coverage or Mutation Testing?”, I described in detail how mutation testing works and why I use it in my daily work. Mutation testing is another technical solution for making sure that a test actually tests something.
A Human Problem
As valuable as these tools are, no technical measure can solve the problem that Goodhart's Law describes. Every metric that is made a target invites manipulation. This is as true for code coverage as it is for the mutation score. Anyone who has to reach a prescribed mutation score will find ways to reach it without the tests getting any better.
Goodhart's Law describes a human problem, or more precisely: an interpersonal one. It is about incentives, about trust, and about whether a team uses a metric as a tool for its own improvement or has it imposed as a target. A metric that a team observes on its own initiative in order to get better remains a good indicator. The same metric imposed from outside will be gamed.
Tools such as PHPUnit and Infection can detect the crudest forms of manipulation and render them ineffective. The decision to treat metrics as indicators rather than targets is not one they can make for us. That decision, we humans have to make ourselves.