How unit tests made the project harder to change — and the ones I’d delete first

Casey Holt

Casey Holt

September 18, 2026

How unit tests made the project harder to change — and the ones I’d delete first

I have been the person who required 90% line coverage and then could not rename a private method without a ceremony. The suite was green. The suite was also a fossil of every constructor we had ever had. Changing a feature meant changing forty tests that did not know the feature. They knew the mocks. I started deleting tests. The project got easier to change. That sentence still makes some people angry. I will stand on it.

Unit tests are not the villain. A certain kind of unit test is: the one that nails the implementation to the floor and calls that safety. Here is how I tell those from the ones I keep, and the first ones I delete when a suite is the reason we are slow.

The suite that made a rename a week

A Java service, Spring, good intentions. Every service class had a test that mocked every collaborator with Mockito, verified times(1) on each call, and asserted on the order of those calls. The tests were “units.” They did not hit the database. They also did not hit a behavior a user would name. When we changed a call from save to upsert, eighty tests failed. The product had not changed. The fossils had.

We spent more time updating verifies than we spent on the upsert. That is the smell. A test that fails when the user’s outcome is the same is a test I will delete or rewrite. Coverage went down. Merge speed went up. Incidents did not go up. We kept the tests that hit the money path through a test database. Those failed when the outcome was wrong. That is the suite I want.

Laptop showing a long list of test results

The ones I delete first

Tests that verify mock call graphs. verify(repo).save(any()) plus three more verifies is a coupling to a conversation between objects. I want a test that says the row exists, or the HTTP status is 422, or the email was enqueued. If I cannot say the outcome without naming the mock, I am testing my own wiring diagram.

Tests that snapshot a whole JSON blob of an internal DTO. One extra field and the snapshot screams. Snapshots of a public API can be a feature. Snapshots of a private mapper are a tax. I delete those and assert on the two fields that matter.

Tests for getters, setters, and generated code. I have seen coverage targets produce tests of Lombok. Delete them. If the coverage tool cannot ignore generated sources, fix the tool.

Tests that exist to hit a branch the compiler already forbids — or a branch that is a catch (Exception e) { log } I should have deleted instead. Testing a swallow is how you keep the swallow.

A second test that is a clone with a different string. Parameterize or delete. A suite of clones is how a change becomes a find-and-replace across a package.

Tests that require a 40-line mock setup to construct the subject. That setup is a design smell. I will fix the constructor or I will test through a narrower entry. I will not keep a novel of stubs.

The ones I keep, even when they are “slow”

A test that starts a real Postgres (Testcontainers) and runs the webhook that marks an invoice paid. Slow. I will pay. That test has saved me from a migration I thought was boring.

A test of money math with examples from finance. Fast. I keep it next to the function. No mocks.

A contract test at the HTTP edge for the two clients we actually have. I keep it. I do not keep a contract test for a client we imagined.

A test that the permission model rejects the cross-tenant id. I keep it even if it is an “integration.” Security outcomes are not a unit-test aesthetic.

I am not anti-unit. I am anti-unit-as-a-coverage-religion. A pure function with a table of cases is the best unit test I know. I write those on purpose. I extract the function so I can. That is the same habit as naming a reason to split: I split for a test I can trust, not for a folder of nouns.

Coverage as a hostage

I have used coverage gates in CI. I still like a report. I do not like a number that forbids deleting a bad test. The gate should not block a PR that removes a verify-times fossil. If your gate does that, the gate is the constraint — and not the good kind. I have made exceptions in the config. I have also lowered the number on purpose and written it down. Nobody died.

Mutation testing (Stryker, PIT) taught me more than line coverage. If I can delete a line and the suite stays green, I did not have a test. I had a decoration. I will run mutation on the money package. I will not run it on the whole monorepo to generate a score for a slide.

Engineer deleting code on a laptop in a cafe

How I delete without becoming a vandal

I delete on the path I am already changing. I do not schedule a “test cleanup quarter” that never ships. When a feature change touches a fossil, I replace the fossil with one outcome test or I delete it and rely on the existing integration. I say so in the PR. “Deleted 12 tests that verified mock order; added one Testcontainers case for the upsert.” Reviewers who care about safety have something to read. Reviewers who care about the number can look at the new case.

If I am scared, I keep the fossil one release and tag it @Deprecated or put it in a legacy source set that does not fail the build. Then I see if production disagrees. It almost never does. Fear is a poor test runner.

I do not delete tests I do not understand on a Friday before a holiday. I read them. If I cannot name the outcome, that is the reason to delete or to rewrite, not to keep.

Flakes, time, and the test that lies on CI

A test that sleeps to wait for a future is a test I rewrite or delete. I have inherited Thread.sleep(2000) as “stability.” It is a lottery. I use a clock I can freeze, or I await a condition with a timeout that fails loud. I do not keep a lottery to protect a number.

Order-dependent tests in a shared in-memory database are the same tarp. If I cannot run the file alone, I do not trust the file. I will isolate or I will delete. Parallel CI made this visible. The suite that only passed on a single runner was not a suite. It was a ritual.

I also delete tests that hit a real third party without a recording (VCR, a recorded HTTP). Those tests fail when their staging is sad. They train the team to ignore red. Ignoring red is how you lose the suite you kept. A recorded interaction is a unit of the wire. A live hope is not.

What I tell a team that is proud of 10,000 unit tests

I ask how long a rename takes and how often the suite fails on a flake. If the answers are “a day” and “often,” you do not have a safety net. You have a tarp. I would trade 2,000 fossils for 50 tests that hit the database on the paths that page. I would keep the table-driven tests of the pure core. I would stop hiring people to write verifies.

TDD is still useful as a way to design a function. It is a poor way to design a mock graph. If your red-green-refactor loop never leaves Mockito, you are designing a conversation between fakes. Design the outcome. Then decide if the test is a unit because the logic is a unit, not because a policy said units are cheaper. Sometimes the cheaper test is the one that looks expensive and tells the truth.

I still write unit tests. I write fewer of the other kind. The project got easier to change when I stopped treating every class as a thing that owed the suite a portrait of its internals. The ones I delete first are the portraits. The ones I keep are the verdicts. If you only remember one rule: a test must fail when the user-visible outcome is wrong, and it must stay quiet when you rename a private collaborator. Most of the suite I inherited failed the second half. Deleting it was the refactor. The code change was the easy part after that.

More articles for you