Unit tests, integration tests, TDD, and mocks: what I’d keep in 2026 and what I’d drop

Tobias Werner

Tobias Werner

September 18, 2026

Unit tests, integration tests, TDD, and mocks: what I’d keep in 2026 and what I’d drop

I used to mock everything that was not a pure function. The suite was fast. The suite was also a fantasy of my own interfaces. Production still found the Hibernate flush order, the Stripe idempotency key I had stubbed into a happy path, and the clock that was not Instant.now() in the way I had imagined.

In 2026 I keep a smaller mix. Agents made it cheap to generate more tests. They did not make it cheap to generate the ones that fail when the system is wrong. Here is what I still write, what I will not, and why TDD is a tool I pick up for one kind of problem and put down for another.

What I keep

Integration tests against real Postgres. Testcontainers, a compose service, a CI job that is allowed to be slower than a unit suite. I want the migration, the constraint, and the query in the same verdict. If you mock the repository, you are testing your typing. I have shipped that typing. It compiled. It double-paid.

A few unit tests on rules that are actually rules. Tax rounding. A state machine. A parser. Things I can name without a network. I do not unit-test a Spring controller that only calls a service. That test is a transcript of the mock setup.

Contract tests at the HTTP edge when we have more than one consumer. Pact or a recorded OpenAPI snapshot. I want to know when I broke the iOS app, not when I broke my own mock of the iOS app.

A smoke path in staging that hits the real payment provider in test mode, or a recorded sandbox. One path. The one that loses money if it is wrong. I will not replace this with 400 unit tests of a PaymentClient wrapper.

Characterization tests on legacy. Golden files. A recorded response. I do not TDD a 2014 module. I pin it, then I change it. Agents are good at proposing pins. I still read the pin.

What I drop

Mocking every collaborator in a service test. Mockito forests that assert verify(repo).save(any()) taught me nothing I did not already believe. I keep mocks for the things I do not own — a bank, a flaky vendor — and I make those mocks thin and obvious. I do not mock my own database.

Coverage as a gate above a vanity number. 80% that is all getters will beat 40% that includes the refund path. I have managed both. I will take the 40% and a comment on the dangerous files. I will not take a JaCoCo argument in stand-up.

TDD as a religion for CRUD. Red-green-refactor on a generated Laravel or Spring resource is theater. I TDD when the next three cases are not obvious and I need the test to tell me the API. I do not TDD a mapper.

End-to-end suites that own the whole product. Playwright for the two journeys that are the business. Not for every button. E2E that takes 40 minutes is a suite people skip. A skipped suite is not a suite.

Tests that assert implementation details the model loves to write. “Expect this private method to be called.” Delete them. They make refactors expensive and they do not catch the bug that escaped.

A CI laptop showing a failing integration test next to a Postgres container

TDD, when I still do it

I still write the test first when I am designing a rule I cannot hold in my head: a complicated eligibility function, a parser, a retry policy. The test is the spec. I have used that with agents — paste the cases, ask for the function, refuse the cases it invents. TDD plus a model is a decent pair if the human owns the cases.

I do not write the test first when I am exploring a dirty integration. I spike, I learn, I pin. Pretending that was TDD is how we got 12 skipped tests named TODO_real_api.

I do not pair-TDD a whole feature in 2026 as a team ceremony. I pair on the invariant test. Then people implement. Then we add the ugly case we found. The ceremony was never the value. The invariant was.

The mix that caught production bugs

On the payouts service, the tests that paid rent were:

  • An integration test that started a real Postgres, applied Flyway, and tried to insert a second payout with the same idempotency key. The unique constraint was the feature. The mock would have passed.
  • A test that replayed a Stripe webhook twice. The second time had to be a no-op. We had mocked Stripe as always-new. Production was not always-new.
  • A time test that used a fixed clock and a batch that “ran at midnight” across DST. The mock clock in a unit test would have been enough here — and was. This is the kind of unit test I keep.
  • A Playwright journey that a finance person could recognize: lock, export, unlock. It failed when we changed a column order. Finance’s VLOOKUP was the user. The unit tests were green.

The tests that did not pay rent were 200 Mockito tests of a service that wrapped the repository. We deleted a third of them in a week and the defect rate did not move. The suite got faster. People started running it.

Printed test cases on a desk with a red pen circling an idempotency scenario

Agents change the volume, not the mix

A model will happily write a test for every public method. I treat that as a draft of a checklist, not as a suite. I ask it for the case I am afraid of, then I make sure that case hits a real database or a real contract. If the generated test only hits a mock, I delete it or I promote it.

I have started adding a CI lint: if a new test file imports Mockito and not Testcontainers or a real HTTP, print a warning. Not a ban. A warning. People still write unit tests. They write fewer decorative ones when the warning is in the log.

Snapshot tests from agents are a special hazard. They pin a hallucination. I allow snapshots for HTML I own and for characterization. I do not allow them for JSON I have not read.

How I structure the repo so the mix survives

I put unit tests next to the rule they pin. I put integration tests in a module that has a real datasource config and a rule: no Mockito in that module. I put e2e in a folder that CI runs on merge to main, not on every push, unless the change touched the two journeys. That split is more important than the framework. JUnit, pytest, RSpec, Pest — I do not care. I care that the dangerous tests cannot silently become mocks because someone copied a template.

I name tests after the failure, not after the method. should_reject_second_payout_with_same_key tells a future reader what to fear. testSave tells them nothing. Agents default to testSave. I rename. If I do not have time to rename, I do not have time to keep the test.

Flakes get a budget. Three flakes and the test is quarantined or deleted. A flake that stays is a permission to ignore the suite. I have managed teams that learned to ignore. It takes a year to unlearn. Deleting the flake is cheaper than a culture of shrug.

What I tell a team that is drowning in tests

List the last five production bugs. For each, ask which test should have caught it and why it did not. You will find mocks, missing time, missing uniqueness, or a journey nobody automated. Write those five. Delete twenty that never failed except when you renamed a method.

Keep the suite short enough that people run it. Keep the dangerous path on a real database. Keep TDD for the rules. Drop the forest of mocks. Coverage will look worse. Production will look better. That is the 2026 mix I will defend to a lead who still wants 90% and a green badge.

If you only change one thing this quarter, stop mocking your own data store. The rest of the religion can wait. The store is where the bugs I still remember were hiding, and no generated unit test was going to walk in there unless I made it.

I would rather have forty tests that embarrass me when they fail than four hundred that compliment my mocks. Agents can write the four hundred before lunch. They cannot pick the forty. That pick is still the job. I keep the pick. I drop the compliment. If a lead asks me for a number, I give them time-to-detect on the last incident, not a coverage percentage. The percentage is a costume. The incident is the curriculum. I would rather the suite be the curriculum. That is what I keep in 2026. That is what I drop when the costume starts to win. Run the suite. Read the failure. Change the system. That loop is older than TDD slogans and it still works when the slogans do not.

More articles for you