Unit tests vs integration tests: I stopped treating it as a debate — here’s the mix I ship
Casey Holt
September 18, 2026
I stopped treating unit versus integration as a debate. The debate is a conference sport. The mix I ship is a few fast tests around pure rules and a smaller set of slower tests around the paths that page. I have deleted the fossils that made a rename a week. I have also kept a Testcontainers case that looked expensive and told the truth. The mix is the job. The labels are optional.
If you came from the piece about tests I delete first, this is the constructive half: what I keep, in what ratio, and when twenty integration tests beat two hundred units — and when that slogan is wrong.
What I mean by the words, so we can stop fencing
A unit test, for me, runs in-process, no network, no real database. It hits a function or a small object I can construct. Table-driven cases. A clock I can freeze. That is a unit.
An integration test talks to a real collaborator I do not mock: Postgres, Redis, the filesystem, a recorded HTTP. I still control the input. I do not call the live Stripe API in CI. I use a recorded fixture or Stripe’s test mode in a nightly if I must.
An end-to-end test drives the UI or the public HTTP the way a client would. I keep a handful. I do not keep a hundred. They flake. They are also the only tests that prove the wiring of cookies, CSRF, and the reverse proxy.
People will argue that a test with Testcontainers is still a unit of the service. I do not care. I care whether it fails when the user’s outcome is wrong and stays quiet when I rename a private method.
The mix I actually put in a repo
Money math, permission checks, state machines: units. Lots of them. They are cheap and they encode the domain. If I cannot unit-test the rule, the rule is stuck in a controller and I extract it. That extraction is design, not coverage.
Webhooks, migrations, “does this query return the lines”: integrations. I want one or two per dangerous path. I use a real Postgres. I do not mock the repository if the bug lives in the SQL.
The happy-path signup and the pay-fail: one or two e2e. Playwright or a HTTP client against a running app. I run them in CI on main, not on every lint. I have paid for that patience.
A ratio I have actually shipped on a billing service: maybe 80 table-driven units, 15 Testcontainers integrations, 3 e2e. Not a pyramid I drew in onboarding. A pile that matched the risk. The marketing site next door had 5 units and 2 e2e. That was also correct. One size is the debate I left.

When I would rather have 20 integrations than 200 units
When the 200 units are mock portraits. When the risk is the join, the transaction, the unique index, the isolation level. When the team cannot keep the mocks honest after a rename. Then twenty tests that hit the database will out-protect two hundred verifies.
When the ORM is in the middle. Hibernate and Active Record lie in units. They tell the truth in SQL. I have been the person who found 400 queries behind one page. A unit test of the service method with a mock repo will never find that. An integration that logs statements will.
When I am the only one who understands the unit suite. A new hire can read an integration that says “given this row, the API returns 422.” They cannot read a mock graph without me. Onboarding is a test-suite feature.
When that slogan is wrong
When the domain is a thick function and the database is a bag. Tax rules, proration, a parser. Two hundred units here are a gift. Twenty integrations will be slow and will not enumerate the cases. I have seen teams “go integration-only” and lose the table of examples finance sent. Do not do that.
When CI is already a coin flip. Adding twenty container tests to a flaky runner makes people ignore red. I fix the runner — or I run the heavy tests on a nightly — before I grow the pile. A suite people ignore is not a mix. It is a ritual.
When the integrations are a second copy of the units with a real DB glued on. That is not a strategy. That is duplication. I pick one layer per risk. I do not pay twice for the same assertion.

CI shape, flakes, and who waits
I split jobs. Units run on every push, in parallel, no Docker if I can avoid it. Integrations run on every push too if they fit in a few minutes with a service container. E2E runs on main and on a nightly. I have merged a unit-only green that hid a migration bug. That is why integrations are on the PR when the PR touches schema. I am willing to wait four minutes. I am not willing to wait twenty for a flake I cannot bisect.
Ownership matters. If the e2e fails because a CSS class moved, a frontend owner fixes it. If I make every backend PR wait on that, I have built a tarp. I tag tests. I do not make the mix a single blob that fails for everyone.
I record HTTP at the edge (vcr, go-vcr, Polly) so an “integration” with Stripe is not a live hope. Live hopes train people to rerun. Reruns are how you ship a red suite. I would rather a recorded interaction that I re-record on purpose when the vendor changes a field.
Contract tests (Pact or a simple OpenAPI snapshot) sit between unit and e2e. I use them when I have two repos. I do not use them when I have one repo and a type I already share. Tools are for the seam that exists.
A story that changed my ratio
We had 1,400 unit tests and 4 integrations. A tax change shipped green and billed the wrong jurisdiction because the unit took a stubbed “address” and the real row had a missing postal code the ORM treated as valid. The integration we added later used a real row and failed. I deleted 200 mock tests in that package over the next month. I added 8 integrations. Incidents on that path stopped. The suite got slower by forty seconds. Nobody missed the verifies. I missed the extra coffee I used to drink while the mocks ran. I got over it.
The inverse story: a parser with 12 integrations and 9 units. Every case booted Postgres. CI was a punishment. I extracted the parser. I wrote 90 table-driven units. I left 2 integrations for “we persist the parse.” CI dropped a minute. That slogan — twenty integrations over two hundred units — would have been malpractice there. Slogans are not mixes. Incidents are.
How I decide for a new path
- Is the risk a rule I can name with examples? Units first. Extract the function.
- Is the risk a query, a transaction, or a vendor adapter? One integration. Log the SQL.
- Is the risk a browser cookie or a redirect? One e2e. Then stop.
- Did I write a mock verify? I delete it or I replace it with an outcome. I said that already. I mean it again.
I do not start with a quota. I start with the last incident. The last incident tells me which layer was missing. I add that layer. I do not add the other layer to feel balanced. Balance is a slide. Incidents are a mix.
I still like a fast suite I can run on save. That suite is units plus the tiny integrations that fit in ten seconds. The rest waits for CI. If I cannot run something on save, I will not run it enough and I will not trust it. Trust is the real ratio. I ship the mix I will run. I do not ship the mix a book would admire.
If a new hire asks “are we a unit-test team or an integration-test team,” I say we are a team that matches the test to the risk. Then I show them the billing folder and the parser folder. The folders disagree on purpose. That disagreement is the mix. A debate that tries to pick one folder for the company is how you get the wrong pile on the next path. I stopped having that debate. I still have the folders. I will keep both. I will keep changing the ratio when the last incident says so. That is the only debate I still attend. If you want a slogan anyway, use this one: test the thing that last lied. The last lie is rarely a private method. It is usually a query, a vendor field, or a rule you stubbed. Put the test there. Leave the conference sport to the conference. I will keep showing the two folders until the question dies.