Docker Compose depends_on service_healthy vs a Sleep: The Race the Healthcheck Misses
Felix Braun
September 21, 2026
depends_on with a hard-coded sleep 10 in an entrypoint is the homelab classic: it works on a quiet evening and fails when the disk is busy, the image is cold, or migrations run long. depends_on: condition: service_healthy looks like the grown-up fix—until the healthcheck itself lies, races the app’s listen socket, or flips healthy before the schema is ready. The race does not disappear. It moves into the check.
This is for solo operators who already know Compose healthchecks exist and still see “started” stacks that are not ready.
What service_healthy actually waits for
With Compose v2 semantics, depends_on can wait until a dependency’s healthcheck reports healthy—not merely until the container process starts. That is the right idea: the API should not boot until Postgres accepts connections, the worker should not start until Redis answers PING.
It is still only as honest as the probe. A TCP check on 5432 can pass while the database rejects your app user. An HTTP / can return 200 from a reverse proxy while the upstream app is still migrating. service_healthy waits for the check you wrote, not for the readiness you meant.

Why sleep felt easier—and why it fails
Sleep is a blunt timer. It ignores load. On a Pi, ten seconds might be fine on Tuesday and fatal on Friday after a reboot when apt and Docker fight for I/O. Sleep also wastes time when the dependency is already ready in two seconds. You pay latency forever to avoid thinking about probes.
Teams keep sleep because it is visible in the entrypoint and requires no healthcheck syntax. Visibility is not correctness. The race is still there; you just cannot see it until traffic arrives.
Sleep also trains bad habits across environments. The value that works on your NVMe homelab box fails on a cheap VPS with noisy neighbors. You end up maintaining different sleep constants per host—an unofficial config language nobody wanted. Healthchecks travel with the compose file; magic numbers do not.
depends_on without condition is not a wait
A common misconception: plain depends_on: [db] means “wait until db is ready.” It does not. It mostly controls start order for dependency graphs and does not block on readiness. If you never set condition: service_healthy (or completed for oneshots where applicable), you still have start-order theater. That misunderstanding is why people add sleep “on top of depends_on” and think they covered it.
Read your Compose version docs for the exact condition names available. Then wire them. Do not assume YAML proximity equals synchronization.
The races healthchecks still miss
- Listen before ready: process binds the port, healthcheck passes, migrations still run.
- Start period too short: flapping failures mark the container unhealthy, Compose gives up, dependents never start—or the opposite, start_period hides real failures.
- Wrong probe:
wgethits a static file while the app worker is dead. - Shared network assumptions: check succeeds from inside the container namespace differently than clients experience.
- One-shot init containers: a healthy DB does not mean your migrate job finished; model that as its own service or job.
If you only switched from sleep to service_healthy without redesigning the probe, you upgraded the waiter, not the signal.
A practical pattern that holds up
Write readiness probes that exercise the dependency the way the app does: authenticate, select 1, hit /readyz that checks DB connectivity. Use start_period / start_interval so slow boots do not flap. Keep liveness stricter or separate if you need restart behavior—do not overload one check to mean everything.
For migrations, prefer an explicit migrate service with restart: on-failure and make the API depends_on that service completing—or run migrate in CI before deploy. Do not pretend a DB healthcheck means schema version N.
Keep a tiny sleep only as a last resort behind a comment that names the bug you could not probe—and file a ticket to delete it.
App-level retry is complementary, not a substitute. Your API should retry DB connections with backoff so a brief blip does not crash the process. That still does not mean you should start five workers that stampede a restoring database because Compose thought “started” meant “ready.” Defense in depth: honest healthchecks to order the boot, retries to survive jitter.
Observability while you wait
When service_healthy blocks forever, you need to see why. docker compose ps and health logs should show failing probes. If your check is silent, fix the check’s output. A waiter without diagnostics recreates the worst part of sleep: staring at a blank terminal wondering which second was the wrong second.
Export container health to your existing notifier if you have one. A stack stuck unhealthy after reboot is an alert, not a curiosity for whenever you next SSH in.
When sleep is still temporarily OK
Throwaway demos, single-user lab stacks, and bring-up experiments can use sleep if you accept flaky reboots. Production-ish homelab services that friends rely on should not. The cost of a real healthcheck is lower than the cost of a Saturday spent restarting compose in the right order.
Decision rule
Replace sleep with service_healthy only when the healthcheck encodes true readiness. If the check is a decorative TCP ping, you still have a race—you just renamed it.
Compose will wait patiently for a lie. Your job is to stop lying in the Dockerfile HEALTHCHECK and the Compose healthcheck block. Then depends_on: condition: service_healthy becomes what the docs promised: an ordered start that matches reality.
Sleep hides timing. Bad healthchecks hide unreadiness. Good healthchecks plus service_healthy hide neither—and that is the point.
If your stack “sometimes” works after reboot, treat that as a failed healthcheck design, not as a cursed machine. Add logging around probe failures, lengthen start_period with intention, and make the ready endpoint print which dependency failed. Future you during an outage wants a sentence, not another sleep.
Order the boot with signals you trust. Everything else is choreography that works until the night the NAS scrub and the container start coincide.
One more operational habit: after you change healthchecks, force a full docker compose down and up on a machine that is busy—run a disk scrub or a backup during the test. Healthy-path boots on an idle SSD teach you nothing about the race you are trying to delete. If the stack still comes up clean under load, you earned the removal of the sleep.
Document the ready contract in the repo README in one paragraph: what endpoint means ready, which service waits on which, and how long start_period is allowed to be. The next person—or next you—will otherwise reintroduce sleep “just to be safe,” and the race returns with a friendlier comment.
Think of service_healthy as a lock that opens only when your probe tells the truth. Sleep is a lock that opens when a timer expires, truth optional. Homelabs that survive messy reboots pick the first kind and accept the upfront work of writing checks that match how the application fails. Homelabs that feel haunted by “random” boot order usually still have a sleep somewhere—or a healthy check that never asked the real question.
Delete the sleep when the probe earns it. Keep the condition. Make readiness boring. Boring boots are the feature.