Docker Compose start_period vs start_interval: When the Stack Looks Up but Nginx Isn’t Ready

Ivo Markham

Ivo Markham

September 18, 2026

Docker Compose start_period vs start_interval: When the Stack Looks Up but Nginx Isn't Ready

I have a Compose file I still treat as the house front door: Nginx as the only published port, Immich or Paperless or a boring Ghost blog behind it, depends_on with condition: service_healthy so the proxy is not supposed to take traffic until something real answers. On a good morning the stack comes up in twenty seconds. On a bad one, Portainer says every container is healthy, Uptime Kuma is green, and the first curl from the laptop is a 502 that lasts long enough for me to open the wrong log.

That gap is almost never “Docker is lying.” It is me asking the wrong timer two different questions. start_period is how long Compose forgives a failing probe. start_interval is how often it asks during that forgiveness window. Mix them up and Nginx looks ready because a process bound :80, or it looks broken because you only sampled it twice while the worker was still compiling templates.

The failure I keep reproducing

The stack is ordinary. Nginx official image, a named volume for conf.d, proxy_pass to an upstream named app on the Compose network. The app is whatever I was impatient about that week: Nextcloud, Immich’s immich-server, or a Node API that runs Prisma migrate on boot. The healthcheck I copied from a gist is some variant of:

healthcheck:
  test: ["CMD", "wget", "-qO-", "http://127.0.0.1/"]
  interval: 30s
  timeout: 5s
  retries: 3
  start_period: 10s

Nginx’s master process is up in under a second. wget to / hits the default server or the first server block, which reverse-proxies to an upstream that is still running migrations, still warming a Rust binary, or still opening Postgres. wget gets 502. During start_period those failures do not count. After ten seconds they do. Three failures at a 30-second interval means I can sit in “unhealthy” for more than a minute after the app is actually fine — or, if I “fix” it by probing nginx -t instead of HTTP, I flip healthy while every browser tab is still a 502.

The dashboard is not wrong. It is answering “did the probe I wrote succeed,” which is a different sentence from “can a human load the site.”

Homelab mini PC and switch with one port still amber while the others are green

What the two knobs actually do

Docker Engine treats a container with a healthcheck as starting until the first success, or until start_period expires and failures start counting toward retries. After that it is healthy or unhealthy. Compose’s depends_on: condition: service_healthy waits on that state. Portainer’s green badge waits on that state. Nothing in that pipeline knows what Nginx considers a good request.

start_period is a mute button on failures. If the first successful probe happens inside the period, Docker immediately calls the container healthy and the mute ends. That last sentence is the one people skip. A lucky 200 on a half-booted app spends your grace window and then the next 502 can mark you unhealthy. A probe that succeeds for the wrong reason — pidof nginx, a TCP open on 80, nginx -t against a config that points at a dead upstream — spends it on purpose.

start_interval showed up later (Compose file support landed with Engine 25 / Compose 2.20-era bits; if your host is older, the key is silently ignored and you are only using interval). It is the poll rate while the container is still in the start window. Default, if you omit it, is not “same as interval.” On current Engine it is five seconds. That is why two files that look identical can behave differently on a 2023 NAS versus a 2026 mini PC: one host is sampling Nginx every five seconds during boot, the other is stuck on your 30-second interval for the whole life of the container.

Once the container is healthy, interval takes over. That is the number you want coarse. Hitting wget every five seconds forever is how a tiny Celeron box spends more CPU on the proxy’s own healthcheck than on the site.

The Nginx-shaped lie

Nginx is a bad default for “is the stack ready” because it is designed to be up before its backends. That is a feature when you want to serve a custom 502 page. It is a trap when Compose is using Nginx as the readiness gate for everything else.

Three probes I have actually shipped, and what they hide:

  • CMD-SHELL pidof nginx || exit 1 — process exists. Healthy while conf.d is still an empty volume and you are on the stock welcome page.
  • CMD nginx -t — config parses. Healthy while every proxy_pass target is host not found at runtime, because Nginx resolves some names at startup and some at request time depending on how you wrote the upstream block.
  • CMD wget -qO- http://127.0.0.1/api/health — honest for an API that has a health route. Useless if that route is implemented as “return 200 if the process is up” and the worker pool is still empty.

The Immich compose samples people copy are better than my old Ghost file: they probe specific containers (immich-server, Redis, Postgres) and do not pretend the proxy is the source of truth. When I put Nginx in front anyway, I stopped making other services wait on Nginx’s health. Nginx waits on the app, or it does not wait at all and I accept a short 502. Making the app wait on Nginx while Nginx’s probe waits on the app is how you get a stack that never leaves starting.

If you want a proxy-level probe that means “a browser would not hate this,” it has to hit a path the upstream actually serves, and it has to treat 502/503 as failure. wget --spider against a static /healthz that Nginx itself serves from alias will go green while Immich is still migrating. I have done that on purpose for a status page. I have also done it by accident and then wondered why Kuma agreed with Compose and not with my phone.

When to lengthen start_period

Lengthen start_period when the first honest probe is allowed to fail for a known, bounded reason. Nginx reloading a large conf.d with a dozen server blocks is not that reason — that is milliseconds. A PHP-FPM container that runs composer install because you built the image wrong is. A Paperless-ngx web container on a spinning rust NAS that needs forty seconds to talk to Tika the first time after a reboot is. Nextcloud’s occ upgrade on a 4 TB volume is.

I time it once with docker compose up and a stopwatch, then add slack, then refuse to add more slack the next time it flakes. If the number crosses two minutes I stop calling it a healthcheck problem. Something in the image is doing real work on every start, and a mute button will not make that work cheaper.

A long start_period with a long interval and no start_interval is how you get the other lie: the app became ready at T+12s, the next probe is at T+30s, and for eighteen seconds every dependent container is still blocked. That is not caution. That is a coarse clock.

Basement homelab rack with mixed mini PCs and loosely dressed cables

When to set start_interval

Set start_interval when the start window is long and you care about the moment of the first success. Nginx in front of a slow app is the textbook case if you insist on probing HTTP through the proxy. Probe every five seconds during a 60-second start_period, then back off to 30s once you are healthy. You flip as soon as the upstream answers, without spending the rest of your life at five-second resolution.

Do not set start_interval shorter than the probe can finish. A wget with timeout: 10s and start_interval: 2s overlaps probes on a loaded ARM box. Docker will not save you from that; you will just see timeouts that look like application failures.

Do not set it if your probe is expensive. curl against Nextcloud’s login page pulls PHP, Redis, and the database. Doing that every two seconds during a three-minute upgrade is a load test you did not schedule. A cheap TCP check is the wrong honesty. A full page fetch at high frequency is the wrong kindness.

On Synology Container Manager and some older Unraid Docker installs I still see Engine builds that ignore start_interval. If docker inspect does not show StartInterval on the container, the YAML was a comment as far as that host is concerned. Fix the Engine or stop pretending the key is doing work.

depends_on is not a startup scheduler

Compose v2 will wait for service_healthy, then start the next container. It will not retry that decision if Nginx later goes unhealthy. It will not restart the app when Nginx starts serving 502s at minute five. It will not order docker compose restart nginx after you recreate the app. People treat depends_on like systemd. It is a one-time gate for up.

That matters for Nginx because the interesting failure is often after the first healthy. You recreate immich-server, the proxy stays up, Kuma still hits Nginx’s own /healthz, and you have a green stack with a dead upstream until the next interval if you even probe the upstream at all. start_period does not come back for a recreate of a different service. If you need that, the healthcheck belongs on the app, or you need a reload hook, not a longer grace on the proxy.

I have also seen restart: unless-stopped on Nginx fight a strict healthcheck. Engine can leave a container running and unhealthy forever. Compose will not recycle it unless you add a restart policy that actually keys off health, which the homelab compose files I copy almost never do. Unhealthy and running is a state. Portainer paints it yellow. I still miss it if I only look at “is the container up.”

A Compose snippet I will actually keep

For a proxy I want honest, on a host that understands both keys:

healthcheck:
  test: ["CMD-SHELL", "wget -qO- http://127.0.0.1/ready | grep -q ok"]
  interval: 30s
  timeout: 3s
  retries: 3
  start_period: 45s
  start_interval: 5s

The /ready location is a tiny Nginx location that proxy_passes to the app’s real readiness route and does not cache. If the app has no such route I add one, or I drop the proxy out of the dependency graph. I do not invent readiness by curling the marketing homepage.

For Nginx itself I often use no healthcheck. The app has one. Postgres has one that runs pg_isready, which is boring and correct. Redis has redis-cli ping. Nginx starts whenever it starts. Browsers eat a few 502s. I care more about not creating a circular wait than about a perfectly ordered up.

If the slow thing is not Nginx at all — a Rails container that runs db:migrate in the same process that later serves HTTP — you want a different shape than a reverse-proxy grace window. A migrate job that flips the healthcheck healthy before the migration finishes is the one-shot problem, not the start_period versus start_interval problem.

What I check before I touch the timers

I run docker inspect --format "{{json .State.Health}}" nginx and read the last few probe outputs. If they are 502s, the timers are not the bug; the upstream is. If they are connection refused, Nginx is not listening yet and a ten-second start_period is fine. If they are 200s on the default welcome page, my probe is pointed at the wrong server_name and I have been measuring Alpine’s stock index.html.

I check image tags. nginx:latest on Tuesday and nginx:1.27-alpine on Thursday will disagree about whether wget is even in the image. Alpine variants often need wget or curl added, or you switch the test to CMD-SHELL with whatever the image actually ships. A failing binary looks identical to a failing upstream if you only watch the badge.

I check which Compose implementation I am on. docker compose v2, the old docker-compose binary, Portainer stacks, and Dockge will parse the same YAML and still drop keys they do not know. If I need start_interval, I verify it on the running container, not in the editor.

I check whether I am probing loopback or the published port. wget http://127.0.0.1 inside the container is not the same as hitting the host’s 192.168.1.12:443 through a Cloudflare tunnel or a Caddy sibling. TLS, HTTP/2, and proxy_protocol have all made a “healthy” Nginx refuse the request I thought I was simulating.

Trade-offs I will not dress up

Honest HTTP through Nginx couples the proxy’s health to every backend. That is what you want for a single-site box. It is wrong for a shared proxy with five server blocks: Immich down should not mark the whole Nginx container unhealthy and page you for the blog that is still fine. Split the probes, or split the proxies. I have done both. Two Nginx containers on one host is ugly and easy to reason about. One Nginx with a /ready that ORs five upstreams is clever and I have regretted it at 1 a.m.

start_period: 0 and a generous retries is valid. You are choosing to count every failure from t=0. That is cleaner on a service that is either up in two seconds or never coming. It is noisy on Nextcloud.

Kubernetes people will say this is what readiness versus liveness was invented for. They are right, and I am still not running k3s for a blog and a photo library. Compose has one health field. You get to pick whether it means “restart me” or “wait for me.” If you use it for both, you will eventually restart a container because it was slow, or wait forever on a container that should have been killed. I use Compose health almost only as a start gate. Restarts stay on process crash.

The short version I wish I had taped to the rack

start_period is how long a failing probe is allowed to be a startup story. start_interval is how often you ask during that story so you notice the first real success. Nginx will happily tell a story about a listening socket while the site is a 502. If the badge and the browser disagree, believe the browser, then change the probe, then touch the timers. The other order is how a stack looks up for a month and still is not ready.

More articles for you