Docker Compose Healthcheck: Stop Trusting depends_on

Short version for the impatient: depends_on on its own only waits for a container to start, not for the thing inside it to be ready. If your app boots before Postgres accepts connections, that is the bug, and a healthcheck plus condition: service_healthy is the fix.

I have shipped this bug more than once. The pattern is always the same. Everything works on my laptop because the database container is warm from yesterday. Then I run docker compose up on a fresh VPS, the API container starts in two seconds, Postgres takes eight to run its init scripts, and the API dies with "connection refused" and a stack trace that looks like it is my fault. With restart: unless-stopped it eventually recovers, so I called it flaky and moved on. That was a mistake. A deploy that only works on the third attempt is a coin flip with extra steps.

This post covers what the healthcheck options do, a Compose file I would actually run on a small server, and the part nobody warns you about: an unhealthy container does not get restarted by Compose. Everything below is checked against the Docker docs, which I link as I go.

What depends_on really waits for

The short syntax is the one most of us write:

services:
  api:
    build: .
    depends_on:
      - db
  db:
    image: postgres:18

Compose starts db first, then api. That is the whole contract. "Started" means the container process exists. Postgres may still be replaying WAL or running your init SQL, and Compose does not care. The Docker startup order guide says this directly: a database needs to start its own services before it can handle incoming connections, and you detect that ready state with the condition attribute.

There are three conditions. service_started is what the short syntax gives you. service_healthy waits until the dependency's healthcheck passes. service_completed_successfully waits for a one-shot container to exit with code 0, which is the right tool for migrations. I will come back to that last one, because it is the underrated option.

Writing a healthcheck that tells the truth

A healthcheck is a command Docker runs inside the container on a timer. Exit code 0 means healthy, 1 means unhealthy. The Compose services reference lists the knobs: test, interval, timeout, retries, start_period and start_interval.

For Postgres, the image ships with pg_isready, so there is nothing to install:

  db:
    image: postgres:18
    environment:
      POSTGRES_USER: app
      POSTGRES_PASSWORD: ${DB_PASSWORD}
      POSTGRES_DB: app
    healthcheck:
      test: ["CMD-SHELL", "pg_isready -U $${POSTGRES_USER} -d $${POSTGRES_DB}"]
      interval: 10s
      timeout: 5s
      retries: 5
      start_period: 30s

That $$ is not a typo. Compose treats a single $ as its own variable interpolation and would try to fill it from your host environment, usually with an empty string. Doubling it passes a literal $ through, so the shell inside the container expands the variable. I lost a long evening to a healthcheck that ran pg_isready -U -d because I wrote one dollar sign.

The start_period option is the one I misunderstood for a long time. It is a grace window. According to the Docker docs, failures during that window do not count against retries. If a check succeeds during the window, the container counts as started and later failures start counting. So a slow first boot does not burn your five retries before Postgres has even finished initialising.

The start_interval option, added in Compose 2.20.2 per the reference, lets you probe more often during the start period. Mine is usually 2s, so the dependent service starts a few seconds sooner on a fast machine. It is optional and I skip it more often than not.

The curl trap in slim images

For your own web service, the instinct is curl -f http://localhost:3000/healthz. It works until you switch to a slim or distroless base image and curl is no longer there. The check fails, the container goes unhealthy, and the logs of the app itself look perfectly fine. I stared at a healthy-looking Node process for twenty minutes before I ran docker inspect and read the health log.

Use whatever the image already has. Alpine images carry BusyBox wget, and a language runtime can probe itself:

  api:
    build: .
    healthcheck:
      test: ["CMD", "wget", "-qO-", "http://localhost:3000/healthz"]
      interval: 15s
      timeout: 3s
      retries: 3
      start_period: 20s

Keep the endpoint cheap. My /healthz returns 200 if the process can answer HTTP, and that is all. I used to make it query the database too, on the theory that "healthy" should mean "fully working". Then a ten-second database stall turned every API container unhealthy at once, and the dependent services lined up behind them. A shallow check answers one question, which is whether this process is alive. Your monitoring can ask the deeper questions separately.

A Compose file I would run on a small server

Here is the shape I use for a typical app with Postgres, a migration step and the API:

services:
  db:
    image: postgres:18
    restart: unless-stopped
    volumes:
      - pgdata:/var/lib/postgresql
    healthcheck:
      test: ["CMD-SHELL", "pg_isready -U $${POSTGRES_USER} -d $${POSTGRES_DB}"]
      interval: 10s
      retries: 5
      start_period: 30s

  migrate:
    build: .
    command: ["./bin/migrate"]
    depends_on:
      db:
        condition: service_healthy
    restart: "no"

  api:
    build: .
    restart: unless-stopped
    depends_on:
      db:
        condition: service_healthy
        restart: true
      migrate:
        condition: service_completed_successfully
    healthcheck:
      test: ["CMD", "wget", "-qO-", "http://localhost:3000/healthz"]
      interval: 15s
      retries: 3
      start_period: 20s

volumes:
  pgdata:

The migrate service is the underrated condition from earlier. It runs once, exits 0, and api does not start until that has happened. No entrypoint script that loops on pg_isready, no sleep 10 hidden in a Dockerfile. I deleted both of those from old projects and nothing was lost.

The restart: true flag under the db dependency is a newer option. Per the reference it needs Compose 2.17.0, and it restarts the dependent service when you restart db through a Compose operation such as docker compose restart. The docs are explicit that this does not cover automatic restarts by the container runtime after a crash. Keep that distinction in mind, because it leads straight into the next problem.

To bring the whole thing up and know whether it worked, use the --wait flag from the docker compose up reference:

docker compose up -d --wait --wait-timeout 120

--wait implies detached mode and blocks until services are running or healthy. It exits non-zero if they are not, which makes it usable in a deploy script or a CI job. That one line replaced a block of sleep and curl retries in my deploy scripts. If you are weighing deploy tooling more broadly, I compared two self-hosted options in Coolify vs Dokploy after migrating a client, which is a different question from this one but the same kind of server.

Unhealthy does not mean restarted

Here is the part I got wrong the longest. I assumed that an unhealthy container would be restarted. In plain Docker Compose it is not. The restart policy reacts to a container exiting. A container whose process is alive but whose healthcheck keeps failing just sits there labelled unhealthy, and traffic keeps going to it if nothing in front of it checks that label.

So the healthcheck gives you three things: ordering at startup, a status in docker ps, and a signal for --wait. It does not give you self-healing. If you want automatic recovery on a single host, you need something that watches the status and acts. A tiny sidecar like willfarrell/autoheal is a common choice, and a reverse proxy that routes only to healthy upstreams covers the traffic side. Alternatively, make the app exit on its own when it detects it is broken, and let restart: unless-stopped do its job.

I have not settled on one answer here. For my own small deployments I let the app crash loudly and rely on the restart policy, plus an uptime monitor that pings from outside. For anything with real customers I add the proxy-level check as well. If your setup is bigger than a couple of containers on one VPS, this is a sign to look at an orchestrator, because Swarm and Kubernetes treat health as an input to scheduling and Compose does not.

What to change this week

Open your current compose.yaml and look for every depends_on that has no condition. For each database or cache, add a healthcheck and switch the dependency to service_healthy. Move your migration command into its own service with service_completed_successfully. Then run docker compose up -d --wait on a machine with an empty volume and time it. If it passes cold, you have fixed the bug. If you want to see how I set up this kind of infrastructure for client projects, my work is on abrarqasim.com.

Originally published at abrarqasim.com. I write there about React, PHP, Rust, Go and the AI tooling around them.

Story originally reported by Dev.to. View at Dev.to →
← Back to all news