web / health checks

I run health checks on my web services during deploy to make sure the new version works before it receives traffic, and every minute to ensure the service is still up.

My servers serve GET /healthz. It returns 200 ok when the service works and 503 with the reason when it does not. Cache-Control: no-store stops a cached answer from fooling the caller. A two-second timeout stops a failed dependency from holding the caller.

The endpoint takes no credentials, because it cannot sit behind the login it might need to check. It reads only what the service needs to answer.

Check the dependency, not the handle

For Postgres or MySQL, SELECT 1 is a real check. It proves that the server runs SQL over the network.

For SQLite, SELECT 1 proves almost nothing. It touches no table, so the engine does not open the database file. A corrupt file, a dropped table, or wrong file permissions still return 200.

I read one row from a core table instead:

SELECT 1 FROM games LIMIT 1

One row proves that the file reads and the schema is intact. I use LIMIT 1 because count(*) scans the whole table and adds no signal.

An empty table is a pass. The query ran, so the file and schema work. An empty games table is a board with no games yet. In Go, that makes sql.ErrNoRows a pass:

func (h *handler) serveHealth(w http.ResponseWriter, r *http.Request) {
	w.Header().Set("Cache-Control", "no-store")
	ctx, cancel := context.WithTimeout(r.Context(), 2*time.Second)
	defer cancel()
	var n int
	err := h.db.QueryRowContext(ctx, "SELECT 1 FROM games LIMIT 1").Scan(&n)
	if err != nil && !errors.Is(err, sql.ErrNoRows) {
		http.Error(w, "db: "+err.Error(), http.StatusServiceUnavailable)
		return
	}
	_, _ = io.WriteString(w, "ok\n")
}

db.Ping proves only that the connection is open. It does not prove that the file reads.

A static server checks liveness only

A server that serves files has no database. Its check is liveness: 200 ok proves that the process serves HTTP.

What reads the answer

A deploy boots the new build and checks its /healthz. It flips the upstream only after a 200.

The deploy checks from the same machine, because that is where the new process runs. It asks the idle slot on its own port, before the proxy sends it any traffic. The check proves that the new process is well. It does not prove that the internet can reach the service.

sudo systemctl restart "$unit$idle"
for i in $(seq 20); do
  if curl -sf -m 2 "http://127.0.0.1:$idle/healthz" > /dev/null; then
    healthy=1
    break
  fi
  sleep 1
done

If no 200 comes, the deploy stops the idle slot and exits with an error. The old slot keeps the traffic. If a 200 comes, the deploy points the proxy at the new slot and stops the old one.

A watchdog checks every minute. This costs little, because the endpoint renders no page and reads at most one row.

← All articles