web / health checks
I run health checks on my web services during deploy to make sure the new version works before it receives traffic, and every minute to ensure the service is still up.
My servers serve GET /healthz. It returns 200 ok when the service
works and 503 with the reason when it does not.
Cache-Control: no-store stops a cached answer from fooling the
caller. A two-second timeout stops a failed dependency from holding
the caller.
The endpoint takes no credentials, because it cannot sit behind the login it might need to check. It reads only what the service needs to answer.
Check the dependency, not the handle
For Postgres or MySQL, SELECT 1 is a real check. It proves that the
server runs SQL over the network.
For SQLite, SELECT 1 proves almost nothing. It touches no table, so
the engine does not open the database file. A corrupt file, a dropped
table, or wrong file permissions still return 200.
I read one row from a core table instead:
SELECT 1 FROM games LIMIT 1
One row proves that the file reads and the schema is intact. I use
LIMIT 1 because count(*) scans the whole table and adds no signal.
An empty table is a pass. The query ran, so the file and schema work.
An empty games table is a board with no games yet. In Go, that makes
sql.ErrNoRows a pass:
func (h *handler) serveHealth(w http.ResponseWriter, r *http.Request) {
w.Header().Set("Cache-Control", "no-store")
ctx, cancel := context.WithTimeout(r.Context(), 2*time.Second)
defer cancel()
var n int
err := h.db.QueryRowContext(ctx, "SELECT 1 FROM games LIMIT 1").Scan(&n)
if err != nil && !errors.Is(err, sql.ErrNoRows) {
http.Error(w, "db: "+err.Error(), http.StatusServiceUnavailable)
return
}
_, _ = io.WriteString(w, "ok\n")
}
db.Ping proves only that the connection is open. It does not prove
that the file reads.
A static server checks liveness only
A server that serves files has no database. Its check is liveness:
200 ok proves that the process serves HTTP.
What reads the answer
A deploy boots the new build and checks its /healthz. It flips the
upstream only after a 200.
The deploy checks from the same machine, because that is where the new process runs. It asks the idle slot on its own port, before the proxy sends it any traffic. The check proves that the new process is well. It does not prove that the internet can reach the service.
sudo systemctl restart "$unit$idle"
for i in $(seq 20); do
if curl -sf -m 2 "http://127.0.0.1:$idle/healthz" > /dev/null; then
healthy=1
break
fi
sleep 1
done
If no 200 comes, the deploy stops the idle slot and exits with an error. The old slot keeps the traffic. If a 200 comes, the deploy points the proxy at the new slot and stops the old one.
A watchdog checks every minute. This costs little, because the endpoint renders no page and reads at most one row.