The week my monitoring showed zero
The workers were alive, the dashboards showed zero downloads, and the only alert that fired pointed in the wrong direction. Why Prometheus service discovery quietly returned an empty list, and why nothing treated that as an error.
Symptom
For seven days the Grafana dashboards showed zero downloads, although the workers were alive.
Exactly one alert fired: “workers missing”, through its absent() branch. But the workers were
there.
Why everything looked healthy
Worker metrics were collected through Docker service discovery: Prometheus asks Docker which containers exist and scrapes them. When discovery returns an empty list, Prometheus has no targets, so it has no failed scrapes either. The job does not fail. It is simply empty.
Cause
The Docker client built into Prometheus, which discovery relies on, talked to the daemon using API version 1.24. The daemon accepted nothing older than 1.40 and refused the request. The refusal stayed inside discovery and never surfaced, neither as a failed scrape nor as a failed job.
Fix
The workers now expose metrics on fixed ports, and Prometheus scrapes them from a static list. The alert rule expects exactly three targets: fewer than that is an outage, not “no data”.
The cost of this fix is written down next to it: changing the replica count means updating the port range, the target list and the number in the rule. Discovery can come back only after Prometheus is upgraded.