monitoring: recover host checks for the whole estate, reported to Gatus
Five checks, 27 endpoints, replacing what Uptime Kuma used to watch:
is it up every 5min, all hosts
is disk full daily, all hosts
is CPU hot every 5min, nodito
is ZFS broken daily, nodito
is UPS online every 5min, nodito
Two roles, kept separate so neither knows about the other - they meet at a URL
and a token, the same way caddy_site and each service meet at a vhost:
roles/gatus_endpoint runs on the observability host, writes ONE file into
/opt/gatus/config/endpoints/. Gatus merges every *.yaml
there and appends lists, so callers compose without
coordinating.
roles/healthcheck runs on the monitored host: a check script, a systemd
service, a timer, and an optional push. Ships a library
of check bodies under templates/checks/.
Everything PUSHES. Gatus never reaches out, which matters because nodito and its
VMs are behind NAT, and because four of the five checks are internal state with
no pollable surface at all. Liveness pushes too, deliberately: a heartbeat proves
the host is running AND can reach the internet, where an ICMP probe from one
vantage point only proves it answers pings from there. And since Gatus alerts
when a heartbeat window expires, a check that stops running raises the alarm by
itself - a dead timer looks exactly like a dead host, which is the correct
reading.
One bearer token per host, generated straight into the vault and never printed.
A token only writes results for its own host's endpoints, so a compromised host
can lie about itself, which it could do anyway.
Three things learned from the source that shaped this:
* Gatus polls its own config every 30s and reloads
(main.listenToConfigurationFileChanges), so gatus_endpoint needs no restart
handler - writing the file IS the deploy.
* ...but on a reload it panics if the new config fails to parse, unless
skip-invalid-config-update is set. Endpoint files are contributed by other
playbooks, so one malformed file would take the monitor down at the worst
possible moment. Now set.
* The push URL uses a key Gatus computes, not the name you write:
sanitize(group) + "_" + sanitize(name), lowercased with / _ . , space # + &
replaced by "-" (config/key/key.go). So knots_box_local is
knots-box-local in the URL. The playbook derives it rather than hand-writing.
storage: maximum-number-of-results 900, up from upstream's 100. Gatus bounds the
database by COUNT and trims inline on insert, so there is no retention job and
no way to fill a disk - but history depth is then a function of check frequency,
and 100 results at a 5-minute interval is 8 hours. 900 is ~3 days of liveness
and ~2.5 years of the daily disk check. The uptime table is separate and its
30-day retention is hard-coded upstream.
A bug worth recording: the first deploy shipped five scripts that all died with
"syntax error: unexpected end of file". Jinja strips an included template's
trailing newline and trim_blocks then eats the newline after {% endif %}, so the
closing brace of check() landed on the same line as the body's last statement -
`return 0}`. Every check was broken and the deploy still reported failed=0,
because the role's "run once" task has failed_when: false and reports the result
as a debug message nobody read. The blank line that fixes it is now load-bearing
and commented as such.
Verified by triggering every unit by hand rather than waiting on timers: all
checks exit 0 on all hosts, and Gatus shows 26 UP / 1 DOWN. The one DOWN is
liveness_watchtower, which is honest - that host currently refuses SSH (TCP
connects, no banner exchange) and is excluded from this deploy. It is also the
box still running Uptime Kuma.
Known waste, not yet fixed: healthcheck installs its dependencies per CHECK
rather than per HOST, so apt runs 29 times estate-wide for a curl that is
already present, and daemon_reload runs 4x per host.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 08:54:10 +02:00
|
|
|
---
|
|
|
|
|
# One health check: a script, a systemd service, a timer, and an optional push.
|
|
|
|
|
#
|
|
|
|
|
# The exit code is the answer and systemd keeps it:
|
|
|
|
|
# systemctl is-failed <name>-healthcheck.service
|
|
|
|
|
# Reporting anywhere else is optional and generic. Point healthcheck_push_url at
|
|
|
|
|
# Gatus, or at whatever replaces it, or at nothing.
|
|
|
|
|
|
|
|
|
|
healthcheck_name: "" # e.g. disk-usage -> disk-usage-healthcheck
|
|
|
|
|
healthcheck_description: ""
|
|
|
|
|
|
|
|
|
|
# The check itself. Pick ONE:
|
|
|
|
|
# healthcheck_check: a template under templates/checks/ (without .sh.j2)
|
|
|
|
|
# healthcheck_command: a shell one-liner that exits 0 for healthy
|
|
|
|
|
healthcheck_check: ""
|
|
|
|
|
healthcheck_command: ""
|
|
|
|
|
|
|
|
|
|
# systemd timer. OnUnitActiveSec unless healthcheck_on_calendar is set.
|
|
|
|
|
healthcheck_interval: "5min"
|
|
|
|
|
healthcheck_on_calendar: ""
|
|
|
|
|
healthcheck_boot_delay: "2min"
|
|
|
|
|
|
|
|
|
|
# ── Reporting ────────────────────────────────────────────────────────────────
|
|
|
|
|
# Gatus external endpoints:
|
|
|
|
|
# POST {url}?success=true|false&error=...
|
|
|
|
|
# Authorization: Bearer {token}
|
|
|
|
|
# Empty url = check and log only, which is a valid state and not an error.
|
|
|
|
|
healthcheck_push_url: ""
|
|
|
|
|
healthcheck_push_token: ""
|
|
|
|
|
|
monitoring: systemd services, domain expiry, DNS correctness, public endpoints
Four more check types, 42 endpoints, taking the estate from 39 to 81.
── systemd services (infra/401) ────────────────────────────────────────────
Every unit we deploy, checked every 5 minutes with a 16-minute heartbeat.
This closes the gap that let a real bug run unnoticed earlier today: a backup
script left forgejo, lnbits, headscale and memos stopped, and NOTHING caught
it. The dumps exited 0, the artefacts were correct, the deploy said failed=0,
and liveness only proves the HOST is up - not that anything on it serves.
One endpoint PER UNIT but only ONE timer per host: the check iterates that
host's units and pushes a result for each, the way check-backups.sh reports per
source. Four units on vipy would otherwise mean four scripts, services and
timers. A single host-level red light would also say "something on vipy is
down" without saying which, which is not the question anyone has.
Keys are host-qualified because unit names collide - caddy runs on four
machines. Which units a host runs lives in host_vars/<host>/monitored_services,
because "what runs here" is a property of the machine, the same reasoning as
the cross-host ports.
nut-driver-enumerator is deliberately excluded: it is a oneshot generator that
is `enabled` but always `inactive`, so it would report down forever. Checked
live before excluding it.
── domain, DNS and public endpoints (infra/402) ────────────────────────────
The first checks that PULL rather than push, and that is the right way round:
all three are about how the outside world sees us, so they must be measured
from outside. Nothing is installed anywhere - no script, no timer, no token.
They also have no heartbeat, because a heartbeat answers "did the reporter
report"; when Gatus does the checking itself, failure is immediate.
domain 1 endpoint, 24h, [DOMAIN_EXPIRATION] > 336h (two weeks)
dns 11 endpoints, 24h, [DNS_RCODE] == NOERROR and [BODY] == the IP
public 14 endpoints, 5m, 11 HTTPS + 3 TCP
Expected A records are derived from inventory (hostvars[host].ansible_host),
not written down again. The estate's recurring bug is an address recorded in a
second place and left behind when the machine moved; asserting against
inventory means a renumbered box is one edit, not two.
Expected HTTP status was checked live per site rather than assumed. Two return
401 - the Gatus dashboard and the DATUM dashboard, both behind basic auth - and
that is what is asserted: expecting 200 there would go green precisely when the
auth broke. headscale asserts /health rather than /, which is a 404 by design.
Every HTTPS check also carries [CERTIFICATE_EXPIRATION] > 168h, which is free
on an endpoint already being polled and catches a renewal that silently stops.
Two things learned the hard way, both now in comments:
* A domain-expiry endpoint needs a URL SCHEME. Gatus derives the endpoint type
from the prefix (endpoint.Type()), so a bare "contrapeso.xyz" is UNKNOWN and
the whole config is rejected. It is https:// plus a DOMAIN_EXPIRATION
condition and no status assertion, so the registrar's parking page at the
apex is irrelevant. Upstream also enforces a 5m minimum interval for that
placeholder, because it uses a free whois service.
* That rejection proved skip-invalid-config-update was worth adding. Gatus
logged "the configuration file was updated, but it is not valid, the old
configuration will continue being used" and kept running. Without it the
reload path calls panic() and one malformed contributed file takes the
monitor down.
Verified: 15/15 service endpoints UP; 26/27 public-facing UP. The single DOWN is
public_memos, correctly - memos-box is powered off, and the condition result
reads [STATUS] (502) == 200. memos-box and arbret-staging-box were shut down
deliberately from the Proxmox UI to test the liveness endpoints; their endpoints
stay registered and will go green when the VMs return.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 09:33:57 +02:00
|
|
|
# Some checks report MORE THAN ONE result - a host with four systemd services
|
|
|
|
|
# needs four endpoints, or a single red light cannot tell you which one died.
|
|
|
|
|
# Those set healthcheck_push_base to the endpoints COLLECTION and the check body
|
|
|
|
|
# appends each key itself, the same way check-backups.sh reports per source.
|
|
|
|
|
healthcheck_push_base: ""
|
|
|
|
|
|
|
|
|
|
# Units for the systemd-units check. Each becomes its own Gatus endpoint.
|
|
|
|
|
healthcheck_units: []
|
|
|
|
|
# Prefix for the per-unit endpoint keys, e.g. "services_vipy" -> services_vipy-caddy.
|
|
|
|
|
healthcheck_units_key_prefix: ""
|
|
|
|
|
|
monitoring: recover host checks for the whole estate, reported to Gatus
Five checks, 27 endpoints, replacing what Uptime Kuma used to watch:
is it up every 5min, all hosts
is disk full daily, all hosts
is CPU hot every 5min, nodito
is ZFS broken daily, nodito
is UPS online every 5min, nodito
Two roles, kept separate so neither knows about the other - they meet at a URL
and a token, the same way caddy_site and each service meet at a vhost:
roles/gatus_endpoint runs on the observability host, writes ONE file into
/opt/gatus/config/endpoints/. Gatus merges every *.yaml
there and appends lists, so callers compose without
coordinating.
roles/healthcheck runs on the monitored host: a check script, a systemd
service, a timer, and an optional push. Ships a library
of check bodies under templates/checks/.
Everything PUSHES. Gatus never reaches out, which matters because nodito and its
VMs are behind NAT, and because four of the five checks are internal state with
no pollable surface at all. Liveness pushes too, deliberately: a heartbeat proves
the host is running AND can reach the internet, where an ICMP probe from one
vantage point only proves it answers pings from there. And since Gatus alerts
when a heartbeat window expires, a check that stops running raises the alarm by
itself - a dead timer looks exactly like a dead host, which is the correct
reading.
One bearer token per host, generated straight into the vault and never printed.
A token only writes results for its own host's endpoints, so a compromised host
can lie about itself, which it could do anyway.
Three things learned from the source that shaped this:
* Gatus polls its own config every 30s and reloads
(main.listenToConfigurationFileChanges), so gatus_endpoint needs no restart
handler - writing the file IS the deploy.
* ...but on a reload it panics if the new config fails to parse, unless
skip-invalid-config-update is set. Endpoint files are contributed by other
playbooks, so one malformed file would take the monitor down at the worst
possible moment. Now set.
* The push URL uses a key Gatus computes, not the name you write:
sanitize(group) + "_" + sanitize(name), lowercased with / _ . , space # + &
replaced by "-" (config/key/key.go). So knots_box_local is
knots-box-local in the URL. The playbook derives it rather than hand-writing.
storage: maximum-number-of-results 900, up from upstream's 100. Gatus bounds the
database by COUNT and trims inline on insert, so there is no retention job and
no way to fill a disk - but history depth is then a function of check frequency,
and 100 results at a 5-minute interval is 8 hours. 900 is ~3 days of liveness
and ~2.5 years of the daily disk check. The uptime table is separate and its
30-day retention is hard-coded upstream.
A bug worth recording: the first deploy shipped five scripts that all died with
"syntax error: unexpected end of file". Jinja strips an included template's
trailing newline and trim_blocks then eats the newline after {% endif %}, so the
closing brace of check() landed on the same line as the body's last statement -
`return 0}`. Every check was broken and the deploy still reported failed=0,
because the role's "run once" task has failed_when: false and reports the result
as a debug message nobody read. The blank line that fixes it is now load-bearing
and commented as such.
Verified by triggering every unit by hand rather than waiting on timers: all
checks exit 0 on all hosts, and Gatus shows 26 UP / 1 DOWN. The one DOWN is
liveness_watchtower, which is honest - that host currently refuses SSH (TCP
connects, no banner exchange) and is excluded from this deploy. It is also the
box still running Uptime Kuma.
Known waste, not yet fixed: healthcheck installs its dependencies per CHECK
rather than per HOST, so apt runs 29 times estate-wide for a curl that is
already present, and daemon_reload runs 4x per host.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 08:54:10 +02:00
|
|
|
healthcheck_script_dir: /usr/local/bin
|
|
|
|
|
healthcheck_log_dir: /var/log/healthchecks
|
|
|
|
|
|
|
|
|
|
# Per-check knobs, consumed by the templates under checks/
|
|
|
|
|
healthcheck_disk_threshold: 85 # percent
|
|
|
|
|
healthcheck_cpu_temp_threshold: 80 # celsius
|
|
|
|
|
healthcheck_zfs_pool: ""
|
|
|
|
|
healthcheck_ups_name: ""
|