personal_infra/ansible/roles/gatus
counterweight ede407ebe4
monitoring: recover host checks for the whole estate, reported to Gatus
Five checks, 27 endpoints, replacing what Uptime Kuma used to watch:

  is it up      every 5min, all hosts
  is disk full  daily,      all hosts
  is CPU hot    every 5min, nodito
  is ZFS broken daily,      nodito
  is UPS online every 5min, nodito

Two roles, kept separate so neither knows about the other - they meet at a URL
and a token, the same way caddy_site and each service meet at a vhost:

  roles/gatus_endpoint  runs on the observability host, writes ONE file into
                        /opt/gatus/config/endpoints/. Gatus merges every *.yaml
                        there and appends lists, so callers compose without
                        coordinating.
  roles/healthcheck     runs on the monitored host: a check script, a systemd
                        service, a timer, and an optional push. Ships a library
                        of check bodies under templates/checks/.

Everything PUSHES. Gatus never reaches out, which matters because nodito and its
VMs are behind NAT, and because four of the five checks are internal state with
no pollable surface at all. Liveness pushes too, deliberately: a heartbeat proves
the host is running AND can reach the internet, where an ICMP probe from one
vantage point only proves it answers pings from there. And since Gatus alerts
when a heartbeat window expires, a check that stops running raises the alarm by
itself - a dead timer looks exactly like a dead host, which is the correct
reading.

One bearer token per host, generated straight into the vault and never printed.
A token only writes results for its own host's endpoints, so a compromised host
can lie about itself, which it could do anyway.

Three things learned from the source that shaped this:

  * Gatus polls its own config every 30s and reloads
    (main.listenToConfigurationFileChanges), so gatus_endpoint needs no restart
    handler - writing the file IS the deploy.
  * ...but on a reload it panics if the new config fails to parse, unless
    skip-invalid-config-update is set. Endpoint files are contributed by other
    playbooks, so one malformed file would take the monitor down at the worst
    possible moment. Now set.
  * The push URL uses a key Gatus computes, not the name you write:
    sanitize(group) + "_" + sanitize(name), lowercased with / _ . , space # + &
    replaced by "-" (config/key/key.go). So knots_box_local is
    knots-box-local in the URL. The playbook derives it rather than hand-writing.

storage: maximum-number-of-results 900, up from upstream's 100. Gatus bounds the
database by COUNT and trims inline on insert, so there is no retention job and
no way to fill a disk - but history depth is then a function of check frequency,
and 100 results at a 5-minute interval is 8 hours. 900 is ~3 days of liveness
and ~2.5 years of the daily disk check. The uptime table is separate and its
30-day retention is hard-coded upstream.

A bug worth recording: the first deploy shipped five scripts that all died with
"syntax error: unexpected end of file". Jinja strips an included template's
trailing newline and trim_blocks then eats the newline after {% endif %}, so the
closing brace of check() landed on the same line as the body's last statement -
`return 0}`. Every check was broken and the deploy still reported failed=0,
because the role's "run once" task has failed_when: false and reports the result
as a debug message nobody read. The blank line that fixes it is now load-bearing
and commented as such.

Verified by triggering every unit by hand rather than waiting on timers: all
checks exit 0 on all hosts, and Gatus shows 26 UP / 1 DOWN. The one DOWN is
liveness_watchtower, which is honest - that host currently refuses SSH (TCP
connects, no banner exchange) and is excluded from this deploy. It is also the
box still running Uptime Kuma.

Known waste, not yet fixed: healthcheck installs its dependencies per CHECK
rather than per HOST, so apt runs 29 times estate-wide for a curl that is
already present, and daemon_reload runs 4x per host.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 08:54:10 +02:00
..
defaults monitoring: recover host checks for the whole estate, reported to Gatus 2026-09-14 08:54:10 +02:00
handlers gatus: deploy on prd-monitoring, behind Caddy basic auth 2026-09-13 22:29:03 +02:00
tasks gatus: deploy on prd-monitoring, behind Caddy basic auth 2026-09-13 22:29:03 +02:00
templates monitoring: recover host checks for the whole estate, reported to Gatus 2026-09-14 08:54:10 +02:00
README.md gatus: deploy on prd-monitoring, behind Caddy basic auth 2026-09-13 22:29:03 +02:00

gatus

Deploys Gatus — health checks, a status page, and alerting — as the upstream container image, on the observability group.

Why the container

Upstream publishes no binary release assets. The image is the only artefact they ship and therefore the only one they test, so it is what docker run in their README gets you, and it is what this role deploys.

Building from source is entirely possible — their Dockerfile is a bare CGO_ENABLED=0 go build, the Vue dashboard is compiled in via //go:embed static in web/static.go, and the sqlite driver is pure-Go modernc.org/sqlite so nothing needs linking. This role did that at first. The reasons it doesn't now:

  • It produces a binary upstream never ran.
  • It means compiling the AWS SDK, gRPC and the Google API libraries on the smallest box in the estate. On this VPS that load was heavy enough that unrelated Ansible tasks timed out while it ran.

The cost of the container is a daemon on the machine whose job is to notice when everything else breaks. That is a real trade, made deliberately.

Pinned by digest, not by tag

gatus_image_digest: "sha256:c5f210d0…"
gatus_image: "ghcr.io/twin/gatus@{{ gatus_image_digest }}"

A tag is mutable — v5.36.0 can be repushed — so pinning the tag alone is a weaker promise than it looks. The digest is a content address: if it resolves, it is byte-for-byte the image reviewed here, and docker compose pull either fetches exactly that or fails. gatus_version is kept alongside it purely so a human can read which release it is; the two must be updated together.

What FROM scratch means for running it

The image has no /etc/passwd, so there is no user to drop to by name and the default is root. The compose file runs it by numeric id (10001:10001) and the data directory on the host is owned to match. Everything else is locked down to approximate what the systemd unit used to do natively:

systemd compose
ProtectSystem=strict read_only: true
NoNewPrivileges=true security_opt: [no-new-privileges:true]
CapabilityBoundingSet= cap_drop: [ALL]
AmbientCapabilities=CAP_NET_RAW cap_add: [NET_RAW] (for icmp://)

Configuration is a directory, not a file

GATUS_CONFIG_PATH points at /opt/gatus/config, and Gatus merges every *.yaml underneath it — maps deep-merge, lists append. This role owns exactly one file:

/opt/gatus/config/00-base.yaml      web, storage, ui, alerting, security   (this role)
/opt/gatus/config/endpoints/*.yaml  one file per service                   (gatus_endpoint)

A primitive defined in two files is ambiguous and upstream refuses it. So anything that is not a list belongs in 00-base.yaml and nowhere else. Endpoints are lists, so each service's file appends cleanly — the same shape as caddy_site, where each service contributes its own vhost.

Pull and push

Gatus polls. For anything with a reachable HTTP or TCP surface that is the better check, because it tests the path a user actually takes. The monitoring host joins the headscale mesh via infra/920, so internal boxes are reachable by MagicDNS name and can be polled directly rather than having to report in.

For state with no pollable surface — ZFS pool health, UPS mains status, disk usage, backup freshness — Gatus has external endpoints, a push API:

POST /api/v1/endpoints/{group}_{name}/external?success=true&error=&duration=
Authorization: Bearer <token>

with heartbeat.interval to alert when nothing reports in. That is the same shape as the generic healthcheck_push_url already wired into every service role, so those scripts need a URL, a POST, and an auth header — not a rewrite.

Variables

See defaults/main.yml. The ones that matter:

Variable Default Note
gatus_version / gatus_image_digest v5.36.0 / sha256:c5f210d0… must move together
gatus_bind_address 127.0.0.1 never bind publicly — the push API shares this listener
gatus_storage_type sqlite memory loses all history on restart
gatus_alerting {} pass-through; any provider Gatus supports
gatus_allow_icmp true adds back NET_RAW for icmp:// checks

gatus_alerting empty is valid and is the current state: every condition is still evaluated and recorded, there is just nowhere to shout yet.

Verifying

docker ps --filter name=gatus
docker logs gatus --tail 50
curl -s localhost:8080/health
ls /opt/gatus/config/endpoints/