Three changes: detection windows tightened, a Signal transport deployed, and
alerts attached to all 84 non-transport endpoints.
── Failing faster ──────────────────────────────────────────────────────────
The heartbeat is a TICKER, not a deadline: Gatus wakes every interval and asks
"did anything arrive in the last interval", so real detection is 1-2x the
window. And a window can only ever be as tight as the push frequency - which is
why two checks moved rather than just having their numbers changed.
liveness, cpu, ups, service-health, probes 16m -> 11m
disk, zfs daily/30h -> 6-hourly/7h
backup store + pull job daily/30h -> 6-hourly/7h
DNS records 24h -> 6h
backup dump 30h -> 26h
backup-dump stays slow because the dump genuinely is daily. The store-side check
catches the same fault within 6h by reading the source's dump timestamp out of
the artefact filename, so 26h is a backstop rather than the primary signal.
── Thresholds differ by check type, deliberately ───────────────────────────
`failure-threshold` counts consecutive failures, but "consecutive" is a
different amount of wall-clock time per check: a push endpoint produces one
failure per heartbeat window, a pulled one per interval. The default of 3 would
mean 33 minutes on an 11m heartbeat and over a day on a 7h one.
More importantly the heartbeat window ALREADY encodes the tolerance - an 11m
window on a 5-minute push is exactly "one missed push forgiven" - so stacking a
threshold of 3 triples a tolerance that was already chosen. Hence:
push/heartbeat endpoints failure-threshold 1
pulled, 5m (public) failure-threshold 3 (= 15 minutes)
pulled, 6h/24h (dns, domain) failure-threshold 1
── The Signal transport ────────────────────────────────────────────────────
roles/signal_api runs signal-cli-rest-api on the observability host, pinned by
digest, MODE=native.
It publishes NO PORTS. The API has no authentication of any kind - anything that
reaches it can send messages as you and read your Signal. Gatus talks to it over
a shared docker network by service name, which is also WHY the network exists:
Gatus runs in a container, so the host's loopback is unreachable from it and a
port published on 127.0.0.1 would not have worked.
MODE=native and not json-rpc because this VPS has 464MB of RAM and already runs
Gatus and Caddy. The json-rpc modes hold a resident JVM; native runs a binary
per request, and alerts are rare enough that startup cost per alert is the right
trade.
Monitored - Gatus polls /v1/health over the same network path the alerts take,
so it proves the delivery route rather than mere container liveness. Deliberately
NOT backed up: the data directory holds Signal private keys and the recovery
path is to link the device again from the phone.
Neither the signal-api endpoint nor Gatus's self-check carries a Signal alert.
If either is down, Signal is precisely what cannot deliver the alert.
── Four traps hit while linking, all now in the role README ────────────────
* /v1/qrcodelink is BROKEN in native mode - returns "no data to encode" while
the binary itself emits a perfectly good URI. Not worked around by switching
MODE, which would put a JVM in the path of every alert permanently.
* The data dir must be owned by uid 1000, not root. A root-owned 0700 dir
cannot be traversed by the container user, so linking silently never
completes and /v1/accounts returns "Failed to read local accounts list".
* `docker exec` runs as ROOT while the service runs as uid 1000, so without
--config the account is written to /root/... on the container's ephemeral
layer. It reports success and is destroyed on the next recreate.
* The phone reporting "network error" was IPv6: chat.signal.org resolves to
dualstack AAAA records first, the container has no IPv6 address, and this
host's IPv6 path is broken - the same edge that 404'd the Go tarball. Fixed
with a mounted gai.conf that prefers IPv4.
Verified: provider loads (configuredProviders=[signal]), a test message was
delivered and confirmed received, 84 endpoints carry alerts, 86 UP / 0 DOWN.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|---|---|---|
| .. | ||
| defaults | ||
| handlers | ||
| tasks | ||
| templates | ||
| README.md | ||
gatus
Deploys Gatus — health checks, a status page,
and alerting — as the upstream container image, on the
observability group.
Why the container
Upstream publishes no binary release assets. The image is the only artefact
they ship and therefore the only one they test, so it is what docker run in
their README gets you, and it is what this role deploys.
Building from source is entirely possible — their Dockerfile is a bare
CGO_ENABLED=0 go build, the Vue dashboard is compiled in via //go:embed static in web/static.go, and the sqlite driver is pure-Go
modernc.org/sqlite so nothing needs linking. This role did that at first. The
reasons it doesn't now:
- It produces a binary upstream never ran.
- It means compiling the AWS SDK, gRPC and the Google API libraries on the smallest box in the estate. On this VPS that load was heavy enough that unrelated Ansible tasks timed out while it ran.
The cost of the container is a daemon on the machine whose job is to notice when everything else breaks. That is a real trade, made deliberately.
Pinned by digest, not by tag
gatus_image_digest: "sha256:c5f210d0…"
gatus_image: "ghcr.io/twin/gatus@{{ gatus_image_digest }}"
A tag is mutable — v5.36.0 can be repushed — so pinning the tag alone is a
weaker promise than it looks. The digest is a content address: if it resolves,
it is byte-for-byte the image reviewed here, and docker compose pull either
fetches exactly that or fails. gatus_version is kept alongside it purely so a
human can read which release it is; the two must be updated together.
What FROM scratch means for running it
The image has no /etc/passwd, so there is no user to drop to by name and the
default is root. The compose file runs it by numeric id (10001:10001) and the
data directory on the host is owned to match. Everything else is locked down to
approximate what the systemd unit used to do natively:
| systemd | compose |
|---|---|
ProtectSystem=strict |
read_only: true |
NoNewPrivileges=true |
security_opt: [no-new-privileges:true] |
CapabilityBoundingSet= |
cap_drop: [ALL] |
AmbientCapabilities=CAP_NET_RAW |
cap_add: [NET_RAW] (for icmp://) |
Configuration is a directory, not a file
GATUS_CONFIG_PATH points at /opt/gatus/config, and Gatus merges every *.yaml
underneath it — maps deep-merge, lists append. This role owns exactly one file:
/opt/gatus/config/00-base.yaml web, storage, ui, alerting, security (this role)
/opt/gatus/config/endpoints/*.yaml one file per service (gatus_endpoint)
A primitive defined in two files is ambiguous and upstream refuses it. So
anything that is not a list belongs in 00-base.yaml and nowhere else. Endpoints
are lists, so each service's file appends cleanly — the same shape as
caddy_site, where each service contributes its own vhost.
Pull and push
Gatus polls. For anything with a reachable HTTP or TCP surface that is the
better check, because it tests the path a user actually takes. The monitoring
host joins the headscale mesh via infra/920, so internal boxes are reachable
by MagicDNS name and can be polled directly rather than having to report in.
For state with no pollable surface — ZFS pool health, UPS mains status, disk usage, backup freshness — Gatus has external endpoints, a push API:
POST /api/v1/endpoints/{group}_{name}/external?success=true&error=&duration=
Authorization: Bearer <token>
with heartbeat.interval to alert when nothing reports in. That is the same
shape as the generic healthcheck_push_url already wired into every service
role, so those scripts need a URL, a POST, and an auth header — not a rewrite.
Variables
See defaults/main.yml. The ones that matter:
| Variable | Default | Note |
|---|---|---|
gatus_version / gatus_image_digest |
v5.36.0 / sha256:c5f210d0… |
must move together |
gatus_bind_address |
127.0.0.1 |
never bind publicly — the push API shares this listener |
gatus_storage_type |
sqlite |
memory loses all history on restart |
gatus_alerting |
{} |
pass-through; any provider Gatus supports |
gatus_allow_icmp |
true |
adds back NET_RAW for icmp:// checks |
gatus_alerting empty is valid and is the current state: every condition is
still evaluated and recorded, there is just nowhere to shout yet.
Verifying
docker ps --filter name=gatus
docker logs gatus --tail 50
curl -s localhost:8080/health
ls /opt/gatus/config/endpoints/