personal_infra/ansible/roles/gatus/templates/docker-compose.yml.j2

54 lines
1.6 KiB
Text
Raw Normal View History

gatus: deploy on prd-monitoring, behind Caddy basic auth First step of replacing Uptime Kuma and ntfy. Gatus runs on the new observability VPS, fronted by Caddy at status.contrapeso.xyz. Deployed as the upstream container image, not built from source. The role did build from source first - their Dockerfile is a bare `CGO_ENABLED=0 go build`, the Vue dashboard is compiled in via `//go:embed static` in web/static.go, and CGO can stay off because the sqlite driver is pure-Go modernc.org/sqlite - but that produces a binary upstream never ran, and it meant compiling the AWS SDK and gRPC on the smallest box in the estate. That load was heavy enough that unrelated Ansible tasks timed out while it ran. The cost of the container is a daemon on the machine whose job is to notice when everything else breaks; that trade is made deliberately and is written down in the role README. Pinned by DIGEST, not tag. A tag is mutable - v5.36.0 can be repushed - so pinning it alone is a weaker promise than it looks: gatus_image: "ghcr.io/twin/gatus@sha256:c5f210d0..." `docker compose pull` now either fetches exactly the reviewed image or fails. gatus_version is kept beside it only so a human can read the release; the two move together. The image is FROM scratch, so it has no /etc/passwd and its default user is root. The container runs as 10001:10001 with the host data dir owned to match, plus read_only, cap_drop ALL, and no-new-privileges. NET_RAW is added back only when gatus_allow_icmp, so the capability for icmp:// checks is a visible grant rather than something inherited from running as root. Config is a DIRECTORY, not a file. Gatus merges every *.yaml under GATUS_CONFIG_PATH - maps deep-merge, lists append - so the role owns 00-base.yaml (web, storage, ui, alerting, security) and each service will drop its own file into endpoints/, the same shape as caddy_site. A primitive defined twice is ambiguous and upstream refuses it, so anything that is not a list lives in the base file and nowhere else. Two bugs the deploy caught: * Gatus panics on a config with no endpoints ("configuration should contain at least one endpoint or suite"), so "install now, add endpoints later" is not a valid state. The role ships endpoints/00-self.yaml checking its own /health. Less circular than it looks: it proves the directory merged, the listener serves, and storage accepted a write. * web.address was carried over from the systemd design as 127.0.0.1. Inside a container that is the CONTAINER's loopback, which docker-proxy cannot reach - gatus came up healthy, self-check passing, while every connection to the published port was refused. It now always binds 0.0.0.0 inside the container; the isolation comes from publishing to 127.0.0.1 on the host. Auth is done at the edge, NOT with Gatus's own security.basic. Reading api/api.go, that middleware protects exactly four routes - the statuses endpoints. Everything else is registered on the unprotected router, including /api/v1/config, every badge, and /api/v1/endpoints/:key/uptimes/:duration and .../response-times/:duration/history, which return real data to anyone who can guess a key ("<group>_<name>"). Verified against the live instance: all seven routes returned 200 unauthenticated, and /uptimes/24h returned "1.000000". So the vhost uses caddy_site_body with a path carve-out rather than caddy_site_basic_auth, which has no way to exempt a path. The external-endpoint push API must NOT sit behind basic auth: it authenticates with `Authorization: Bearer <token>`, and basic auth wants the same header. It is not unauthenticated - the handler 401s on a missing prefix, an empty token, or a token that does not match that endpoint's own. Verified end to end. All seven previously-open routes now 401. The push path distinguishes cleanly: POST with no auth gets Gatus's own "invalid Authorization header" with NO WWW-Authenticate; POST with a bogus Bearer gets 404 (key looked up, no external endpoints yet); GET on the same path gets Caddy's 401 with WWW-Authenticate: Basic, so the exemption is scoped to POST alone. The self-check still passes because it polls localhost inside the container and never traverses Caddy. The host itself was rebuilt from scratch: 01 (ok=9 changed=8), 02 (ok=12 changed=6), 910_docker (--limit, since that playbook still wrongly claims all of `managed` needs Docker), caddy (ok=13 changed=8), gatus (ok=18 changed=2). Not done here: gatus_alerting is still {} - valid, and every condition is evaluated and recorded, there is just nowhere to shout until a provider is chosen to replace ntfy. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 22:29:03 +02:00
# Managed by Ansible (roles/gatus)
services:
gatus:
image: {{ gatus_image }}
container_name: gatus
restart: unless-stopped
# The image is FROM scratch: no /etc/passwd, so there is no user to drop to
# by name and the default is root. Run it by numeric id instead.
user: "{{ gatus_uid }}:{{ gatus_gid }}"
ports:
# Loopback on purpose. Caddy fronts this, and the external-endpoint push
# API is served from the same listener as the dashboard - publishing
# 0.0.0.0 would put both straight on the public internet.
- "{{ gatus_bind_address }}:{{ gatus_port }}:{{ gatus_port }}"
environment:
GATUS_CONFIG_PATH: /config
GATUS_LOG_LEVEL: "{{ gatus_log_level }}"
volumes:
- {{ gatus_config_dir }}:/config:ro
- {{ gatus_data_dir }}:/data
# Hardening. The systemd unit this replaced got most of it from
# ProtectSystem/NoNewPrivileges/etc; these are the container equivalents.
read_only: true
security_opt:
- no-new-privileges:true
cap_drop:
- ALL
{% if gatus_allow_icmp %}
cap_add:
# icmp:// endpoints need raw sockets. Dropped above with ALL, added back
# explicitly so the grant is visible rather than inherited from root.
- NET_RAW
{% endif %}
alerting: Signal via signal-cli-rest-api, and faster failure detection Three changes: detection windows tightened, a Signal transport deployed, and alerts attached to all 84 non-transport endpoints. ── Failing faster ────────────────────────────────────────────────────────── The heartbeat is a TICKER, not a deadline: Gatus wakes every interval and asks "did anything arrive in the last interval", so real detection is 1-2x the window. And a window can only ever be as tight as the push frequency - which is why two checks moved rather than just having their numbers changed. liveness, cpu, ups, service-health, probes 16m -> 11m disk, zfs daily/30h -> 6-hourly/7h backup store + pull job daily/30h -> 6-hourly/7h DNS records 24h -> 6h backup dump 30h -> 26h backup-dump stays slow because the dump genuinely is daily. The store-side check catches the same fault within 6h by reading the source's dump timestamp out of the artefact filename, so 26h is a backstop rather than the primary signal. ── Thresholds differ by check type, deliberately ─────────────────────────── `failure-threshold` counts consecutive failures, but "consecutive" is a different amount of wall-clock time per check: a push endpoint produces one failure per heartbeat window, a pulled one per interval. The default of 3 would mean 33 minutes on an 11m heartbeat and over a day on a 7h one. More importantly the heartbeat window ALREADY encodes the tolerance - an 11m window on a 5-minute push is exactly "one missed push forgiven" - so stacking a threshold of 3 triples a tolerance that was already chosen. Hence: push/heartbeat endpoints failure-threshold 1 pulled, 5m (public) failure-threshold 3 (= 15 minutes) pulled, 6h/24h (dns, domain) failure-threshold 1 ── The Signal transport ──────────────────────────────────────────────────── roles/signal_api runs signal-cli-rest-api on the observability host, pinned by digest, MODE=native. It publishes NO PORTS. The API has no authentication of any kind - anything that reaches it can send messages as you and read your Signal. Gatus talks to it over a shared docker network by service name, which is also WHY the network exists: Gatus runs in a container, so the host's loopback is unreachable from it and a port published on 127.0.0.1 would not have worked. MODE=native and not json-rpc because this VPS has 464MB of RAM and already runs Gatus and Caddy. The json-rpc modes hold a resident JVM; native runs a binary per request, and alerts are rare enough that startup cost per alert is the right trade. Monitored - Gatus polls /v1/health over the same network path the alerts take, so it proves the delivery route rather than mere container liveness. Deliberately NOT backed up: the data directory holds Signal private keys and the recovery path is to link the device again from the phone. Neither the signal-api endpoint nor Gatus's self-check carries a Signal alert. If either is down, Signal is precisely what cannot deliver the alert. ── Four traps hit while linking, all now in the role README ──────────────── * /v1/qrcodelink is BROKEN in native mode - returns "no data to encode" while the binary itself emits a perfectly good URI. Not worked around by switching MODE, which would put a JVM in the path of every alert permanently. * The data dir must be owned by uid 1000, not root. A root-owned 0700 dir cannot be traversed by the container user, so linking silently never completes and /v1/accounts returns "Failed to read local accounts list". * `docker exec` runs as ROOT while the service runs as uid 1000, so without --config the account is written to /root/... on the container's ephemeral layer. It reports success and is destroyed on the next recreate. * The phone reporting "network error" was IPv6: chat.signal.org resolves to dualstack AAAA records first, the container has no IPv6 address, and this host's IPv6 path is broken - the same edge that 404'd the Go tarball. Fixed with a mounted gai.conf that prefers IPv4. Verified: provider loads (configuredProviders=[signal]), a test message was delivered and confirmed received, 84 endpoints carry alerts, 86 UP / 0 DOWN. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 21:40:37 +02:00
networks:
# Shared with signal-api, so alerts can be delivered by service name.
# 127.0.0.1 inside this container is the container, not the host.
- {{ gatus_network }}
gatus: deploy on prd-monitoring, behind Caddy basic auth First step of replacing Uptime Kuma and ntfy. Gatus runs on the new observability VPS, fronted by Caddy at status.contrapeso.xyz. Deployed as the upstream container image, not built from source. The role did build from source first - their Dockerfile is a bare `CGO_ENABLED=0 go build`, the Vue dashboard is compiled in via `//go:embed static` in web/static.go, and CGO can stay off because the sqlite driver is pure-Go modernc.org/sqlite - but that produces a binary upstream never ran, and it meant compiling the AWS SDK and gRPC on the smallest box in the estate. That load was heavy enough that unrelated Ansible tasks timed out while it ran. The cost of the container is a daemon on the machine whose job is to notice when everything else breaks; that trade is made deliberately and is written down in the role README. Pinned by DIGEST, not tag. A tag is mutable - v5.36.0 can be repushed - so pinning it alone is a weaker promise than it looks: gatus_image: "ghcr.io/twin/gatus@sha256:c5f210d0..." `docker compose pull` now either fetches exactly the reviewed image or fails. gatus_version is kept beside it only so a human can read the release; the two move together. The image is FROM scratch, so it has no /etc/passwd and its default user is root. The container runs as 10001:10001 with the host data dir owned to match, plus read_only, cap_drop ALL, and no-new-privileges. NET_RAW is added back only when gatus_allow_icmp, so the capability for icmp:// checks is a visible grant rather than something inherited from running as root. Config is a DIRECTORY, not a file. Gatus merges every *.yaml under GATUS_CONFIG_PATH - maps deep-merge, lists append - so the role owns 00-base.yaml (web, storage, ui, alerting, security) and each service will drop its own file into endpoints/, the same shape as caddy_site. A primitive defined twice is ambiguous and upstream refuses it, so anything that is not a list lives in the base file and nowhere else. Two bugs the deploy caught: * Gatus panics on a config with no endpoints ("configuration should contain at least one endpoint or suite"), so "install now, add endpoints later" is not a valid state. The role ships endpoints/00-self.yaml checking its own /health. Less circular than it looks: it proves the directory merged, the listener serves, and storage accepted a write. * web.address was carried over from the systemd design as 127.0.0.1. Inside a container that is the CONTAINER's loopback, which docker-proxy cannot reach - gatus came up healthy, self-check passing, while every connection to the published port was refused. It now always binds 0.0.0.0 inside the container; the isolation comes from publishing to 127.0.0.1 on the host. Auth is done at the edge, NOT with Gatus's own security.basic. Reading api/api.go, that middleware protects exactly four routes - the statuses endpoints. Everything else is registered on the unprotected router, including /api/v1/config, every badge, and /api/v1/endpoints/:key/uptimes/:duration and .../response-times/:duration/history, which return real data to anyone who can guess a key ("<group>_<name>"). Verified against the live instance: all seven routes returned 200 unauthenticated, and /uptimes/24h returned "1.000000". So the vhost uses caddy_site_body with a path carve-out rather than caddy_site_basic_auth, which has no way to exempt a path. The external-endpoint push API must NOT sit behind basic auth: it authenticates with `Authorization: Bearer <token>`, and basic auth wants the same header. It is not unauthenticated - the handler 401s on a missing prefix, an empty token, or a token that does not match that endpoint's own. Verified end to end. All seven previously-open routes now 401. The push path distinguishes cleanly: POST with no auth gets Gatus's own "invalid Authorization header" with NO WWW-Authenticate; POST with a bogus Bearer gets 404 (key looked up, no external endpoints yet); GET on the same path gets Caddy's 401 with WWW-Authenticate: Basic, so the exemption is scoped to POST alone. The self-check still passes because it polls localhost inside the container and never traverses Caddy. The host itself was rebuilt from scratch: 01 (ok=9 changed=8), 02 (ok=12 changed=6), 910_docker (--limit, since that playbook still wrongly claims all of `managed` needs Docker), caddy (ok=13 changed=8), gatus (ok=18 changed=2). Not done here: gatus_alerting is still {} - valid, and every condition is evaluated and recorded, there is just nowhere to shout until a provider is chosen to replace ntfy. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 22:29:03 +02:00
logging:
driver: json-file
options:
max-size: "10m"
max-file: "3"
alerting: Signal via signal-cli-rest-api, and faster failure detection Three changes: detection windows tightened, a Signal transport deployed, and alerts attached to all 84 non-transport endpoints. ── Failing faster ────────────────────────────────────────────────────────── The heartbeat is a TICKER, not a deadline: Gatus wakes every interval and asks "did anything arrive in the last interval", so real detection is 1-2x the window. And a window can only ever be as tight as the push frequency - which is why two checks moved rather than just having their numbers changed. liveness, cpu, ups, service-health, probes 16m -> 11m disk, zfs daily/30h -> 6-hourly/7h backup store + pull job daily/30h -> 6-hourly/7h DNS records 24h -> 6h backup dump 30h -> 26h backup-dump stays slow because the dump genuinely is daily. The store-side check catches the same fault within 6h by reading the source's dump timestamp out of the artefact filename, so 26h is a backstop rather than the primary signal. ── Thresholds differ by check type, deliberately ─────────────────────────── `failure-threshold` counts consecutive failures, but "consecutive" is a different amount of wall-clock time per check: a push endpoint produces one failure per heartbeat window, a pulled one per interval. The default of 3 would mean 33 minutes on an 11m heartbeat and over a day on a 7h one. More importantly the heartbeat window ALREADY encodes the tolerance - an 11m window on a 5-minute push is exactly "one missed push forgiven" - so stacking a threshold of 3 triples a tolerance that was already chosen. Hence: push/heartbeat endpoints failure-threshold 1 pulled, 5m (public) failure-threshold 3 (= 15 minutes) pulled, 6h/24h (dns, domain) failure-threshold 1 ── The Signal transport ──────────────────────────────────────────────────── roles/signal_api runs signal-cli-rest-api on the observability host, pinned by digest, MODE=native. It publishes NO PORTS. The API has no authentication of any kind - anything that reaches it can send messages as you and read your Signal. Gatus talks to it over a shared docker network by service name, which is also WHY the network exists: Gatus runs in a container, so the host's loopback is unreachable from it and a port published on 127.0.0.1 would not have worked. MODE=native and not json-rpc because this VPS has 464MB of RAM and already runs Gatus and Caddy. The json-rpc modes hold a resident JVM; native runs a binary per request, and alerts are rare enough that startup cost per alert is the right trade. Monitored - Gatus polls /v1/health over the same network path the alerts take, so it proves the delivery route rather than mere container liveness. Deliberately NOT backed up: the data directory holds Signal private keys and the recovery path is to link the device again from the phone. Neither the signal-api endpoint nor Gatus's self-check carries a Signal alert. If either is down, Signal is precisely what cannot deliver the alert. ── Four traps hit while linking, all now in the role README ──────────────── * /v1/qrcodelink is BROKEN in native mode - returns "no data to encode" while the binary itself emits a perfectly good URI. Not worked around by switching MODE, which would put a JVM in the path of every alert permanently. * The data dir must be owned by uid 1000, not root. A root-owned 0700 dir cannot be traversed by the container user, so linking silently never completes and /v1/accounts returns "Failed to read local accounts list". * `docker exec` runs as ROOT while the service runs as uid 1000, so without --config the account is written to /root/... on the container's ephemeral layer. It reports success and is destroyed on the next recreate. * The phone reporting "network error" was IPv6: chat.signal.org resolves to dualstack AAAA records first, the container has no IPv6 address, and this host's IPv6 path is broken - the same edge that 404'd the Go tarball. Fixed with a mounted gai.conf that prefers IPv4. Verified: provider loads (configuredProviders=[signal]), a test message was delivered and confirmed received, 84 endpoints carry alerts, 86 UP / 0 DOWN. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 21:40:37 +02:00
networks:
{{ gatus_network }}:
external: true