personal_infra/ansible/services/gatus/deploy_gatus_playbook.yml
counterweight 3a9e1d5851
alerting: Signal via signal-cli-rest-api, and faster failure detection
Three changes: detection windows tightened, a Signal transport deployed, and
alerts attached to all 84 non-transport endpoints.

── Failing faster ──────────────────────────────────────────────────────────
The heartbeat is a TICKER, not a deadline: Gatus wakes every interval and asks
"did anything arrive in the last interval", so real detection is 1-2x the
window. And a window can only ever be as tight as the push frequency - which is
why two checks moved rather than just having their numbers changed.

  liveness, cpu, ups, service-health, probes   16m -> 11m
  disk, zfs                       daily/30h -> 6-hourly/7h
  backup store + pull job         daily/30h -> 6-hourly/7h
  DNS records                           24h -> 6h
  backup dump                             30h -> 26h

backup-dump stays slow because the dump genuinely is daily. The store-side check
catches the same fault within 6h by reading the source's dump timestamp out of
the artefact filename, so 26h is a backstop rather than the primary signal.

── Thresholds differ by check type, deliberately ───────────────────────────
`failure-threshold` counts consecutive failures, but "consecutive" is a
different amount of wall-clock time per check: a push endpoint produces one
failure per heartbeat window, a pulled one per interval. The default of 3 would
mean 33 minutes on an 11m heartbeat and over a day on a 7h one.

More importantly the heartbeat window ALREADY encodes the tolerance - an 11m
window on a 5-minute push is exactly "one missed push forgiven" - so stacking a
threshold of 3 triples a tolerance that was already chosen. Hence:

  push/heartbeat endpoints   failure-threshold 1
  pulled, 5m (public)        failure-threshold 3   (= 15 minutes)
  pulled, 6h/24h (dns, domain) failure-threshold 1

── The Signal transport ────────────────────────────────────────────────────
roles/signal_api runs signal-cli-rest-api on the observability host, pinned by
digest, MODE=native.

It publishes NO PORTS. The API has no authentication of any kind - anything that
reaches it can send messages as you and read your Signal. Gatus talks to it over
a shared docker network by service name, which is also WHY the network exists:
Gatus runs in a container, so the host's loopback is unreachable from it and a
port published on 127.0.0.1 would not have worked.

MODE=native and not json-rpc because this VPS has 464MB of RAM and already runs
Gatus and Caddy. The json-rpc modes hold a resident JVM; native runs a binary
per request, and alerts are rare enough that startup cost per alert is the right
trade.

Monitored - Gatus polls /v1/health over the same network path the alerts take,
so it proves the delivery route rather than mere container liveness. Deliberately
NOT backed up: the data directory holds Signal private keys and the recovery
path is to link the device again from the phone.

Neither the signal-api endpoint nor Gatus's self-check carries a Signal alert.
If either is down, Signal is precisely what cannot deliver the alert.

── Four traps hit while linking, all now in the role README ────────────────
  * /v1/qrcodelink is BROKEN in native mode - returns "no data to encode" while
    the binary itself emits a perfectly good URI. Not worked around by switching
    MODE, which would put a JVM in the path of every alert permanently.
  * The data dir must be owned by uid 1000, not root. A root-owned 0700 dir
    cannot be traversed by the container user, so linking silently never
    completes and /v1/accounts returns "Failed to read local accounts list".
  * `docker exec` runs as ROOT while the service runs as uid 1000, so without
    --config the account is written to /root/... on the container's ephemeral
    layer. It reports success and is destroyed on the next recreate.
  * The phone reporting "network error" was IPv6: chat.signal.org resolves to
    dualstack AAAA records first, the container has no IPv6 address, and this
    host's IPv6 path is broken - the same edge that 404'd the Go tarball. Fixed
    with a mounted gai.conf that prefers IPv4.

Verified: provider loads (configuredProviders=[signal]), a test message was
delivered and confirmed received, 84 endpoints carry alerts, 86 UP / 0 DOWN.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 21:40:37 +02:00

106 lines
4.6 KiB
YAML

---
# Gatus: health checks, status page and alerting for the whole estate.
#
# Built from source and run under systemd - upstream publishes no binaries, and
# their Dockerfile shows the runtime needs nothing but the static binary and a
# CA bundle. See roles/gatus/README.md.
- name: Deploy Gatus on the observability host
hosts: observability
become: yes
vars:
gatus_alerting:
signal:
# NOTE: the key is `api-url`, not `url` as upstream's own README table
# says - see alerting/provider/signal/signal.go. Gatus appends /v2/send
# itself if the suffix is missing.
#
# Reached by service name over the shared docker network. Gatus runs in
# a container, so 127.0.0.1 here would be the Gatus container, not the
# host - and the Signal API deliberately publishes no ports because it
# has no authentication.
api-url: "http://signal-api:8080"
number: "{{ signal_number }}"
recipients: "{{ signal_recipients }}"
default-alert:
# Overridden per group by the registration playbooks; these are the
# values that apply if a caller sets nothing.
failure-threshold: 1
success-threshold: 2
send-on-resolved: true
# An ongoing outage should not become an ongoing phone buzz.
minimum-reminder-interval: 6h
roles:
- gatus
# The dashboard is bound to loopback; Caddy publishes it.
#
# Auth is done HERE, at the edge, and not with Gatus's own `security.basic`.
# Gatus's security middleware protects exactly four routes (api/api.go):
#
# /api/v1/endpoints/statuses
# /api/v1/endpoints/:key/statuses
# /api/v1/suites/statuses
# /api/v1/suites/:key/statuses
#
# Everything else is registered on the UNPROTECTED router, including
# /api/v1/config, every badge, and - the part that matters -
# /api/v1/endpoints/:key/uptimes/:duration and .../response-times/:duration/history,
# which return real per-endpoint data to anyone who can guess a key. Keys are
# just "<group>_<name>". So Gatus's own auth makes the dashboard render empty
# while leaving the data readable, which is worse than it looks.
#
# The one route that must NOT sit behind basic auth is the external-endpoint
# push API. It authenticates with `Authorization: Bearer <token>`, and basic
# auth wants `Authorization: Basic <...>` - same header, two schemes, and the
# push clients lose. It is not actually unauthenticated: the handler 401s on a
# missing prefix, an empty token, or a token that does not match that endpoint's
# own. Upstream's comment on the route says exactly that.
- name: Publish the Gatus status page through Caddy
hosts: observability
become: yes
tasks:
- name: Require the dashboard credentials to be set
ansible.builtin.assert:
that:
- gatus_dashboard_username is defined
- gatus_dashboard_username | length > 0
- gatus_dashboard_password_hash is defined
- gatus_dashboard_password_hash.startswith('$2')
fail_msg: >-
gatus_dashboard_username and gatus_dashboard_password_hash must be in
the vault. Generate the hash on the observability host, which runs
Caddy natively, so the bcrypt cost and format match what verifies it:
caddy hash-password --plaintext 'your-password'
then: ansible-vault edit group_vars/all/vault.yml
- name: Configure the Caddy vhost for Gatus
ansible.builtin.include_role:
name: caddy_site
vars:
caddy_site_name: gatus
caddy_site_domain: "{{ subdomains.gatus }}.{{ root_domain }}"
# caddy_site_body rather than caddy_site_upstream + caddy_site_basic_auth,
# because that pair applies auth to the whole site with no way to carve
# out the push path. `handle` blocks are mutually exclusive and first
# match wins, so the push API gets a route of its own.
caddy_site_body: |
@push {
path /api/v1/endpoints/*/external
method POST
}
# Push API: Bearer-authenticated by Gatus itself. No basic auth here,
# or the Authorization header collides.
handle @push {
reverse_proxy 127.0.0.1:{{ gatus_port | default(8080) }}
}
# Everything else: the dashboard, the config endpoint, the badges and
# the uptime/response-time history.
handle {
basic_auth {
{{ gatus_dashboard_username }} {{ gatus_dashboard_password_hash }}
}
reverse_proxy 127.0.0.1:{{ gatus_port | default(8080) }}
}