Three changes: detection windows tightened, a Signal transport deployed, and
alerts attached to all 84 non-transport endpoints.
── Failing faster ──────────────────────────────────────────────────────────
The heartbeat is a TICKER, not a deadline: Gatus wakes every interval and asks
"did anything arrive in the last interval", so real detection is 1-2x the
window. And a window can only ever be as tight as the push frequency - which is
why two checks moved rather than just having their numbers changed.
liveness, cpu, ups, service-health, probes 16m -> 11m
disk, zfs daily/30h -> 6-hourly/7h
backup store + pull job daily/30h -> 6-hourly/7h
DNS records 24h -> 6h
backup dump 30h -> 26h
backup-dump stays slow because the dump genuinely is daily. The store-side check
catches the same fault within 6h by reading the source's dump timestamp out of
the artefact filename, so 26h is a backstop rather than the primary signal.
── Thresholds differ by check type, deliberately ───────────────────────────
`failure-threshold` counts consecutive failures, but "consecutive" is a
different amount of wall-clock time per check: a push endpoint produces one
failure per heartbeat window, a pulled one per interval. The default of 3 would
mean 33 minutes on an 11m heartbeat and over a day on a 7h one.
More importantly the heartbeat window ALREADY encodes the tolerance - an 11m
window on a 5-minute push is exactly "one missed push forgiven" - so stacking a
threshold of 3 triples a tolerance that was already chosen. Hence:
push/heartbeat endpoints failure-threshold 1
pulled, 5m (public) failure-threshold 3 (= 15 minutes)
pulled, 6h/24h (dns, domain) failure-threshold 1
── The Signal transport ────────────────────────────────────────────────────
roles/signal_api runs signal-cli-rest-api on the observability host, pinned by
digest, MODE=native.
It publishes NO PORTS. The API has no authentication of any kind - anything that
reaches it can send messages as you and read your Signal. Gatus talks to it over
a shared docker network by service name, which is also WHY the network exists:
Gatus runs in a container, so the host's loopback is unreachable from it and a
port published on 127.0.0.1 would not have worked.
MODE=native and not json-rpc because this VPS has 464MB of RAM and already runs
Gatus and Caddy. The json-rpc modes hold a resident JVM; native runs a binary
per request, and alerts are rare enough that startup cost per alert is the right
trade.
Monitored - Gatus polls /v1/health over the same network path the alerts take,
so it proves the delivery route rather than mere container liveness. Deliberately
NOT backed up: the data directory holds Signal private keys and the recovery
path is to link the device again from the phone.
Neither the signal-api endpoint nor Gatus's self-check carries a Signal alert.
If either is down, Signal is precisely what cannot deliver the alert.
── Four traps hit while linking, all now in the role README ────────────────
* /v1/qrcodelink is BROKEN in native mode - returns "no data to encode" while
the binary itself emits a perfectly good URI. Not worked around by switching
MODE, which would put a JVM in the path of every alert permanently.
* The data dir must be owned by uid 1000, not root. A root-owned 0700 dir
cannot be traversed by the container user, so linking silently never
completes and /v1/accounts returns "Failed to read local accounts list".
* `docker exec` runs as ROOT while the service runs as uid 1000, so without
--config the account is written to /root/... on the container's ephemeral
layer. It reports success and is destroyed on the next recreate.
* The phone reporting "network error" was IPv6: chat.signal.org resolves to
dualstack AAAA records first, the container has no IPv6 address, and this
host's IPv6 path is broken - the same edge that 404'd the Go tarball. Fixed
with a mounted gai.conf that prefers IPv4.
Verified: provider loads (configuredProviders=[signal]), a test message was
delivered and confirmed received, 84 endpoints carry alerts, 86 UP / 0 DOWN.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
106 lines
4.6 KiB
YAML
106 lines
4.6 KiB
YAML
---
|
|
# Gatus: health checks, status page and alerting for the whole estate.
|
|
#
|
|
# Built from source and run under systemd - upstream publishes no binaries, and
|
|
# their Dockerfile shows the runtime needs nothing but the static binary and a
|
|
# CA bundle. See roles/gatus/README.md.
|
|
- name: Deploy Gatus on the observability host
|
|
hosts: observability
|
|
become: yes
|
|
vars:
|
|
gatus_alerting:
|
|
signal:
|
|
# NOTE: the key is `api-url`, not `url` as upstream's own README table
|
|
# says - see alerting/provider/signal/signal.go. Gatus appends /v2/send
|
|
# itself if the suffix is missing.
|
|
#
|
|
# Reached by service name over the shared docker network. Gatus runs in
|
|
# a container, so 127.0.0.1 here would be the Gatus container, not the
|
|
# host - and the Signal API deliberately publishes no ports because it
|
|
# has no authentication.
|
|
api-url: "http://signal-api:8080"
|
|
number: "{{ signal_number }}"
|
|
recipients: "{{ signal_recipients }}"
|
|
default-alert:
|
|
# Overridden per group by the registration playbooks; these are the
|
|
# values that apply if a caller sets nothing.
|
|
failure-threshold: 1
|
|
success-threshold: 2
|
|
send-on-resolved: true
|
|
# An ongoing outage should not become an ongoing phone buzz.
|
|
minimum-reminder-interval: 6h
|
|
|
|
roles:
|
|
- gatus
|
|
|
|
# The dashboard is bound to loopback; Caddy publishes it.
|
|
#
|
|
# Auth is done HERE, at the edge, and not with Gatus's own `security.basic`.
|
|
# Gatus's security middleware protects exactly four routes (api/api.go):
|
|
#
|
|
# /api/v1/endpoints/statuses
|
|
# /api/v1/endpoints/:key/statuses
|
|
# /api/v1/suites/statuses
|
|
# /api/v1/suites/:key/statuses
|
|
#
|
|
# Everything else is registered on the UNPROTECTED router, including
|
|
# /api/v1/config, every badge, and - the part that matters -
|
|
# /api/v1/endpoints/:key/uptimes/:duration and .../response-times/:duration/history,
|
|
# which return real per-endpoint data to anyone who can guess a key. Keys are
|
|
# just "<group>_<name>". So Gatus's own auth makes the dashboard render empty
|
|
# while leaving the data readable, which is worse than it looks.
|
|
#
|
|
# The one route that must NOT sit behind basic auth is the external-endpoint
|
|
# push API. It authenticates with `Authorization: Bearer <token>`, and basic
|
|
# auth wants `Authorization: Basic <...>` - same header, two schemes, and the
|
|
# push clients lose. It is not actually unauthenticated: the handler 401s on a
|
|
# missing prefix, an empty token, or a token that does not match that endpoint's
|
|
# own. Upstream's comment on the route says exactly that.
|
|
- name: Publish the Gatus status page through Caddy
|
|
hosts: observability
|
|
become: yes
|
|
tasks:
|
|
- name: Require the dashboard credentials to be set
|
|
ansible.builtin.assert:
|
|
that:
|
|
- gatus_dashboard_username is defined
|
|
- gatus_dashboard_username | length > 0
|
|
- gatus_dashboard_password_hash is defined
|
|
- gatus_dashboard_password_hash.startswith('$2')
|
|
fail_msg: >-
|
|
gatus_dashboard_username and gatus_dashboard_password_hash must be in
|
|
the vault. Generate the hash on the observability host, which runs
|
|
Caddy natively, so the bcrypt cost and format match what verifies it:
|
|
caddy hash-password --plaintext 'your-password'
|
|
then: ansible-vault edit group_vars/all/vault.yml
|
|
|
|
- name: Configure the Caddy vhost for Gatus
|
|
ansible.builtin.include_role:
|
|
name: caddy_site
|
|
vars:
|
|
caddy_site_name: gatus
|
|
caddy_site_domain: "{{ subdomains.gatus }}.{{ root_domain }}"
|
|
# caddy_site_body rather than caddy_site_upstream + caddy_site_basic_auth,
|
|
# because that pair applies auth to the whole site with no way to carve
|
|
# out the push path. `handle` blocks are mutually exclusive and first
|
|
# match wins, so the push API gets a route of its own.
|
|
caddy_site_body: |
|
|
@push {
|
|
path /api/v1/endpoints/*/external
|
|
method POST
|
|
}
|
|
|
|
# Push API: Bearer-authenticated by Gatus itself. No basic auth here,
|
|
# or the Authorization header collides.
|
|
handle @push {
|
|
reverse_proxy 127.0.0.1:{{ gatus_port | default(8080) }}
|
|
}
|
|
|
|
# Everything else: the dashboard, the config endpoint, the badges and
|
|
# the uptime/response-time history.
|
|
handle {
|
|
basic_auth {
|
|
{{ gatus_dashboard_username }} {{ gatus_dashboard_password_hash }}
|
|
}
|
|
reverse_proxy 127.0.0.1:{{ gatus_port | default(8080) }}
|
|
}
|