personal_infra/ansible/infra/401_service_monitoring.yml
counterweight 3a9e1d5851
alerting: Signal via signal-cli-rest-api, and faster failure detection
Three changes: detection windows tightened, a Signal transport deployed, and
alerts attached to all 84 non-transport endpoints.

── Failing faster ──────────────────────────────────────────────────────────
The heartbeat is a TICKER, not a deadline: Gatus wakes every interval and asks
"did anything arrive in the last interval", so real detection is 1-2x the
window. And a window can only ever be as tight as the push frequency - which is
why two checks moved rather than just having their numbers changed.

  liveness, cpu, ups, service-health, probes   16m -> 11m
  disk, zfs                       daily/30h -> 6-hourly/7h
  backup store + pull job         daily/30h -> 6-hourly/7h
  DNS records                           24h -> 6h
  backup dump                             30h -> 26h

backup-dump stays slow because the dump genuinely is daily. The store-side check
catches the same fault within 6h by reading the source's dump timestamp out of
the artefact filename, so 26h is a backstop rather than the primary signal.

── Thresholds differ by check type, deliberately ───────────────────────────
`failure-threshold` counts consecutive failures, but "consecutive" is a
different amount of wall-clock time per check: a push endpoint produces one
failure per heartbeat window, a pulled one per interval. The default of 3 would
mean 33 minutes on an 11m heartbeat and over a day on a 7h one.

More importantly the heartbeat window ALREADY encodes the tolerance - an 11m
window on a 5-minute push is exactly "one missed push forgiven" - so stacking a
threshold of 3 triples a tolerance that was already chosen. Hence:

  push/heartbeat endpoints   failure-threshold 1
  pulled, 5m (public)        failure-threshold 3   (= 15 minutes)
  pulled, 6h/24h (dns, domain) failure-threshold 1

── The Signal transport ────────────────────────────────────────────────────
roles/signal_api runs signal-cli-rest-api on the observability host, pinned by
digest, MODE=native.

It publishes NO PORTS. The API has no authentication of any kind - anything that
reaches it can send messages as you and read your Signal. Gatus talks to it over
a shared docker network by service name, which is also WHY the network exists:
Gatus runs in a container, so the host's loopback is unreachable from it and a
port published on 127.0.0.1 would not have worked.

MODE=native and not json-rpc because this VPS has 464MB of RAM and already runs
Gatus and Caddy. The json-rpc modes hold a resident JVM; native runs a binary
per request, and alerts are rare enough that startup cost per alert is the right
trade.

Monitored - Gatus polls /v1/health over the same network path the alerts take,
so it proves the delivery route rather than mere container liveness. Deliberately
NOT backed up: the data directory holds Signal private keys and the recovery
path is to link the device again from the phone.

Neither the signal-api endpoint nor Gatus's self-check carries a Signal alert.
If either is down, Signal is precisely what cannot deliver the alert.

── Four traps hit while linking, all now in the role README ────────────────
  * /v1/qrcodelink is BROKEN in native mode - returns "no data to encode" while
    the binary itself emits a perfectly good URI. Not worked around by switching
    MODE, which would put a JVM in the path of every alert permanently.
  * The data dir must be owned by uid 1000, not root. A root-owned 0700 dir
    cannot be traversed by the container user, so linking silently never
    completes and /v1/accounts returns "Failed to read local accounts list".
  * `docker exec` runs as ROOT while the service runs as uid 1000, so without
    --config the account is written to /root/... on the container's ephemeral
    layer. It reports success and is destroyed on the next recreate.
  * The phone reporting "network error" was IPv6: chat.signal.org resolves to
    dualstack AAAA records first, the container has no IPv6 address, and this
    host's IPv6 path is broken - the same edge that 404'd the Go tarball. Fixed
    with a mounted gai.conf that prefers IPv4.

Verified: provider loads (configuredProviders=[signal]), a test message was
delivered and confirmed received, 84 endpoints carry alerts, 86 UP / 0 DOWN.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 21:40:37 +02:00

107 lines
5.7 KiB
YAML

---
# Is each systemd-deployed service actually running?
#
# Every 5 minutes, with an 11-minute Gatus heartbeat - one missed run before
# it alarms, so a reboot or a slow check does not page anyone, but a host that
# stops reporting does.
#
# This closes the gap that let a real bug run unnoticed: a backup script left
# forgejo, lnbits, headscale and memos stopped, and NOTHING caught it. The dumps
# exited 0, the artefacts were correct, the deploy said failed=0, and liveness
# only proves the HOST is up - not that anything on it is serving.
#
# One endpoint PER UNIT, not per host. A host running four services needs four
# endpoints, or a single red light says "something on vipy is down" without
# saying which - and that is the question you actually have at 3am. But only ONE
# timer per host: the check iterates that host's units and pushes a result for
# each, the same way check-backups.sh reports per source. Four units on vipy
# would otherwise mean four scripts, four services and four timers.
#
# Which units each host runs is in host_vars/<host>/main.yml as
# monitored_services, because "what runs here" is a property of the machine.
#
# Keys are host-qualified because unit names collide - caddy runs on four
# machines. Gatus computes sanitize(group)_sanitize(name), so group "services"
# and name "vipy/caddy" give services_vipy-caddy.
# ─────────────────────────────────────────────────────────────────────────────
# Register one endpoint per unit. Runs first: Gatus reloads within 30s, and the
# host play above takes minutes, so every endpoint exists before its first push.
# ─────────────────────────────────────────────────────────────────────────────
# ─────────────────────────────────────────────────────────────────────────────
# Alerting thresholds, and why they differ by check type.
#
# `failure-threshold` counts CONSECUTIVE failures, but "consecutive" means a
# different amount of wall-clock time per check:
#
# push/heartbeat endpoints a failure is produced once per heartbeat window
# pulled endpoints a failure is produced once per interval
#
# So the default of 3 would mean 33 minutes on an 11m heartbeat and over a day
# on a 7h one - and the heartbeat window ALREADY encodes the tolerance. An 11m
# window on a 5-minute push is precisely "one missed push forgiven"; stacking a
# threshold of 3 on top triples a tolerance that was already chosen.
#
# Hence: push endpoints alert on the FIRST heartbeat failure. Pulled endpoints
# have no built-in tolerance, so the threshold is where it belongs for them.
# ─────────────────────────────────────────────────────────────────────────────
- name: Register the service checks with Gatus
hosts: observability
become: yes
tasks:
# Two plain steps rather than one clever expression: first collect which
# units each host declares, then flatten that into endpoints.
- name: Collect the units each host declares
ansible.builtin.set_fact:
host_units: "{{ host_units | default([]) + [{'host': item, 'units': hostvars[item].monitored_services}] }}"
loop: "{{ groups['managed'] | sort }}"
when: hostvars[item].monitored_services | default([]) | length > 0
- name: Build one endpoint per unit
ansible.builtin.set_fact:
service_endpoints: "{{ service_endpoints | default([]) + [{
'name': (item.0.host | lower | regex_replace('[/_.,# +&]', '-')) ~ '/' ~ item.1,
'group': 'services',
'token': gatus_push_tokens[item.0.host],
'heartbeat': '11m'}] }}"
loop: "{{ host_units | subelements('units') }}"
- name: Register the service endpoints
ansible.builtin.include_role:
name: gatus_endpoint
vars:
gatus_endpoint_default_alerts:
- type: signal
# 1, not 3: the heartbeat window is the tolerance. See the note above.
failure-threshold: 1
success-threshold: 2
send-on-resolved: true
minimum-reminder-interval: 6h
gatus_endpoint_name: services
gatus_endpoint_external: "{{ service_endpoints }}"
- name: Monitor systemd services on every host that has them
hosts: managed
become: yes
vars:
gatus_api: "https://{{ subdomains.gatus }}.{{ root_domain }}/api/v1/endpoints"
host_key: "{{ inventory_hostname | lower | regex_replace('[/_.,# +&]', '-') }}"
tasks:
- name: Is every deployed service running?
ansible.builtin.include_role:
name: healthcheck
vars:
healthcheck_name: service-health
healthcheck_description: "systemd services on {{ inventory_hostname }}"
healthcheck_check: systemd-units
healthcheck_units: "{{ monitored_services }}"
healthcheck_units_key_prefix: "services_{{ host_key }}"
healthcheck_interval: "5min"
healthcheck_boot_delay: "2min"
# The per-unit results go to keys under this collection; the role's own
# single-result push is unused here, so only the base is set.
healthcheck_push_base: "{{ gatus_api }}"
healthcheck_push_token: "{{ gatus_push_tokens[inventory_hostname] }}"
when: monitored_services | default([]) | length > 0