Three changes: detection windows tightened, a Signal transport deployed, and
alerts attached to all 84 non-transport endpoints.
── Failing faster ──────────────────────────────────────────────────────────
The heartbeat is a TICKER, not a deadline: Gatus wakes every interval and asks
"did anything arrive in the last interval", so real detection is 1-2x the
window. And a window can only ever be as tight as the push frequency - which is
why two checks moved rather than just having their numbers changed.
liveness, cpu, ups, service-health, probes 16m -> 11m
disk, zfs daily/30h -> 6-hourly/7h
backup store + pull job daily/30h -> 6-hourly/7h
DNS records 24h -> 6h
backup dump 30h -> 26h
backup-dump stays slow because the dump genuinely is daily. The store-side check
catches the same fault within 6h by reading the source's dump timestamp out of
the artefact filename, so 26h is a backstop rather than the primary signal.
── Thresholds differ by check type, deliberately ───────────────────────────
`failure-threshold` counts consecutive failures, but "consecutive" is a
different amount of wall-clock time per check: a push endpoint produces one
failure per heartbeat window, a pulled one per interval. The default of 3 would
mean 33 minutes on an 11m heartbeat and over a day on a 7h one.
More importantly the heartbeat window ALREADY encodes the tolerance - an 11m
window on a 5-minute push is exactly "one missed push forgiven" - so stacking a
threshold of 3 triples a tolerance that was already chosen. Hence:
push/heartbeat endpoints failure-threshold 1
pulled, 5m (public) failure-threshold 3 (= 15 minutes)
pulled, 6h/24h (dns, domain) failure-threshold 1
── The Signal transport ────────────────────────────────────────────────────
roles/signal_api runs signal-cli-rest-api on the observability host, pinned by
digest, MODE=native.
It publishes NO PORTS. The API has no authentication of any kind - anything that
reaches it can send messages as you and read your Signal. Gatus talks to it over
a shared docker network by service name, which is also WHY the network exists:
Gatus runs in a container, so the host's loopback is unreachable from it and a
port published on 127.0.0.1 would not have worked.
MODE=native and not json-rpc because this VPS has 464MB of RAM and already runs
Gatus and Caddy. The json-rpc modes hold a resident JVM; native runs a binary
per request, and alerts are rare enough that startup cost per alert is the right
trade.
Monitored - Gatus polls /v1/health over the same network path the alerts take,
so it proves the delivery route rather than mere container liveness. Deliberately
NOT backed up: the data directory holds Signal private keys and the recovery
path is to link the device again from the phone.
Neither the signal-api endpoint nor Gatus's self-check carries a Signal alert.
If either is down, Signal is precisely what cannot deliver the alert.
── Four traps hit while linking, all now in the role README ────────────────
* /v1/qrcodelink is BROKEN in native mode - returns "no data to encode" while
the binary itself emits a perfectly good URI. Not worked around by switching
MODE, which would put a JVM in the path of every alert permanently.
* The data dir must be owned by uid 1000, not root. A root-owned 0700 dir
cannot be traversed by the container user, so linking silently never
completes and /v1/accounts returns "Failed to read local accounts list".
* `docker exec` runs as ROOT while the service runs as uid 1000, so without
--config the account is written to /root/... on the container's ephemeral
layer. It reports success and is destroyed on the next recreate.
* The phone reporting "network error" was IPv6: chat.signal.org resolves to
dualstack AAAA records first, the container has no IPv6 address, and this
host's IPv6 path is broken - the same edge that 404'd the Go tarball. Fixed
with a mounted gai.conf that prefers IPv4.
Verified: provider loads (configuredProviders=[signal]), a test message was
delivered and confirmed received, 84 endpoints carry alerts, 86 UP / 0 DOWN.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
107 lines
4.9 KiB
YAML
107 lines
4.9 KiB
YAML
- name: Configure the offsite backup pull
|
|
hosts: backup_store
|
|
gather_facts: yes
|
|
|
|
tasks:
|
|
- name: Ensure the box pulls every source on a timer
|
|
ansible.builtin.include_role:
|
|
name: backup_store
|
|
vars:
|
|
# check-backups.sh reports one result per source plus one for the store
|
|
# itself, so it needs the collection URL and appends each key.
|
|
backup_store_check_push_base: "https://{{ subdomains.gatus }}.{{ root_domain }}/api/v1/endpoints"
|
|
backup_store_check_push_token: "{{ gatus_push_tokens[inventory_hostname] }}"
|
|
backup_store_sources:
|
|
- name: arbret
|
|
source: "arbret@prd-arbret:/opt/arbret/backups/"
|
|
retention_days: 90
|
|
- name: headscale
|
|
source: "backup-pull@headscale.contrapeso.xyz:/opt/backups/headscale/"
|
|
retention_days: 90
|
|
- name: memos
|
|
source: "backup-pull@memos-box:/opt/backups/memos/"
|
|
retention_days: 90
|
|
- name: vaultwarden
|
|
source: "backup-pull@prd-vipy:/opt/backups/vaultwarden/"
|
|
retention_days: 90
|
|
- name: lnbits
|
|
source: "backup-pull@prd-vipy:/opt/backups/lnbits/"
|
|
retention_days: 90
|
|
- name: forgejo
|
|
source: "backup-pull@prd-vipy:/opt/backups/forgejo/"
|
|
retention_days: 14
|
|
|
|
# ─────────────────────────────────────────────────────────────────────────────
|
|
# Register the backup checks with Gatus.
|
|
#
|
|
# Two groups on purpose, because they answer different questions and fail for
|
|
# different reasons:
|
|
#
|
|
# backup-dump did the SOURCE produce an artefact? Pushed by each dump right
|
|
# after it runs, so a broken dump is visible within minutes.
|
|
# backup-store did it ARRIVE, is it fresh, non-zero, plausibly sized, and is
|
|
# retention pruning? Pushed by check-backups.sh at 05:30.
|
|
#
|
|
# The store alone could catch almost everything, because the artefact filename
|
|
# carries the source's dump timestamp - a source whose timer died still pulls
|
|
# "ok" forever, but the timestamp gives it away. What the source side adds is
|
|
# LATENCY and DIAGNOSIS: the store only learns at the next 04:00 pull, and it
|
|
# cannot tell you whether the dump broke or the pull did.
|
|
#
|
|
# arbret has no dump endpoint: prd-arbret lives in [arbret], which `managed`
|
|
# deliberately excludes, so nothing of ours runs there. It is store-checked only.
|
|
# ─────────────────────────────────────────────────────────────────────────────
|
|
- name: Register the backup checks with Gatus
|
|
hosts: observability
|
|
become: yes
|
|
vars:
|
|
# Sources we deploy the dump for, and the host each one runs on.
|
|
dump_sources:
|
|
- {name: headscale, host: spacey}
|
|
- {name: memos, host: memos_box_local}
|
|
- {name: vaultwarden, host: vipy}
|
|
- {name: lnbits, host: vipy}
|
|
- {name: forgejo, host: vipy}
|
|
store_sources: [arbret, headscale, memos, vaultwarden, lnbits, forgejo]
|
|
|
|
tasks:
|
|
# 26h, not 7h: the DUMP is genuinely daily, so the window cannot be tighter
|
|
# than a day plus slack. The store-side check catches the same fault within
|
|
# 6h by reading the artefact's dump timestamp out of the filename, so this is
|
|
# the slow backstop rather than the primary signal.
|
|
- name: Build the dump endpoint list
|
|
ansible.builtin.set_fact:
|
|
dump_endpoints: "{{ dump_endpoints | default([]) + [{
|
|
'name': item.name,
|
|
'group': 'backup-dump',
|
|
'token': gatus_push_tokens[item.host],
|
|
'heartbeat': '26h'}] }}"
|
|
loop: "{{ dump_sources }}"
|
|
|
|
- name: Build the store endpoint list
|
|
ansible.builtin.set_fact:
|
|
store_endpoints: "{{ store_endpoints | default([]) + [{
|
|
'name': item,
|
|
'group': 'backup-store',
|
|
'token': gatus_push_tokens['small_backups_local'],
|
|
'heartbeat': '7h'}] }}"
|
|
loop: "{{ store_sources }}"
|
|
|
|
- name: Register the backup endpoints
|
|
ansible.builtin.include_role:
|
|
name: gatus_endpoint
|
|
vars:
|
|
# Push endpoints: the heartbeat window is the tolerance, so alert on
|
|
# the first failure rather than waiting for three 7h windows to pass.
|
|
gatus_endpoint_default_alerts:
|
|
- type: signal
|
|
failure-threshold: 1
|
|
success-threshold: 2
|
|
send-on-resolved: true
|
|
minimum-reminder-interval: 12h
|
|
gatus_endpoint_name: backups
|
|
gatus_endpoint_external: "{{ dump_endpoints + store_endpoints + [{
|
|
'name': 'pull job',
|
|
'group': 'backup-store',
|
|
'token': gatus_push_tokens['small_backups_local'],
|
|
'heartbeat': '7h'}] }}"
|