personal_infra/ansible/roles/backup_store
counterweight 3a9e1d5851
alerting: Signal via signal-cli-rest-api, and faster failure detection
Three changes: detection windows tightened, a Signal transport deployed, and
alerts attached to all 84 non-transport endpoints.

── Failing faster ──────────────────────────────────────────────────────────
The heartbeat is a TICKER, not a deadline: Gatus wakes every interval and asks
"did anything arrive in the last interval", so real detection is 1-2x the
window. And a window can only ever be as tight as the push frequency - which is
why two checks moved rather than just having their numbers changed.

  liveness, cpu, ups, service-health, probes   16m -> 11m
  disk, zfs                       daily/30h -> 6-hourly/7h
  backup store + pull job         daily/30h -> 6-hourly/7h
  DNS records                           24h -> 6h
  backup dump                             30h -> 26h

backup-dump stays slow because the dump genuinely is daily. The store-side check
catches the same fault within 6h by reading the source's dump timestamp out of
the artefact filename, so 26h is a backstop rather than the primary signal.

── Thresholds differ by check type, deliberately ───────────────────────────
`failure-threshold` counts consecutive failures, but "consecutive" is a
different amount of wall-clock time per check: a push endpoint produces one
failure per heartbeat window, a pulled one per interval. The default of 3 would
mean 33 minutes on an 11m heartbeat and over a day on a 7h one.

More importantly the heartbeat window ALREADY encodes the tolerance - an 11m
window on a 5-minute push is exactly "one missed push forgiven" - so stacking a
threshold of 3 triples a tolerance that was already chosen. Hence:

  push/heartbeat endpoints   failure-threshold 1
  pulled, 5m (public)        failure-threshold 3   (= 15 minutes)
  pulled, 6h/24h (dns, domain) failure-threshold 1

── The Signal transport ────────────────────────────────────────────────────
roles/signal_api runs signal-cli-rest-api on the observability host, pinned by
digest, MODE=native.

It publishes NO PORTS. The API has no authentication of any kind - anything that
reaches it can send messages as you and read your Signal. Gatus talks to it over
a shared docker network by service name, which is also WHY the network exists:
Gatus runs in a container, so the host's loopback is unreachable from it and a
port published on 127.0.0.1 would not have worked.

MODE=native and not json-rpc because this VPS has 464MB of RAM and already runs
Gatus and Caddy. The json-rpc modes hold a resident JVM; native runs a binary
per request, and alerts are rare enough that startup cost per alert is the right
trade.

Monitored - Gatus polls /v1/health over the same network path the alerts take,
so it proves the delivery route rather than mere container liveness. Deliberately
NOT backed up: the data directory holds Signal private keys and the recovery
path is to link the device again from the phone.

Neither the signal-api endpoint nor Gatus's self-check carries a Signal alert.
If either is down, Signal is precisely what cannot deliver the alert.

── Four traps hit while linking, all now in the role README ────────────────
  * /v1/qrcodelink is BROKEN in native mode - returns "no data to encode" while
    the binary itself emits a perfectly good URI. Not worked around by switching
    MODE, which would put a JVM in the path of every alert permanently.
  * The data dir must be owned by uid 1000, not root. A root-owned 0700 dir
    cannot be traversed by the container user, so linking silently never
    completes and /v1/accounts returns "Failed to read local accounts list".
  * `docker exec` runs as ROOT while the service runs as uid 1000, so without
    --config the account is written to /root/... on the container's ephemeral
    layer. It reports success and is destroyed on the next recreate.
  * The phone reporting "network error" was IPv6: chat.signal.org resolves to
    dualstack AAAA records first, the container has no IPv6 address, and this
    host's IPv6 path is broken - the same edge that 404'd the Go tarball. Fixed
    with a mounted gai.conf that prefers IPv4.

Verified: provider loads (configuredProviders=[signal]), a test message was
delivered and confirmed received, 84 endpoints carry alerts, 86 UP / 0 DOWN.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 21:40:37 +02:00
..
defaults alerting: Signal via signal-cli-rest-api, and faster failure detection 2026-09-14 21:40:37 +02:00
handlers backup stuff 2026-09-12 16:02:00 +02:00
tasks backups: monitor both the dump and the pull, per source 2026-09-14 09:22:12 +02:00
templates backups: monitor both the dump and the pull, per source 2026-09-14 09:22:12 +02:00
README.md backup stuff 2026-09-12 16:02:00 +02:00

backup_store

Pulls already-encrypted backup artefacts from every source host onto small-backups-box, on a timer, and expires them per source.

Generalises the hand-written pull-backups.sh that had one hardcoded source (arbret). That job's behaviour is preserved exactly: same source path, same 90 days, same destination directory.

This host holds no key

Everything pulled here is ciphertext produced by backup_source on the source host. The box cannot read any of it — the age identity lives only on lapy. That is deliberate: the machine holding every backup should not also be able to open them.

One failing source must not stop the others

The script is set -uo pipefail, not -e. Each source runs in its own function, failures are counted, and the script exits non-zero at the end so systemd marks the unit failed. A dead host costs you that one source, not the whole run.

This is the specific failure the whole plan exists to prevent: the laptop jobs aborted on first error and then silently produced empty directories for nine months.

Trust points one way

The box authenticates with ~/.ssh/id_pull to an unprivileged, dedicated account on each source (backup-pull, or arbret on prd-arbret), authorised with restrict. That account can read one directory and do nothing else — no sudo, no pty, no forwarding. A compromised backup box cannot reach into production.

Addressing: names, never IPs

Sources are addressed by name. The job this replaced hardcoded spacey's IP; the droplet was later rebuilt, the address was recycled to a stranger, and the backup failed silently from 2025-12-01 while the directory listing still looked healthy.

Two kinds of name are in play:

  • Tailnet members (vipy, memos-box, …) → MagicDNS names. These require a headscale ACL grant from tag:small-backups-box to the source's :22; without it the box cannot even resolve the peer, let alone reach it.
  • spacey is not a tailnet member — it is the headscale control server — so its backup is pulled over the public internet via headscale.contrapeso.xyz, which follows the host if the droplet is rebuilt.

Retention here is the long tail

Sources keep a few days locally; this box keeps 90 (or whatever the source entry says). Losing the source's local copy is expected.