personal_infra/ansible/roles/backup_store
counterweight c2de6dbbd9
backups: monitor both the dump and the pull, per source
Twelve endpoints in two groups, because they answer different questions and
fail for different reasons:

  backup-dump_<svc>       pushed by the SOURCE right after its dump runs
  backup-store_<svc>      pushed by the BOX at 05:30, per source
  backup-store_pull-job   pushed by the BOX, about the box itself

The store alone could catch almost everything, because the artefact filename
carries the source's dump timestamp - a source whose timer died still pulls
"ok" forever, but the timestamp gives it away. What the source side adds is
LATENCY and DIAGNOSIS: the store only learns at the next 04:00 pull, and it
cannot tell you whether the dump broke or the pull did.

pull-job is separate from the per-source checks because it is a LEADING
indicator where those are lagging ones. A disabled pull timer, a failed pull
job, or a filling disk are all visible immediately, while the per-source checks
only fire once an artefact is >26h stale - about a day later. Disable the timer
at 10:00 and every source stays green until tomorrow; pull-job goes red this
morning and names the cause instead of showing six stale sources.

Frequencies: dumps 02:00-02:30 staggered, pull 04:00, verification 05:30, all
daily and Persistent. 26h staleness decides red; a 30h Gatus heartbeat catches
the verification itself having stopped, so a dead check-backups.timer cannot
hide a stale backup.

arbret has no dump endpoint: prd-arbret is in [arbret], which `managed`
deliberately excludes, so nothing of ours runs there. Store-checked only.

check-backups.sh was manual-only; it now runs on a timer and reports per source
rather than only printing. The human-readable report is unchanged.

A SERIOUS bug introduced and fixed in this change, recorded because the shape
is easy to repeat: the reporting hook was added to backup.sh as a second
`trap ... EXIT`. Bash REPLACES the EXIT handler rather than adding to it, so
that silently deleted the trap which restarts the stopped service - the one the
script's own comment calls "the point", and the bug the role was written to
eliminate. Every backup then stopped its service and left it stopped. It took
forgejo, lnbits, headscale and memos down for several minutes each, and nothing
caught it: the dumps exit 0, the artefacts are correct, the deploy reports
failed=0, and liveness only proves the HOST is up. There is now ONE EXIT
handler doing both jobs, armed BEFORE the stop so a failure during the stop
still restarts. Verified by rendering both variants, asserting exactly one EXIT
trap in each, and simulating a mid-way failure to confirm the restart fires.

Verified: all 12 endpoints UP; six sources pulled cleanly (arbret 31M,
headscale 198K, memos 7.8M, vaultwarden 2.1M, lnbits 30M, forgejo 2.6G), store
16% full, "RESULT: all checks passed".

Known gaps, deliberately not closed here:
  * `yell` warnings - disk 75-90%, an artefact under half the previous size,
    retention not pruning - never reach Gatus, because a push is binary.
  * Nothing verifies a backed-up service came back UP. That is the gap that let
    the trap bug run unnoticed, and it is what the next change addresses.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 09:22:12 +02:00
..
defaults backups: monitor both the dump and the pull, per source 2026-09-14 09:22:12 +02:00
handlers backup stuff 2026-09-12 16:02:00 +02:00
tasks backups: monitor both the dump and the pull, per source 2026-09-14 09:22:12 +02:00
templates backups: monitor both the dump and the pull, per source 2026-09-14 09:22:12 +02:00
README.md backup stuff 2026-09-12 16:02:00 +02:00

backup_store

Pulls already-encrypted backup artefacts from every source host onto small-backups-box, on a timer, and expires them per source.

Generalises the hand-written pull-backups.sh that had one hardcoded source (arbret). That job's behaviour is preserved exactly: same source path, same 90 days, same destination directory.

This host holds no key

Everything pulled here is ciphertext produced by backup_source on the source host. The box cannot read any of it — the age identity lives only on lapy. That is deliberate: the machine holding every backup should not also be able to open them.

One failing source must not stop the others

The script is set -uo pipefail, not -e. Each source runs in its own function, failures are counted, and the script exits non-zero at the end so systemd marks the unit failed. A dead host costs you that one source, not the whole run.

This is the specific failure the whole plan exists to prevent: the laptop jobs aborted on first error and then silently produced empty directories for nine months.

Trust points one way

The box authenticates with ~/.ssh/id_pull to an unprivileged, dedicated account on each source (backup-pull, or arbret on prd-arbret), authorised with restrict. That account can read one directory and do nothing else — no sudo, no pty, no forwarding. A compromised backup box cannot reach into production.

Addressing: names, never IPs

Sources are addressed by name. The job this replaced hardcoded spacey's IP; the droplet was later rebuilt, the address was recycled to a stranger, and the backup failed silently from 2025-12-01 while the directory listing still looked healthy.

Two kinds of name are in play:

  • Tailnet members (vipy, memos-box, …) → MagicDNS names. These require a headscale ACL grant from tag:small-backups-box to the source's :22; without it the box cannot even resolve the peer, let alone reach it.
  • spacey is not a tailnet member — it is the headscale control server — so its backup is pulled over the public internet via headscale.contrapeso.xyz, which follows the host if the droplet is rebuilt.

Retention here is the long tail

Sources keep a few days locally; this box keeps 90 (or whatever the source entry says). Losing the source's local copy is expected.