Twelve endpoints in two groups, because they answer different questions and
fail for different reasons:
backup-dump_<svc> pushed by the SOURCE right after its dump runs
backup-store_<svc> pushed by the BOX at 05:30, per source
backup-store_pull-job pushed by the BOX, about the box itself
The store alone could catch almost everything, because the artefact filename
carries the source's dump timestamp - a source whose timer died still pulls
"ok" forever, but the timestamp gives it away. What the source side adds is
LATENCY and DIAGNOSIS: the store only learns at the next 04:00 pull, and it
cannot tell you whether the dump broke or the pull did.
pull-job is separate from the per-source checks because it is a LEADING
indicator where those are lagging ones. A disabled pull timer, a failed pull
job, or a filling disk are all visible immediately, while the per-source checks
only fire once an artefact is >26h stale - about a day later. Disable the timer
at 10:00 and every source stays green until tomorrow; pull-job goes red this
morning and names the cause instead of showing six stale sources.
Frequencies: dumps 02:00-02:30 staggered, pull 04:00, verification 05:30, all
daily and Persistent. 26h staleness decides red; a 30h Gatus heartbeat catches
the verification itself having stopped, so a dead check-backups.timer cannot
hide a stale backup.
arbret has no dump endpoint: prd-arbret is in [arbret], which `managed`
deliberately excludes, so nothing of ours runs there. Store-checked only.
check-backups.sh was manual-only; it now runs on a timer and reports per source
rather than only printing. The human-readable report is unchanged.
A SERIOUS bug introduced and fixed in this change, recorded because the shape
is easy to repeat: the reporting hook was added to backup.sh as a second
`trap ... EXIT`. Bash REPLACES the EXIT handler rather than adding to it, so
that silently deleted the trap which restarts the stopped service - the one the
script's own comment calls "the point", and the bug the role was written to
eliminate. Every backup then stopped its service and left it stopped. It took
forgejo, lnbits, headscale and memos down for several minutes each, and nothing
caught it: the dumps exit 0, the artefacts are correct, the deploy reports
failed=0, and liveness only proves the HOST is up. There is now ONE EXIT
handler doing both jobs, armed BEFORE the stop so a failure during the stop
still restarts. Verified by rendering both variants, asserting exactly one EXIT
trap in each, and simulating a mid-way failure to confirm the restart fires.
Verified: all 12 endpoints UP; six sources pulled cleanly (arbret 31M,
headscale 198K, memos 7.8M, vaultwarden 2.1M, lnbits 30M, forgejo 2.6G), store
16% full, "RESULT: all checks passed".
Known gaps, deliberately not closed here:
* `yell` warnings - disk 75-90%, an artefact under half the previous size,
retention not pruning - never reach Gatus, because a push is binary.
* Nothing verifies a backed-up service came back UP. That is the gap that let
the trap bug run unnoticed, and it is what the next change addresses.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|---|---|---|
| .. | ||
| defaults | ||
| handlers | ||
| tasks | ||
| templates | ||
| README.md | ||
backup_store
Pulls already-encrypted backup artefacts from every source host onto
small-backups-box, on a timer, and expires them per source.
Generalises the hand-written pull-backups.sh that had one hardcoded source
(arbret). That job's behaviour is preserved exactly: same source path, same
90 days, same destination directory.
This host holds no key
Everything pulled here is ciphertext produced by backup_source on the source
host. The box cannot read any of it — the age identity lives only on lapy. That
is deliberate: the machine holding every backup should not also be able to open
them.
One failing source must not stop the others
The script is set -uo pipefail, not -e. Each source runs in its own
function, failures are counted, and the script exits non-zero at the end so
systemd marks the unit failed. A dead host costs you that one source, not the
whole run.
This is the specific failure the whole plan exists to prevent: the laptop jobs aborted on first error and then silently produced empty directories for nine months.
Trust points one way
The box authenticates with ~/.ssh/id_pull to an unprivileged, dedicated
account on each source (backup-pull, or arbret on prd-arbret), authorised
with restrict. That account can read one directory and do nothing else — no
sudo, no pty, no forwarding. A compromised backup box cannot reach into
production.
Addressing: names, never IPs
Sources are addressed by name. The job this replaced hardcoded spacey's IP; the droplet was later rebuilt, the address was recycled to a stranger, and the backup failed silently from 2025-12-01 while the directory listing still looked healthy.
Two kinds of name are in play:
- Tailnet members (vipy, memos-box, …) → MagicDNS names. These require a
headscale ACL grant from
tag:small-backups-boxto the source's:22; without it the box cannot even resolve the peer, let alone reach it. - spacey is not a tailnet member — it is the headscale control server — so
its backup is pulled over the public internet via
headscale.contrapeso.xyz, which follows the host if the droplet is rebuilt.
Retention here is the long tail
Sources keep a few days locally; this box keeps 90 (or whatever the source entry says). Losing the source's local copy is expected.