personal_infra/ansible/roles/backup_store
counterweight 0112bb2c69
disk check: it never worked. Fix it, and the false DOWNs it hid behind
Investigating a dashboard full of DOWN services that were plainly healthy turned
up three bugs, all introduced by me.

── 1. The disk check has been lying since it was written ───────────────────
    df: options -P and --output are mutually exclusive

`df -P ... --output=pcent,target` errors out. `2>/dev/null` swallowed it, the
read loop got nothing, `worst` stayed 0, and the check reported "max 0% on /"
and exited 0 on EVERY host regardless of real usage. vipy is at 58%. It would
never have caught a full disk - a green light wired to nothing, which is worse
than no check at all.

Dropped -P. More importantly, "df returned no filesystems" is now a FAILURE
rather than being read as 0% - the original bug was only invisible because an
empty result was indistinguishable from an empty disk.

Testing the fix immediately surfaced a second problem it had been masking:
/sys/firmware/efi/efivars sits at 87% on a healthy machine, so a working check
would have paged daily. efivarfs and ramfs join tmpfs/devtmpfs/squashfs/overlay
in the exclusions, all of them pseudo-filesystems whose "usage" is not a fact an
operator can act on. It now reports "max 73% on / (3 filesystems)".

── 2. The false DOWNs were a half-finished deploy ──────────────────────────
The previous commit tightened heartbeats from 30h to 7h AND moved the checks
from daily to 6-hourly - but only the registration side was deployed, with
--limit observability. The hosts kept pushing once a day against a 7h window, so
after seven hours every disk, ZFS and backup-store endpoint went red.

Nothing was wrong with the services and nothing was wrong with Gatus: it
correctly reported that pushes were not arriving. Changing a heartbeat window
without deploying the matching cadence is a guaranteed false alarm, and the two
have to ship together.

── 3. The new OnCalendar was invalid ───────────────────────────────────────
"*-*-* 05:30:00,11:30:00,17:30:00,23:30:00" is not valid systemd syntax - a
comma-separated list of FULL TIMES is rejected with "bad unit file setting", and
the timer silently failed to install. Corrected to "*-*-* 05/6:30:00" and
validated with `systemd-analyze calendar` before deploying, which is how this
should have been written in the first place.

Verified: 11 disk checks and the ZFS check all HEALTHY on a manual run; timers
confirmed on the hosts as 00/6:00:00, 00/6:20:00 and 05/6:30:00; 86 UP / 0 DOWN.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-15 20:42:59 +02:00
..
defaults disk check: it never worked. Fix it, and the false DOWNs it hid behind 2026-09-15 20:42:59 +02:00
handlers backup stuff 2026-09-12 16:02:00 +02:00
tasks backups: monitor both the dump and the pull, per source 2026-09-14 09:22:12 +02:00
templates backups: monitor both the dump and the pull, per source 2026-09-14 09:22:12 +02:00
README.md backup stuff 2026-09-12 16:02:00 +02:00

backup_store

Pulls already-encrypted backup artefacts from every source host onto small-backups-box, on a timer, and expires them per source.

Generalises the hand-written pull-backups.sh that had one hardcoded source (arbret). That job's behaviour is preserved exactly: same source path, same 90 days, same destination directory.

This host holds no key

Everything pulled here is ciphertext produced by backup_source on the source host. The box cannot read any of it — the age identity lives only on lapy. That is deliberate: the machine holding every backup should not also be able to open them.

One failing source must not stop the others

The script is set -uo pipefail, not -e. Each source runs in its own function, failures are counted, and the script exits non-zero at the end so systemd marks the unit failed. A dead host costs you that one source, not the whole run.

This is the specific failure the whole plan exists to prevent: the laptop jobs aborted on first error and then silently produced empty directories for nine months.

Trust points one way

The box authenticates with ~/.ssh/id_pull to an unprivileged, dedicated account on each source (backup-pull, or arbret on prd-arbret), authorised with restrict. That account can read one directory and do nothing else — no sudo, no pty, no forwarding. A compromised backup box cannot reach into production.

Addressing: names, never IPs

Sources are addressed by name. The job this replaced hardcoded spacey's IP; the droplet was later rebuilt, the address was recycled to a stranger, and the backup failed silently from 2025-12-01 while the directory listing still looked healthy.

Two kinds of name are in play:

  • Tailnet members (vipy, memos-box, …) → MagicDNS names. These require a headscale ACL grant from tag:small-backups-box to the source's :22; without it the box cannot even resolve the peer, let alone reach it.
  • spacey is not a tailnet member — it is the headscale control server — so its backup is pulled over the public internet via headscale.contrapeso.xyz, which follows the host if the droplet is rebuilt.

Retention here is the long tail

Sources keep a few days locally; this box keeps 90 (or whatever the source entry says). Losing the source's local copy is expected.