5.6 KiB
backup_source
Makes a host back itself up: dump to stdout, encrypt with age, write to a
local directory, prune, on a systemd timer. small-backups-box pulls the
directory later (see backup_store).
Modelled on prd-arbret, which has been doing exactly this correctly since
before the rest of the estate was migrated.
Usage
- ansible.builtin.include_role:
name: backup_source
vars:
backup_source_name: headscale
backup_source_description: "Headscale"
backup_source_dump_command: "tar -czf - -C / var/lib/headscale etc/headscale"
backup_source_stop_service: headscale
backup_source_retention_days: 7
Produces /opt/backups/headscale/headscale_<YYYYmmdd_HHMMSS>.tar.gz.age,
headscale-backup.{service,timer}, and /usr/local/bin/headscale-backup.sh.
Why the source encrypts, not the destination
age -r <recipient> is asymmetric and the host holds only the public key, so
a compromised host cannot read its own backups — or anyone else's. The scripts
this replaces encrypted with GPG on the laptop, after the data had already
crossed the network, which protects the artefact at rest but not in transit.
The matching identity lives only on lapy and is escrowed. Lose it and every artefact everywhere becomes noise, including arbret's.
backup_source_dump_command writes to STDOUT
The role pipes it into age, so plaintext never touches the disk. Use -C /
with relative paths in tar rather than absolute ones: it avoids tar's "removing
leading /" and makes the restore target explicit.
Services that are not systemd
backup_source_stop_service runs systemctl stop/start. For anything else,
give the pair explicitly — vaultwarden is a docker compose stack, so
systemctl stop vaultwarden silently does nothing:
backup_source_stop_command: "docker compose -f /opt/vaultwarden/docker-compose.yml stop"
backup_source_start_command: "docker compose -f /opt/vaultwarden/docker-compose.yml start"
The same EXIT trap wraps both forms. The assert refuses a stop command without a matching start command, because that combination fails in the one way you would not notice: the service stops and never comes back.
More than one thing to back up
tar takes several paths, so multiple files or directories are normally one
artefact — headscale captures /var/lib/headscale and /etc/headscale together,
lnbits captures its data directory and its .env.
Prefer one artefact. A backup should be a consistent snapshot, and two artefacts
written by two runs can drift — you can end up restoring an .env that does not
match the database it configures. Pulling a single file back out needs no
unpacking:
age -d -i <identity> <artefact> | tar -xzO opt/lnbits/lnbits/.env
If you genuinely need separate artefacts, call the role twice with different
backup_source_names rather than extending it — but only one call may set
backup_source_stop_service, or the service is stopped twice per night.
The case this shape cannot express is a database dump plus a file tree
(pg_dump and a media directory, say): you cannot merge those into one stream
without staging plaintext on disk, which is exactly what this design avoids.
None of the current services need it — all are file trees, all stopped for the
dump. A future one that does should use two role calls.
Everything here is sqlite, so everything stops
All five services are sqlite-backed, several in WAL mode (-wal/-shm files
present). A live copy of a WAL-mode database can be torn or stale, so each is
stopped for the duration. Measured downtime: under a second for headscale and
memos, ~6 s vaultwarden, ~11 s lnbits, and 2m36s for forgejo — 2.7 G of repos
and database. That last one is the real cost of a consistent snapshot; if it
becomes unacceptable the answer is sqlite3 .backup plus an online repo copy,
not skipping the stop.
The trap is the reason this role exists
When backup_source_stop_service is set, the script stops the unit and installs
an EXIT trap that starts it again. Without it, a failed dump leaves the service
down until the next timer fires — every hand-written script this replaced had
that bug, and it was only ever masked because their systemctl stop failed
first, before anything was stopped.
Verified on spacey: with the dump forced to fail, the log shows
Stopping → Writing → Restarting, the script exits 1 (so systemd marks the unit
failed rather than hiding it), and headscale is active afterwards.
If systemctl stop itself fails, set -e exits before the trap is installed —
which is correct, because nothing was stopped.
.partial
The dump writes <artifact>.partial and only mvs it into place on success, so
a truncated file is never mistaken for a backup. A failure inside the pipeline
does leave one behind, and the prune glob cannot match it (it ends .partial,
not .tar.gz.age), so the script clears stale partials at the start of each
run. Tested by failing mid-pipeline: 1 partial left, 0 after the next run.
backup_source_stop_service may be a bare name
headscale and headscale.service both work. The unit template normalises it,
because systemd rejects a bare name in After= with
Failed to add dependency ... Invalid argument — which it logs and then ignores,
so the unit appears to work while carrying no ordering at all.
Retention is two-tier
backup_source_retention_days is local and short — these hosts are
disk-constrained. The long tail lives on small-backups-box, which decides its
own retention per source. Losing the local copy is expected and fine.