No description
Find a file
counterweight 0f03c503c8
nodito: de-Uptime-Kuma the ZFS and NUT playbooks, extract templates
These two are host-specific by design - nodito is a pet, not cattle - so they
stay playbooks rather than becoming roles. But they had rotted.

De-Kuma, following the pattern of the six service roles:

  Both plays opened with an assert on uptime_kuma_username/password, which were
  removed from the vault, so both failed before doing anything. Dropped that,
  the two embedded Python monitor-creation scripts, and their /tmp cleanup.
  Kept every check, threshold and systemd timer - those are the durable part.

  Reporting is now generic: `healthcheck_push_url` goes into the unit as
  Environment=HEALTHCHECK_PUSH_URL and the script reads ${HEALTHCHECK_PUSH_URL:-},
  treating empty as normal rather than an error. The exit code is the real
  answer; systemd keeps it. Both scripts now also report status=down on failure
  instead of only going silent. Live push URLs harvested into the vault so
  nothing observable changes for ZFS.

Three live bugs found while check-diffing:

  1. 32_zfs would have DE-REGISTERED the Proxmox storage. `pvesm remove` was
     gated on the storage existing and `pvesm add` on it NOT existing - mutually
     exclusive - so a real run removed the proxmox-tank-1 entry backing every VM
     and never put it back. It would also have dropped `mountpoint /var/lib/vz`,
     which the live entry has and `pvesm add` does not set. Registration is now
     add-only.

  2. zfs_disk_1 named ata-...WX11TN0Z, a disk no longer in the machine. The live
     mirror is WX120LHQ + WX11TN2P; a leg was replaced and the repo never caught
     up. Inert behind `when: zfs_pool_exists.rc != 0`, but wrong on any
     disaster-recovery run. This is the seventh instance of an identifier
     written down once whose hardware later moved.

  3. 34_nut has NEVER been applied to nodito - no /etc/nut file carries the
     "Managed by Ansible" marker; they were written by hand in January 2026. The
     vault held the literal CHANGE_ME_TO_SECURE_PASSWORD, so applying it would
     have overwritten a working upsd/upsmon auth pair with a placeholder and
     restarted NUT, leaving the hypervisor's UPS unable to trigger a clean
     shutdown on mains loss. The Kuma assert was the only thing stopping that,
     so removing it without a replacement would have armed the gun: there is now
     an explicit assert that refuses to run on the placeholder. The real
     password is in the vault and `Configure upsd users` check-diffs clean.

  Templates reconciled with the live files first, so applying 34_nut is close to
  a no-op: added `maxretry = 3` to ups.conf and OFFDURATION / RBWARNTIME /
  NOCOMMWARNTIME / FINALDELAY plus quoted POWERDOWNFLAG to upsmon.conf. The only
  substantive additions left are the NOTIFYMSG/NOTIFYFLAG syslog lines.

  /usr/local/bin/ups-heartbeat.sh on the box is an orphan - mode 0644, not
  executable, referenced by no unit and no cron entry - but its push token
  belongs to a monitor that still exists and answers, so that monitor has had no
  heartbeat since January. Harvested as healthcheck_push_urls.ups; applying this
  play is what will finally feed it.

Finish the host_vars migration:

  infra/nodito/nodito_vars.yml was byte-identical to host_vars/nodito/main.yml.
  Deleted it, moved nodito_secrets.yml to host_vars/nodito/vault.yml, and
  stripped the dead vars_files entries from all three playbooks.

Extract the 13 inline `content: |` blocks to infra/nodito/templates/, pulled via
a YAML load rather than retyped. 681+582 lines become 313+344 plus templates.
Ownership parity against HEAD checked mechanically: no owner/group/mode drift on
any surviving task.

Neither playbook has been applied. ZFS play 2 check-runs failed=0 with two
changes: the script rewrite (the deployed one is a hand-edited DEBUG VERSION
with `set -x`) and the Environment line.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 19:07:50 +02:00
ansible nodito: de-Uptime-Kuma the ZFS and NUT playbooks, extract templates 2026-09-13 19:07:50 +02:00
archive/uptime_kuma archive: record Uptime Kuma monitors and setup before decommissioning 2026-09-11 22:43:15 +02:00
tofu/nodito tofu: stop gitignoring the lock file and the VM inventory 2026-09-12 18:43:14 +02:00
.gitignore tofu: stop gitignoring the lock file and the VM inventory 2026-09-12 18:43:14 +02:00
01_infra_setup.md docs: mark Uptime Kuma as decommissioned 2026-09-11 22:43:56 +02:00
02_vps_core_services_setup.md docs: mark Uptime Kuma as decommissioned 2026-09-11 22:43:56 +02:00
03_vm_disk_enlargement.md little thingies 2026-03-22 21:25:08 +01:00
README.md docs: mark Uptime Kuma as decommissioned 2026-09-11 22:43:56 +02:00
requirements.txt uptime-kuma: annotate config and drop the unused collection 2026-09-11 22:43:56 +02:00

Personal infra

My repo documenting my personal infra, along with artifacts, scripts, etc.

How to use

Go through the different numbered markdowns in the repo root to do the different parts.

How to edit secrets

ansible-vault edit ansible/your_file_with_secrets.yml

Assumes that you've set ansible/.vault_pass with chmod 600.

Overview

Services

  • Reverse Proxy
    • Deployed on Vipy
    • Caddy
    • Plan install
    • File based config
    • Crossbackup to Desky via rsync
  • Uptime Kuma — decommissioned 2026-09-11, see archive/uptime_kuma/
    • Deployed on Vipy
    • Crossbackup to Desky via rsync
  • Vaultwarden
    • Deployed on Desky
    • Crossbackup to Vipy via rsync
  • Gitea
    • Deployed on Desky
    • Crossbackup to Vipy via rsync
  • Immich
    • Deployed on Desky
  • VPN
    • All set up on Vipy
  • Bitcoin Knots
    • Deployed on Desky
  • electrs
  • Synapse Server
  • Phoenix D + LNBits
  • Backups

Infra

  • Laptop (Lapy)
  • One beefy desktop (Desky)
  • One VPS (Vipy)