No description
Find a file
counterweight 853e62a19c
monitoring: retire the Uptime-Kuma-era checks, add ZFS pool capacity
Five things deprecated, each verified against the DEPLOYED script before being
deleted rather than assumed superseded:

  infra/410_disk_usage_alerts.yml   -> disk-usage check   (infra/400)
  infra/420_system_healthcheck.yml  -> liveness check     (infra/400)
  infra/430_cpu_temp_alerts.yml     -> cpu-temp check     (infra/400)
  32_zfs play 2 (monitoring half)   -> zfs-health check   (infra/400)
  34_nut play 2 (entirely)          -> ups-status check   (infra/400)

Nothing is lost by the swap. The old system_healthcheck.sh only computed uptime
and pushed, which is exactly a liveness heartbeat. The old disk monitor was
WEAKER than its replacement: it checked "/" alone at 80%, where the new one
walks every real filesystem at 85%.

Deleting the playbooks was not the hard part. The units they installed live on
the hosts, enabled, and keep firing regardless of what the repo says - two of
them were still pushing to uptime.contrapeso.xyz every 15 minutes across nine
machines. A playbook deleted without a cleanup leaves its output running
forever with nothing left to explain it. So infra/409_remove_legacy_monitoring
stops, disables and removes the units, deletes /opt/{disk-monitoring,
system-healthcheck,nodito-monitoring,zfs-monitoring}, and removes the orphaned
hand-written ups-heartbeat.sh. It ends by grepping for any surviving Kuma
reference and reporting it. Kept permanently and idempotent, so a rebuilt or
restored host cannot quietly bring them back.

A trap avoided: the monthly ZFS scrub lived INSIDE 32_zfs play 2. Deleting the
play wholesale would have silently stopped scrubbing the pool - and an
unscrubbed pool makes the health check meaningless, because it would have
nothing true to report. That play is now scrub-only and check-runs ok=5
changed=0.

ZFS pool capacity added as a sixth condition to the zfs-health check. `zpool
status` reports a 95% full pool as perfectly ONLINE, so capacity has to be read
separately with `zpool list` - and it is the failure you get warning of rather
than the one you discover. Threshold 80%, because ZFS allocation degrades badly
past roughly that and fragmentation is hard to undo. The pool is at 47%.
Verified both directions: passes on the real pool, and a simulated 91% exits 1.

Also fixed the waste recorded in c2de6db: the healthcheck role installed its
dependencies once per CHECK rather than per HOST - 29 apt transactions
estate-wide for a curl already present, and the slowest part of every deploy.
It now deduplicates within a play run, and the redundant standalone
daemon_reload is gone (the systemd task already does one).

site.yml updated, which exposed that services/gatus was never in it. It now runs
before the three registration playbooks, since registering endpoints against a
Gatus that is not yet serving would simply fail.

Verified: no legacy timer remains on any host; 83 endpoints, 83 UP, 0 DOWN.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 10:18:57 +02:00
ansible monitoring: retire the Uptime-Kuma-era checks, add ZFS pool capacity 2026-09-14 10:18:57 +02:00
archive/uptime_kuma archive: record Uptime Kuma monitors and setup before decommissioning 2026-09-11 22:43:15 +02:00
tofu/nodito tofu: stop gitignoring the lock file and the VM inventory 2026-09-12 18:43:14 +02:00
.gitignore tofu: stop gitignoring the lock file and the VM inventory 2026-09-12 18:43:14 +02:00
01_infra_setup.md docs: mark Uptime Kuma as decommissioned 2026-09-11 22:43:56 +02:00
02_vps_core_services_setup.md docs: mark Uptime Kuma as decommissioned 2026-09-11 22:43:56 +02:00
03_vm_disk_enlargement.md little thingies 2026-03-22 21:25:08 +01:00
README.md docs: mark Uptime Kuma as decommissioned 2026-09-11 22:43:56 +02:00
requirements.txt uptime-kuma: annotate config and drop the unused collection 2026-09-11 22:43:56 +02:00

Personal infra

My repo documenting my personal infra, along with artifacts, scripts, etc.

How to use

Go through the different numbered markdowns in the repo root to do the different parts.

How to edit secrets

ansible-vault edit ansible/your_file_with_secrets.yml

Assumes that you've set ansible/.vault_pass with chmod 600.

Overview

Services

  • Reverse Proxy
    • Deployed on Vipy
    • Caddy
    • Plan install
    • File based config
    • Crossbackup to Desky via rsync
  • Uptime Kuma — decommissioned 2026-09-11, see archive/uptime_kuma/
    • Deployed on Vipy
    • Crossbackup to Desky via rsync
  • Vaultwarden
    • Deployed on Desky
    • Crossbackup to Vipy via rsync
  • Gitea
    • Deployed on Desky
    • Crossbackup to Vipy via rsync
  • Immich
    • Deployed on Desky
  • VPN
    • All set up on Vipy
  • Bitcoin Knots
    • Deployed on Desky
  • electrs
  • Synapse Server
  • Phoenix D + LNBits
  • Backups

Infra

  • Laptop (Lapy)
  • One beefy desktop (Desky)
  • One VPS (Vipy)