personal_infra/ansible/host_vars/nodito/main.yml

46 lines
1.8 KiB
YAML
Raw Normal View History

2026-09-11 22:17:13 +02:00
# Nodito CPU Temperature Monitoring Configuration
# Temperature Monitoring Configuration
temp_threshold_celsius: 80
temp_check_interval_minutes: 1
# Script Configuration
monitoring_script_dir: /opt/nodito-monitoring
monitoring_script_path: "{{ monitoring_script_dir }}/cpu_temp_monitor.sh"
log_file: "{{ monitoring_script_dir }}/cpu_temp_monitor.log"
# System Configuration
systemd_service_name: nodito-cpu-temp-monitor
# ZFS Pool Configuration
zfs_pool_name: "proxmox-tank-1"
nodito: de-Uptime-Kuma the ZFS and NUT playbooks, extract templates These two are host-specific by design - nodito is a pet, not cattle - so they stay playbooks rather than becoming roles. But they had rotted. De-Kuma, following the pattern of the six service roles: Both plays opened with an assert on uptime_kuma_username/password, which were removed from the vault, so both failed before doing anything. Dropped that, the two embedded Python monitor-creation scripts, and their /tmp cleanup. Kept every check, threshold and systemd timer - those are the durable part. Reporting is now generic: `healthcheck_push_url` goes into the unit as Environment=HEALTHCHECK_PUSH_URL and the script reads ${HEALTHCHECK_PUSH_URL:-}, treating empty as normal rather than an error. The exit code is the real answer; systemd keeps it. Both scripts now also report status=down on failure instead of only going silent. Live push URLs harvested into the vault so nothing observable changes for ZFS. Three live bugs found while check-diffing: 1. 32_zfs would have DE-REGISTERED the Proxmox storage. `pvesm remove` was gated on the storage existing and `pvesm add` on it NOT existing - mutually exclusive - so a real run removed the proxmox-tank-1 entry backing every VM and never put it back. It would also have dropped `mountpoint /var/lib/vz`, which the live entry has and `pvesm add` does not set. Registration is now add-only. 2. zfs_disk_1 named ata-...WX11TN0Z, a disk no longer in the machine. The live mirror is WX120LHQ + WX11TN2P; a leg was replaced and the repo never caught up. Inert behind `when: zfs_pool_exists.rc != 0`, but wrong on any disaster-recovery run. This is the seventh instance of an identifier written down once whose hardware later moved. 3. 34_nut has NEVER been applied to nodito - no /etc/nut file carries the "Managed by Ansible" marker; they were written by hand in January 2026. The vault held the literal CHANGE_ME_TO_SECURE_PASSWORD, so applying it would have overwritten a working upsd/upsmon auth pair with a placeholder and restarted NUT, leaving the hypervisor's UPS unable to trigger a clean shutdown on mains loss. The Kuma assert was the only thing stopping that, so removing it without a replacement would have armed the gun: there is now an explicit assert that refuses to run on the placeholder. The real password is in the vault and `Configure upsd users` check-diffs clean. Templates reconciled with the live files first, so applying 34_nut is close to a no-op: added `maxretry = 3` to ups.conf and OFFDURATION / RBWARNTIME / NOCOMMWARNTIME / FINALDELAY plus quoted POWERDOWNFLAG to upsmon.conf. The only substantive additions left are the NOTIFYMSG/NOTIFYFLAG syslog lines. /usr/local/bin/ups-heartbeat.sh on the box is an orphan - mode 0644, not executable, referenced by no unit and no cron entry - but its push token belongs to a monitor that still exists and answers, so that monitor has had no heartbeat since January. Harvested as healthcheck_push_urls.ups; applying this play is what will finally feed it. Finish the host_vars migration: infra/nodito/nodito_vars.yml was byte-identical to host_vars/nodito/main.yml. Deleted it, moved nodito_secrets.yml to host_vars/nodito/vault.yml, and stripped the dead vars_files entries from all three playbooks. Extract the 13 inline `content: |` blocks to infra/nodito/templates/, pulled via a YAML load rather than retyped. 681+582 lines become 313+344 plus templates. Ownership parity against HEAD checked mechanically: no owner/group/mode drift on any surviving task. Neither playbook has been applied. ZFS play 2 check-runs failed=0 with two changes: the script rewrite (the deployed one is a hand-edited DEBUG VERSION with `set -x`) and the Environment line. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 19:07:50 +02:00
# Corrected 2026-09-13: this said WX11TN0Z, a disk that is no longer in the
# machine. The live mirror is WX120LHQ + WX11TN2P - a leg was evidently
# replaced and the repo never caught up. Pool creation is guarded by
# `when: zfs_pool_exists.rc != 0` so it was inert, but it would have been
# wrong on any disaster-recovery run.
zfs_disk_1: "/dev/disk/by-id/ata-ST4000NT001-3M2101_WX120LHQ" # First disk for RAID 1 mirror
2026-09-11 22:17:13 +02:00
zfs_disk_2: "/dev/disk/by-id/ata-ST4000NT001-3M2101_WX11TN2P" # Second disk for RAID 1 mirror
zfs_pool_mountpoint: "/var/lib/vz"
# UPS Configuration (CyberPower CP900EPFCLCD via USB)
ups_name: cyberpower
ups_desc: "CyberPower CP900EPFCLCD"
ups_driver: usbhid-ups
ups_port: auto
ups_user: counterweight
ups_offdelay: 120 # Seconds after shutdown before UPS cuts outlet power
ups_ondelay: 30 # Seconds after mains returns before UPS restores outlet power
monitoring: systemd services, domain expiry, DNS correctness, public endpoints Four more check types, 42 endpoints, taking the estate from 39 to 81. ── systemd services (infra/401) ──────────────────────────────────────────── Every unit we deploy, checked every 5 minutes with a 16-minute heartbeat. This closes the gap that let a real bug run unnoticed earlier today: a backup script left forgejo, lnbits, headscale and memos stopped, and NOTHING caught it. The dumps exited 0, the artefacts were correct, the deploy said failed=0, and liveness only proves the HOST is up - not that anything on it serves. One endpoint PER UNIT but only ONE timer per host: the check iterates that host's units and pushes a result for each, the way check-backups.sh reports per source. Four units on vipy would otherwise mean four scripts, services and timers. A single host-level red light would also say "something on vipy is down" without saying which, which is not the question anyone has. Keys are host-qualified because unit names collide - caddy runs on four machines. Which units a host runs lives in host_vars/<host>/monitored_services, because "what runs here" is a property of the machine, the same reasoning as the cross-host ports. nut-driver-enumerator is deliberately excluded: it is a oneshot generator that is `enabled` but always `inactive`, so it would report down forever. Checked live before excluding it. ── domain, DNS and public endpoints (infra/402) ──────────────────────────── The first checks that PULL rather than push, and that is the right way round: all three are about how the outside world sees us, so they must be measured from outside. Nothing is installed anywhere - no script, no timer, no token. They also have no heartbeat, because a heartbeat answers "did the reporter report"; when Gatus does the checking itself, failure is immediate. domain 1 endpoint, 24h, [DOMAIN_EXPIRATION] > 336h (two weeks) dns 11 endpoints, 24h, [DNS_RCODE] == NOERROR and [BODY] == the IP public 14 endpoints, 5m, 11 HTTPS + 3 TCP Expected A records are derived from inventory (hostvars[host].ansible_host), not written down again. The estate's recurring bug is an address recorded in a second place and left behind when the machine moved; asserting against inventory means a renumbered box is one edit, not two. Expected HTTP status was checked live per site rather than assumed. Two return 401 - the Gatus dashboard and the DATUM dashboard, both behind basic auth - and that is what is asserted: expecting 200 there would go green precisely when the auth broke. headscale asserts /health rather than /, which is a 404 by design. Every HTTPS check also carries [CERTIFICATE_EXPIRATION] > 168h, which is free on an endpoint already being polled and catches a renewal that silently stops. Two things learned the hard way, both now in comments: * A domain-expiry endpoint needs a URL SCHEME. Gatus derives the endpoint type from the prefix (endpoint.Type()), so a bare "contrapeso.xyz" is UNKNOWN and the whole config is rejected. It is https:// plus a DOMAIN_EXPIRATION condition and no status assertion, so the registrar's parking page at the apex is irrelevant. Upstream also enforces a 5m minimum interval for that placeholder, because it uses a free whois service. * That rejection proved skip-invalid-config-update was worth adding. Gatus logged "the configuration file was updated, but it is not valid, the old configuration will continue being used" and kept running. Without it the reload path calls panic() and one malformed contributed file takes the monitor down. Verified: 15/15 service endpoints UP; 26/27 public-facing UP. The single DOWN is public_memos, correctly - memos-box is powered off, and the condition result reads [STATUS] (502) == 200. memos-box and arbret-staging-box were shut down deliberately from the Proxmox UI to test the liveness endpoints; their endpoints stay registered and will go green when the VMs return. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 09:33:57 +02:00
# Systemd services deployed on this host, monitored every 5 minutes.
#
# The fact lives with the machine rather than in a central map, for the same
# reason the cross-host ports do: "what runs here" is a property of the host,
# and a central list is one more thing to forget to update when a service moves.
#
# Only units WE deploy belong here. Distro units (ssh, cron) have their own
# supervision and would be noise.
monitored_services:
- nut-server
- nut-monitor