2026-01-11 22:43:27 +01:00
|
|
|
- name: Setup NUT (Network UPS Tools) for CyberPower UPS
|
2026-09-11 21:56:24 +02:00
|
|
|
hosts: hypervisor
|
2026-01-11 22:43:27 +01:00
|
|
|
become: true
|
|
|
|
|
|
|
|
|
|
tasks:
|
nodito: de-Uptime-Kuma the ZFS and NUT playbooks, extract templates
These two are host-specific by design - nodito is a pet, not cattle - so they
stay playbooks rather than becoming roles. But they had rotted.
De-Kuma, following the pattern of the six service roles:
Both plays opened with an assert on uptime_kuma_username/password, which were
removed from the vault, so both failed before doing anything. Dropped that,
the two embedded Python monitor-creation scripts, and their /tmp cleanup.
Kept every check, threshold and systemd timer - those are the durable part.
Reporting is now generic: `healthcheck_push_url` goes into the unit as
Environment=HEALTHCHECK_PUSH_URL and the script reads ${HEALTHCHECK_PUSH_URL:-},
treating empty as normal rather than an error. The exit code is the real
answer; systemd keeps it. Both scripts now also report status=down on failure
instead of only going silent. Live push URLs harvested into the vault so
nothing observable changes for ZFS.
Three live bugs found while check-diffing:
1. 32_zfs would have DE-REGISTERED the Proxmox storage. `pvesm remove` was
gated on the storage existing and `pvesm add` on it NOT existing - mutually
exclusive - so a real run removed the proxmox-tank-1 entry backing every VM
and never put it back. It would also have dropped `mountpoint /var/lib/vz`,
which the live entry has and `pvesm add` does not set. Registration is now
add-only.
2. zfs_disk_1 named ata-...WX11TN0Z, a disk no longer in the machine. The live
mirror is WX120LHQ + WX11TN2P; a leg was replaced and the repo never caught
up. Inert behind `when: zfs_pool_exists.rc != 0`, but wrong on any
disaster-recovery run. This is the seventh instance of an identifier
written down once whose hardware later moved.
3. 34_nut has NEVER been applied to nodito - no /etc/nut file carries the
"Managed by Ansible" marker; they were written by hand in January 2026. The
vault held the literal CHANGE_ME_TO_SECURE_PASSWORD, so applying it would
have overwritten a working upsd/upsmon auth pair with a placeholder and
restarted NUT, leaving the hypervisor's UPS unable to trigger a clean
shutdown on mains loss. The Kuma assert was the only thing stopping that,
so removing it without a replacement would have armed the gun: there is now
an explicit assert that refuses to run on the placeholder. The real
password is in the vault and `Configure upsd users` check-diffs clean.
Templates reconciled with the live files first, so applying 34_nut is close to
a no-op: added `maxretry = 3` to ups.conf and OFFDURATION / RBWARNTIME /
NOCOMMWARNTIME / FINALDELAY plus quoted POWERDOWNFLAG to upsmon.conf. The only
substantive additions left are the NOTIFYMSG/NOTIFYFLAG syslog lines.
/usr/local/bin/ups-heartbeat.sh on the box is an orphan - mode 0644, not
executable, referenced by no unit and no cron entry - but its push token
belongs to a monitor that still exists and answers, so that monitor has had no
heartbeat since January. Harvested as healthcheck_push_urls.ups; applying this
play is what will finally feed it.
Finish the host_vars migration:
infra/nodito/nodito_vars.yml was byte-identical to host_vars/nodito/main.yml.
Deleted it, moved nodito_secrets.yml to host_vars/nodito/vault.yml, and
stripped the dead vars_files entries from all three playbooks.
Extract the 13 inline `content: |` blocks to infra/nodito/templates/, pulled via
a YAML load rather than retyped. 681+582 lines become 313+344 plus templates.
Ownership parity against HEAD checked mechanically: no owner/group/mode drift on
any surviving task.
Neither playbook has been applied. ZFS play 2 check-runs failed=0 with two
changes: the script rewrite (the deployed one is a hand-edited DEBUG VERSION
with `set -x`) and the Environment line.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 19:07:50 +02:00
|
|
|
# ------------------------------------------------------------------
|
|
|
|
|
# Safety catch
|
|
|
|
|
#
|
|
|
|
|
# /etc/nut/upsd.users and /etc/nut/upsmon.conf on nodito were written by
|
|
|
|
|
# hand in January 2026 and carry a working password. host_vars/nodito/vault.yml
|
|
|
|
|
# (formerly infra/nodito/nodito_secrets.yml) still holds the literal string
|
|
|
|
|
# CHANGE_ME_TO_SECURE_PASSWORD, so running this play would overwrite that
|
|
|
|
|
# working pair with a placeholder and restart NUT - leaving the hypervisor's
|
|
|
|
|
# UPS unmonitored and unable to trigger a clean shutdown on mains loss.
|
|
|
|
|
#
|
|
|
|
|
# Until the real password is put in the vault, stop here.
|
|
|
|
|
# ansible-vault edit host_vars/nodito/vault.yml
|
|
|
|
|
# ------------------------------------------------------------------
|
|
|
|
|
- name: Refuse to run with a placeholder UPS password
|
|
|
|
|
assert:
|
|
|
|
|
that:
|
|
|
|
|
- ups_password is defined
|
|
|
|
|
- ups_password | length > 0
|
|
|
|
|
- ups_password != "CHANGE_ME_TO_SECURE_PASSWORD"
|
|
|
|
|
fail_msg: >-
|
|
|
|
|
ups_password is unset or still the placeholder. Applying this play would
|
|
|
|
|
overwrite the working /etc/nut/upsd.users and /etc/nut/upsmon.conf on
|
|
|
|
|
nodito and restart NUT. Put the real password in the vault first:
|
|
|
|
|
ansible-vault edit host_vars/nodito/vault.yml
|
|
|
|
|
|
2026-01-11 22:43:27 +01:00
|
|
|
# ------------------------------------------------------------------
|
|
|
|
|
# Installation
|
|
|
|
|
# ------------------------------------------------------------------
|
|
|
|
|
- name: Install NUT packages
|
|
|
|
|
apt:
|
|
|
|
|
name:
|
|
|
|
|
- nut
|
|
|
|
|
- nut-client
|
|
|
|
|
- nut-server
|
|
|
|
|
state: present
|
|
|
|
|
update_cache: true
|
|
|
|
|
|
|
|
|
|
# ------------------------------------------------------------------
|
|
|
|
|
# Verify UPS is detected
|
|
|
|
|
# ------------------------------------------------------------------
|
|
|
|
|
- name: Check if UPS is detected via USB
|
|
|
|
|
shell: lsusb | grep -i cyber
|
|
|
|
|
register: lsusb_output
|
|
|
|
|
changed_when: false
|
|
|
|
|
failed_when: false
|
|
|
|
|
|
|
|
|
|
- name: Display USB detection result
|
|
|
|
|
debug:
|
|
|
|
|
msg: "{{ lsusb_output.stdout | default('UPS not detected via USB - ensure it is plugged in') }}"
|
|
|
|
|
|
|
|
|
|
- name: Fail if UPS not detected
|
|
|
|
|
fail:
|
|
|
|
|
msg: "CyberPower UPS not detected via USB. Ensure the USB cable is connected."
|
|
|
|
|
when: lsusb_output.rc != 0
|
|
|
|
|
|
|
|
|
|
- name: Reload udev rules for USB permissions
|
|
|
|
|
shell: |
|
|
|
|
|
udevadm control --reload-rules
|
|
|
|
|
udevadm trigger --subsystem-match=usb --action=add
|
|
|
|
|
changed_when: true
|
|
|
|
|
|
|
|
|
|
- name: Verify USB device has nut group permissions
|
|
|
|
|
shell: |
|
|
|
|
|
BUS_DEV=$(lsusb | grep -i cyber | grep -oP 'Bus \K\d+|Device \K\d+' | tr '\n' '/' | sed 's/\/$//')
|
|
|
|
|
if [ -n "$BUS_DEV" ]; then
|
|
|
|
|
BUS=$(echo $BUS_DEV | cut -d'/' -f1)
|
|
|
|
|
DEV=$(echo $BUS_DEV | cut -d'/' -f2)
|
|
|
|
|
ls -la /dev/bus/usb/$BUS/$DEV
|
|
|
|
|
else
|
|
|
|
|
echo "UPS device not found"
|
|
|
|
|
exit 1
|
|
|
|
|
fi
|
|
|
|
|
register: usb_permissions
|
|
|
|
|
changed_when: false
|
|
|
|
|
|
|
|
|
|
- name: Display USB permissions
|
|
|
|
|
debug:
|
|
|
|
|
msg: "{{ usb_permissions.stdout }} (should show 'root nut', not 'root root')"
|
|
|
|
|
|
|
|
|
|
- name: Scan for UPS with nut-scanner
|
|
|
|
|
command: nut-scanner -U
|
|
|
|
|
register: nut_scanner_output
|
|
|
|
|
changed_when: false
|
|
|
|
|
failed_when: false
|
|
|
|
|
|
|
|
|
|
- name: Display nut-scanner result
|
|
|
|
|
debug:
|
|
|
|
|
msg: "{{ nut_scanner_output.stdout_lines }}"
|
|
|
|
|
|
|
|
|
|
# ------------------------------------------------------------------
|
|
|
|
|
# Configuration files
|
|
|
|
|
# ------------------------------------------------------------------
|
|
|
|
|
- name: Configure NUT mode (standalone)
|
nodito: de-Uptime-Kuma the ZFS and NUT playbooks, extract templates
These two are host-specific by design - nodito is a pet, not cattle - so they
stay playbooks rather than becoming roles. But they had rotted.
De-Kuma, following the pattern of the six service roles:
Both plays opened with an assert on uptime_kuma_username/password, which were
removed from the vault, so both failed before doing anything. Dropped that,
the two embedded Python monitor-creation scripts, and their /tmp cleanup.
Kept every check, threshold and systemd timer - those are the durable part.
Reporting is now generic: `healthcheck_push_url` goes into the unit as
Environment=HEALTHCHECK_PUSH_URL and the script reads ${HEALTHCHECK_PUSH_URL:-},
treating empty as normal rather than an error. The exit code is the real
answer; systemd keeps it. Both scripts now also report status=down on failure
instead of only going silent. Live push URLs harvested into the vault so
nothing observable changes for ZFS.
Three live bugs found while check-diffing:
1. 32_zfs would have DE-REGISTERED the Proxmox storage. `pvesm remove` was
gated on the storage existing and `pvesm add` on it NOT existing - mutually
exclusive - so a real run removed the proxmox-tank-1 entry backing every VM
and never put it back. It would also have dropped `mountpoint /var/lib/vz`,
which the live entry has and `pvesm add` does not set. Registration is now
add-only.
2. zfs_disk_1 named ata-...WX11TN0Z, a disk no longer in the machine. The live
mirror is WX120LHQ + WX11TN2P; a leg was replaced and the repo never caught
up. Inert behind `when: zfs_pool_exists.rc != 0`, but wrong on any
disaster-recovery run. This is the seventh instance of an identifier
written down once whose hardware later moved.
3. 34_nut has NEVER been applied to nodito - no /etc/nut file carries the
"Managed by Ansible" marker; they were written by hand in January 2026. The
vault held the literal CHANGE_ME_TO_SECURE_PASSWORD, so applying it would
have overwritten a working upsd/upsmon auth pair with a placeholder and
restarted NUT, leaving the hypervisor's UPS unable to trigger a clean
shutdown on mains loss. The Kuma assert was the only thing stopping that,
so removing it without a replacement would have armed the gun: there is now
an explicit assert that refuses to run on the placeholder. The real
password is in the vault and `Configure upsd users` check-diffs clean.
Templates reconciled with the live files first, so applying 34_nut is close to
a no-op: added `maxretry = 3` to ups.conf and OFFDURATION / RBWARNTIME /
NOCOMMWARNTIME / FINALDELAY plus quoted POWERDOWNFLAG to upsmon.conf. The only
substantive additions left are the NOTIFYMSG/NOTIFYFLAG syslog lines.
/usr/local/bin/ups-heartbeat.sh on the box is an orphan - mode 0644, not
executable, referenced by no unit and no cron entry - but its push token
belongs to a monitor that still exists and answers, so that monitor has had no
heartbeat since January. Harvested as healthcheck_push_urls.ups; applying this
play is what will finally feed it.
Finish the host_vars migration:
infra/nodito/nodito_vars.yml was byte-identical to host_vars/nodito/main.yml.
Deleted it, moved nodito_secrets.yml to host_vars/nodito/vault.yml, and
stripped the dead vars_files entries from all three playbooks.
Extract the 13 inline `content: |` blocks to infra/nodito/templates/, pulled via
a YAML load rather than retyped. 681+582 lines become 313+344 plus templates.
Ownership parity against HEAD checked mechanically: no owner/group/mode drift on
any surviving task.
Neither playbook has been applied. ZFS play 2 check-runs failed=0 with two
changes: the script rewrite (the deployed one is a hand-edited DEBUG VERSION
with `set -x`) and the Environment line.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 19:07:50 +02:00
|
|
|
template:
|
2026-01-11 22:43:27 +01:00
|
|
|
dest: /etc/nut/nut.conf
|
nodito: de-Uptime-Kuma the ZFS and NUT playbooks, extract templates
These two are host-specific by design - nodito is a pet, not cattle - so they
stay playbooks rather than becoming roles. But they had rotted.
De-Kuma, following the pattern of the six service roles:
Both plays opened with an assert on uptime_kuma_username/password, which were
removed from the vault, so both failed before doing anything. Dropped that,
the two embedded Python monitor-creation scripts, and their /tmp cleanup.
Kept every check, threshold and systemd timer - those are the durable part.
Reporting is now generic: `healthcheck_push_url` goes into the unit as
Environment=HEALTHCHECK_PUSH_URL and the script reads ${HEALTHCHECK_PUSH_URL:-},
treating empty as normal rather than an error. The exit code is the real
answer; systemd keeps it. Both scripts now also report status=down on failure
instead of only going silent. Live push URLs harvested into the vault so
nothing observable changes for ZFS.
Three live bugs found while check-diffing:
1. 32_zfs would have DE-REGISTERED the Proxmox storage. `pvesm remove` was
gated on the storage existing and `pvesm add` on it NOT existing - mutually
exclusive - so a real run removed the proxmox-tank-1 entry backing every VM
and never put it back. It would also have dropped `mountpoint /var/lib/vz`,
which the live entry has and `pvesm add` does not set. Registration is now
add-only.
2. zfs_disk_1 named ata-...WX11TN0Z, a disk no longer in the machine. The live
mirror is WX120LHQ + WX11TN2P; a leg was replaced and the repo never caught
up. Inert behind `when: zfs_pool_exists.rc != 0`, but wrong on any
disaster-recovery run. This is the seventh instance of an identifier
written down once whose hardware later moved.
3. 34_nut has NEVER been applied to nodito - no /etc/nut file carries the
"Managed by Ansible" marker; they were written by hand in January 2026. The
vault held the literal CHANGE_ME_TO_SECURE_PASSWORD, so applying it would
have overwritten a working upsd/upsmon auth pair with a placeholder and
restarted NUT, leaving the hypervisor's UPS unable to trigger a clean
shutdown on mains loss. The Kuma assert was the only thing stopping that,
so removing it without a replacement would have armed the gun: there is now
an explicit assert that refuses to run on the placeholder. The real
password is in the vault and `Configure upsd users` check-diffs clean.
Templates reconciled with the live files first, so applying 34_nut is close to
a no-op: added `maxretry = 3` to ups.conf and OFFDURATION / RBWARNTIME /
NOCOMMWARNTIME / FINALDELAY plus quoted POWERDOWNFLAG to upsmon.conf. The only
substantive additions left are the NOTIFYMSG/NOTIFYFLAG syslog lines.
/usr/local/bin/ups-heartbeat.sh on the box is an orphan - mode 0644, not
executable, referenced by no unit and no cron entry - but its push token
belongs to a monitor that still exists and answers, so that monitor has had no
heartbeat since January. Harvested as healthcheck_push_urls.ups; applying this
play is what will finally feed it.
Finish the host_vars migration:
infra/nodito/nodito_vars.yml was byte-identical to host_vars/nodito/main.yml.
Deleted it, moved nodito_secrets.yml to host_vars/nodito/vault.yml, and
stripped the dead vars_files entries from all three playbooks.
Extract the 13 inline `content: |` blocks to infra/nodito/templates/, pulled via
a YAML load rather than retyped. 681+582 lines become 313+344 plus templates.
Ownership parity against HEAD checked mechanically: no owner/group/mode drift on
any surviving task.
Neither playbook has been applied. ZFS play 2 check-runs failed=0 with two
changes: the script rewrite (the deployed one is a hand-edited DEBUG VERSION
with `set -x`) and the Environment line.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 19:07:50 +02:00
|
|
|
src: templates/nut.conf.j2
|
2026-01-11 22:43:27 +01:00
|
|
|
owner: root
|
|
|
|
|
group: nut
|
|
|
|
|
mode: "0640"
|
|
|
|
|
notify: Restart NUT services
|
|
|
|
|
|
|
|
|
|
- name: Configure UPS device
|
nodito: de-Uptime-Kuma the ZFS and NUT playbooks, extract templates
These two are host-specific by design - nodito is a pet, not cattle - so they
stay playbooks rather than becoming roles. But they had rotted.
De-Kuma, following the pattern of the six service roles:
Both plays opened with an assert on uptime_kuma_username/password, which were
removed from the vault, so both failed before doing anything. Dropped that,
the two embedded Python monitor-creation scripts, and their /tmp cleanup.
Kept every check, threshold and systemd timer - those are the durable part.
Reporting is now generic: `healthcheck_push_url` goes into the unit as
Environment=HEALTHCHECK_PUSH_URL and the script reads ${HEALTHCHECK_PUSH_URL:-},
treating empty as normal rather than an error. The exit code is the real
answer; systemd keeps it. Both scripts now also report status=down on failure
instead of only going silent. Live push URLs harvested into the vault so
nothing observable changes for ZFS.
Three live bugs found while check-diffing:
1. 32_zfs would have DE-REGISTERED the Proxmox storage. `pvesm remove` was
gated on the storage existing and `pvesm add` on it NOT existing - mutually
exclusive - so a real run removed the proxmox-tank-1 entry backing every VM
and never put it back. It would also have dropped `mountpoint /var/lib/vz`,
which the live entry has and `pvesm add` does not set. Registration is now
add-only.
2. zfs_disk_1 named ata-...WX11TN0Z, a disk no longer in the machine. The live
mirror is WX120LHQ + WX11TN2P; a leg was replaced and the repo never caught
up. Inert behind `when: zfs_pool_exists.rc != 0`, but wrong on any
disaster-recovery run. This is the seventh instance of an identifier
written down once whose hardware later moved.
3. 34_nut has NEVER been applied to nodito - no /etc/nut file carries the
"Managed by Ansible" marker; they were written by hand in January 2026. The
vault held the literal CHANGE_ME_TO_SECURE_PASSWORD, so applying it would
have overwritten a working upsd/upsmon auth pair with a placeholder and
restarted NUT, leaving the hypervisor's UPS unable to trigger a clean
shutdown on mains loss. The Kuma assert was the only thing stopping that,
so removing it without a replacement would have armed the gun: there is now
an explicit assert that refuses to run on the placeholder. The real
password is in the vault and `Configure upsd users` check-diffs clean.
Templates reconciled with the live files first, so applying 34_nut is close to
a no-op: added `maxretry = 3` to ups.conf and OFFDURATION / RBWARNTIME /
NOCOMMWARNTIME / FINALDELAY plus quoted POWERDOWNFLAG to upsmon.conf. The only
substantive additions left are the NOTIFYMSG/NOTIFYFLAG syslog lines.
/usr/local/bin/ups-heartbeat.sh on the box is an orphan - mode 0644, not
executable, referenced by no unit and no cron entry - but its push token
belongs to a monitor that still exists and answers, so that monitor has had no
heartbeat since January. Harvested as healthcheck_push_urls.ups; applying this
play is what will finally feed it.
Finish the host_vars migration:
infra/nodito/nodito_vars.yml was byte-identical to host_vars/nodito/main.yml.
Deleted it, moved nodito_secrets.yml to host_vars/nodito/vault.yml, and
stripped the dead vars_files entries from all three playbooks.
Extract the 13 inline `content: |` blocks to infra/nodito/templates/, pulled via
a YAML load rather than retyped. 681+582 lines become 313+344 plus templates.
Ownership parity against HEAD checked mechanically: no owner/group/mode drift on
any surviving task.
Neither playbook has been applied. ZFS play 2 check-runs failed=0 with two
changes: the script rewrite (the deployed one is a hand-edited DEBUG VERSION
with `set -x`) and the Environment line.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 19:07:50 +02:00
|
|
|
template:
|
2026-01-11 22:43:27 +01:00
|
|
|
dest: /etc/nut/ups.conf
|
nodito: de-Uptime-Kuma the ZFS and NUT playbooks, extract templates
These two are host-specific by design - nodito is a pet, not cattle - so they
stay playbooks rather than becoming roles. But they had rotted.
De-Kuma, following the pattern of the six service roles:
Both plays opened with an assert on uptime_kuma_username/password, which were
removed from the vault, so both failed before doing anything. Dropped that,
the two embedded Python monitor-creation scripts, and their /tmp cleanup.
Kept every check, threshold and systemd timer - those are the durable part.
Reporting is now generic: `healthcheck_push_url` goes into the unit as
Environment=HEALTHCHECK_PUSH_URL and the script reads ${HEALTHCHECK_PUSH_URL:-},
treating empty as normal rather than an error. The exit code is the real
answer; systemd keeps it. Both scripts now also report status=down on failure
instead of only going silent. Live push URLs harvested into the vault so
nothing observable changes for ZFS.
Three live bugs found while check-diffing:
1. 32_zfs would have DE-REGISTERED the Proxmox storage. `pvesm remove` was
gated on the storage existing and `pvesm add` on it NOT existing - mutually
exclusive - so a real run removed the proxmox-tank-1 entry backing every VM
and never put it back. It would also have dropped `mountpoint /var/lib/vz`,
which the live entry has and `pvesm add` does not set. Registration is now
add-only.
2. zfs_disk_1 named ata-...WX11TN0Z, a disk no longer in the machine. The live
mirror is WX120LHQ + WX11TN2P; a leg was replaced and the repo never caught
up. Inert behind `when: zfs_pool_exists.rc != 0`, but wrong on any
disaster-recovery run. This is the seventh instance of an identifier
written down once whose hardware later moved.
3. 34_nut has NEVER been applied to nodito - no /etc/nut file carries the
"Managed by Ansible" marker; they were written by hand in January 2026. The
vault held the literal CHANGE_ME_TO_SECURE_PASSWORD, so applying it would
have overwritten a working upsd/upsmon auth pair with a placeholder and
restarted NUT, leaving the hypervisor's UPS unable to trigger a clean
shutdown on mains loss. The Kuma assert was the only thing stopping that,
so removing it without a replacement would have armed the gun: there is now
an explicit assert that refuses to run on the placeholder. The real
password is in the vault and `Configure upsd users` check-diffs clean.
Templates reconciled with the live files first, so applying 34_nut is close to
a no-op: added `maxretry = 3` to ups.conf and OFFDURATION / RBWARNTIME /
NOCOMMWARNTIME / FINALDELAY plus quoted POWERDOWNFLAG to upsmon.conf. The only
substantive additions left are the NOTIFYMSG/NOTIFYFLAG syslog lines.
/usr/local/bin/ups-heartbeat.sh on the box is an orphan - mode 0644, not
executable, referenced by no unit and no cron entry - but its push token
belongs to a monitor that still exists and answers, so that monitor has had no
heartbeat since January. Harvested as healthcheck_push_urls.ups; applying this
play is what will finally feed it.
Finish the host_vars migration:
infra/nodito/nodito_vars.yml was byte-identical to host_vars/nodito/main.yml.
Deleted it, moved nodito_secrets.yml to host_vars/nodito/vault.yml, and
stripped the dead vars_files entries from all three playbooks.
Extract the 13 inline `content: |` blocks to infra/nodito/templates/, pulled via
a YAML load rather than retyped. 681+582 lines become 313+344 plus templates.
Ownership parity against HEAD checked mechanically: no owner/group/mode drift on
any surviving task.
Neither playbook has been applied. ZFS play 2 check-runs failed=0 with two
changes: the script rewrite (the deployed one is a hand-edited DEBUG VERSION
with `set -x`) and the Environment line.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 19:07:50 +02:00
|
|
|
src: templates/ups.conf.j2
|
2026-01-11 22:43:27 +01:00
|
|
|
owner: root
|
|
|
|
|
group: nut
|
|
|
|
|
mode: "0640"
|
|
|
|
|
notify: Restart NUT services
|
|
|
|
|
|
|
|
|
|
- name: Configure upsd to listen on localhost
|
nodito: de-Uptime-Kuma the ZFS and NUT playbooks, extract templates
These two are host-specific by design - nodito is a pet, not cattle - so they
stay playbooks rather than becoming roles. But they had rotted.
De-Kuma, following the pattern of the six service roles:
Both plays opened with an assert on uptime_kuma_username/password, which were
removed from the vault, so both failed before doing anything. Dropped that,
the two embedded Python monitor-creation scripts, and their /tmp cleanup.
Kept every check, threshold and systemd timer - those are the durable part.
Reporting is now generic: `healthcheck_push_url` goes into the unit as
Environment=HEALTHCHECK_PUSH_URL and the script reads ${HEALTHCHECK_PUSH_URL:-},
treating empty as normal rather than an error. The exit code is the real
answer; systemd keeps it. Both scripts now also report status=down on failure
instead of only going silent. Live push URLs harvested into the vault so
nothing observable changes for ZFS.
Three live bugs found while check-diffing:
1. 32_zfs would have DE-REGISTERED the Proxmox storage. `pvesm remove` was
gated on the storage existing and `pvesm add` on it NOT existing - mutually
exclusive - so a real run removed the proxmox-tank-1 entry backing every VM
and never put it back. It would also have dropped `mountpoint /var/lib/vz`,
which the live entry has and `pvesm add` does not set. Registration is now
add-only.
2. zfs_disk_1 named ata-...WX11TN0Z, a disk no longer in the machine. The live
mirror is WX120LHQ + WX11TN2P; a leg was replaced and the repo never caught
up. Inert behind `when: zfs_pool_exists.rc != 0`, but wrong on any
disaster-recovery run. This is the seventh instance of an identifier
written down once whose hardware later moved.
3. 34_nut has NEVER been applied to nodito - no /etc/nut file carries the
"Managed by Ansible" marker; they were written by hand in January 2026. The
vault held the literal CHANGE_ME_TO_SECURE_PASSWORD, so applying it would
have overwritten a working upsd/upsmon auth pair with a placeholder and
restarted NUT, leaving the hypervisor's UPS unable to trigger a clean
shutdown on mains loss. The Kuma assert was the only thing stopping that,
so removing it without a replacement would have armed the gun: there is now
an explicit assert that refuses to run on the placeholder. The real
password is in the vault and `Configure upsd users` check-diffs clean.
Templates reconciled with the live files first, so applying 34_nut is close to
a no-op: added `maxretry = 3` to ups.conf and OFFDURATION / RBWARNTIME /
NOCOMMWARNTIME / FINALDELAY plus quoted POWERDOWNFLAG to upsmon.conf. The only
substantive additions left are the NOTIFYMSG/NOTIFYFLAG syslog lines.
/usr/local/bin/ups-heartbeat.sh on the box is an orphan - mode 0644, not
executable, referenced by no unit and no cron entry - but its push token
belongs to a monitor that still exists and answers, so that monitor has had no
heartbeat since January. Harvested as healthcheck_push_urls.ups; applying this
play is what will finally feed it.
Finish the host_vars migration:
infra/nodito/nodito_vars.yml was byte-identical to host_vars/nodito/main.yml.
Deleted it, moved nodito_secrets.yml to host_vars/nodito/vault.yml, and
stripped the dead vars_files entries from all three playbooks.
Extract the 13 inline `content: |` blocks to infra/nodito/templates/, pulled via
a YAML load rather than retyped. 681+582 lines become 313+344 plus templates.
Ownership parity against HEAD checked mechanically: no owner/group/mode drift on
any surviving task.
Neither playbook has been applied. ZFS play 2 check-runs failed=0 with two
changes: the script rewrite (the deployed one is a hand-edited DEBUG VERSION
with `set -x`) and the Environment line.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 19:07:50 +02:00
|
|
|
template:
|
2026-01-11 22:43:27 +01:00
|
|
|
dest: /etc/nut/upsd.conf
|
nodito: de-Uptime-Kuma the ZFS and NUT playbooks, extract templates
These two are host-specific by design - nodito is a pet, not cattle - so they
stay playbooks rather than becoming roles. But they had rotted.
De-Kuma, following the pattern of the six service roles:
Both plays opened with an assert on uptime_kuma_username/password, which were
removed from the vault, so both failed before doing anything. Dropped that,
the two embedded Python monitor-creation scripts, and their /tmp cleanup.
Kept every check, threshold and systemd timer - those are the durable part.
Reporting is now generic: `healthcheck_push_url` goes into the unit as
Environment=HEALTHCHECK_PUSH_URL and the script reads ${HEALTHCHECK_PUSH_URL:-},
treating empty as normal rather than an error. The exit code is the real
answer; systemd keeps it. Both scripts now also report status=down on failure
instead of only going silent. Live push URLs harvested into the vault so
nothing observable changes for ZFS.
Three live bugs found while check-diffing:
1. 32_zfs would have DE-REGISTERED the Proxmox storage. `pvesm remove` was
gated on the storage existing and `pvesm add` on it NOT existing - mutually
exclusive - so a real run removed the proxmox-tank-1 entry backing every VM
and never put it back. It would also have dropped `mountpoint /var/lib/vz`,
which the live entry has and `pvesm add` does not set. Registration is now
add-only.
2. zfs_disk_1 named ata-...WX11TN0Z, a disk no longer in the machine. The live
mirror is WX120LHQ + WX11TN2P; a leg was replaced and the repo never caught
up. Inert behind `when: zfs_pool_exists.rc != 0`, but wrong on any
disaster-recovery run. This is the seventh instance of an identifier
written down once whose hardware later moved.
3. 34_nut has NEVER been applied to nodito - no /etc/nut file carries the
"Managed by Ansible" marker; they were written by hand in January 2026. The
vault held the literal CHANGE_ME_TO_SECURE_PASSWORD, so applying it would
have overwritten a working upsd/upsmon auth pair with a placeholder and
restarted NUT, leaving the hypervisor's UPS unable to trigger a clean
shutdown on mains loss. The Kuma assert was the only thing stopping that,
so removing it without a replacement would have armed the gun: there is now
an explicit assert that refuses to run on the placeholder. The real
password is in the vault and `Configure upsd users` check-diffs clean.
Templates reconciled with the live files first, so applying 34_nut is close to
a no-op: added `maxretry = 3` to ups.conf and OFFDURATION / RBWARNTIME /
NOCOMMWARNTIME / FINALDELAY plus quoted POWERDOWNFLAG to upsmon.conf. The only
substantive additions left are the NOTIFYMSG/NOTIFYFLAG syslog lines.
/usr/local/bin/ups-heartbeat.sh on the box is an orphan - mode 0644, not
executable, referenced by no unit and no cron entry - but its push token
belongs to a monitor that still exists and answers, so that monitor has had no
heartbeat since January. Harvested as healthcheck_push_urls.ups; applying this
play is what will finally feed it.
Finish the host_vars migration:
infra/nodito/nodito_vars.yml was byte-identical to host_vars/nodito/main.yml.
Deleted it, moved nodito_secrets.yml to host_vars/nodito/vault.yml, and
stripped the dead vars_files entries from all three playbooks.
Extract the 13 inline `content: |` blocks to infra/nodito/templates/, pulled via
a YAML load rather than retyped. 681+582 lines become 313+344 plus templates.
Ownership parity against HEAD checked mechanically: no owner/group/mode drift on
any surviving task.
Neither playbook has been applied. ZFS play 2 check-runs failed=0 with two
changes: the script rewrite (the deployed one is a hand-edited DEBUG VERSION
with `set -x`) and the Environment line.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 19:07:50 +02:00
|
|
|
src: templates/upsd.conf.j2
|
2026-01-11 22:43:27 +01:00
|
|
|
owner: root
|
|
|
|
|
group: nut
|
|
|
|
|
mode: "0640"
|
|
|
|
|
notify: Restart NUT services
|
|
|
|
|
|
|
|
|
|
- name: Configure upsd users
|
nodito: de-Uptime-Kuma the ZFS and NUT playbooks, extract templates
These two are host-specific by design - nodito is a pet, not cattle - so they
stay playbooks rather than becoming roles. But they had rotted.
De-Kuma, following the pattern of the six service roles:
Both plays opened with an assert on uptime_kuma_username/password, which were
removed from the vault, so both failed before doing anything. Dropped that,
the two embedded Python monitor-creation scripts, and their /tmp cleanup.
Kept every check, threshold and systemd timer - those are the durable part.
Reporting is now generic: `healthcheck_push_url` goes into the unit as
Environment=HEALTHCHECK_PUSH_URL and the script reads ${HEALTHCHECK_PUSH_URL:-},
treating empty as normal rather than an error. The exit code is the real
answer; systemd keeps it. Both scripts now also report status=down on failure
instead of only going silent. Live push URLs harvested into the vault so
nothing observable changes for ZFS.
Three live bugs found while check-diffing:
1. 32_zfs would have DE-REGISTERED the Proxmox storage. `pvesm remove` was
gated on the storage existing and `pvesm add` on it NOT existing - mutually
exclusive - so a real run removed the proxmox-tank-1 entry backing every VM
and never put it back. It would also have dropped `mountpoint /var/lib/vz`,
which the live entry has and `pvesm add` does not set. Registration is now
add-only.
2. zfs_disk_1 named ata-...WX11TN0Z, a disk no longer in the machine. The live
mirror is WX120LHQ + WX11TN2P; a leg was replaced and the repo never caught
up. Inert behind `when: zfs_pool_exists.rc != 0`, but wrong on any
disaster-recovery run. This is the seventh instance of an identifier
written down once whose hardware later moved.
3. 34_nut has NEVER been applied to nodito - no /etc/nut file carries the
"Managed by Ansible" marker; they were written by hand in January 2026. The
vault held the literal CHANGE_ME_TO_SECURE_PASSWORD, so applying it would
have overwritten a working upsd/upsmon auth pair with a placeholder and
restarted NUT, leaving the hypervisor's UPS unable to trigger a clean
shutdown on mains loss. The Kuma assert was the only thing stopping that,
so removing it without a replacement would have armed the gun: there is now
an explicit assert that refuses to run on the placeholder. The real
password is in the vault and `Configure upsd users` check-diffs clean.
Templates reconciled with the live files first, so applying 34_nut is close to
a no-op: added `maxretry = 3` to ups.conf and OFFDURATION / RBWARNTIME /
NOCOMMWARNTIME / FINALDELAY plus quoted POWERDOWNFLAG to upsmon.conf. The only
substantive additions left are the NOTIFYMSG/NOTIFYFLAG syslog lines.
/usr/local/bin/ups-heartbeat.sh on the box is an orphan - mode 0644, not
executable, referenced by no unit and no cron entry - but its push token
belongs to a monitor that still exists and answers, so that monitor has had no
heartbeat since January. Harvested as healthcheck_push_urls.ups; applying this
play is what will finally feed it.
Finish the host_vars migration:
infra/nodito/nodito_vars.yml was byte-identical to host_vars/nodito/main.yml.
Deleted it, moved nodito_secrets.yml to host_vars/nodito/vault.yml, and
stripped the dead vars_files entries from all three playbooks.
Extract the 13 inline `content: |` blocks to infra/nodito/templates/, pulled via
a YAML load rather than retyped. 681+582 lines become 313+344 plus templates.
Ownership parity against HEAD checked mechanically: no owner/group/mode drift on
any surviving task.
Neither playbook has been applied. ZFS play 2 check-runs failed=0 with two
changes: the script rewrite (the deployed one is a hand-edited DEBUG VERSION
with `set -x`) and the Environment line.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 19:07:50 +02:00
|
|
|
template:
|
2026-01-11 22:43:27 +01:00
|
|
|
dest: /etc/nut/upsd.users
|
nodito: de-Uptime-Kuma the ZFS and NUT playbooks, extract templates
These two are host-specific by design - nodito is a pet, not cattle - so they
stay playbooks rather than becoming roles. But they had rotted.
De-Kuma, following the pattern of the six service roles:
Both plays opened with an assert on uptime_kuma_username/password, which were
removed from the vault, so both failed before doing anything. Dropped that,
the two embedded Python monitor-creation scripts, and their /tmp cleanup.
Kept every check, threshold and systemd timer - those are the durable part.
Reporting is now generic: `healthcheck_push_url` goes into the unit as
Environment=HEALTHCHECK_PUSH_URL and the script reads ${HEALTHCHECK_PUSH_URL:-},
treating empty as normal rather than an error. The exit code is the real
answer; systemd keeps it. Both scripts now also report status=down on failure
instead of only going silent. Live push URLs harvested into the vault so
nothing observable changes for ZFS.
Three live bugs found while check-diffing:
1. 32_zfs would have DE-REGISTERED the Proxmox storage. `pvesm remove` was
gated on the storage existing and `pvesm add` on it NOT existing - mutually
exclusive - so a real run removed the proxmox-tank-1 entry backing every VM
and never put it back. It would also have dropped `mountpoint /var/lib/vz`,
which the live entry has and `pvesm add` does not set. Registration is now
add-only.
2. zfs_disk_1 named ata-...WX11TN0Z, a disk no longer in the machine. The live
mirror is WX120LHQ + WX11TN2P; a leg was replaced and the repo never caught
up. Inert behind `when: zfs_pool_exists.rc != 0`, but wrong on any
disaster-recovery run. This is the seventh instance of an identifier
written down once whose hardware later moved.
3. 34_nut has NEVER been applied to nodito - no /etc/nut file carries the
"Managed by Ansible" marker; they were written by hand in January 2026. The
vault held the literal CHANGE_ME_TO_SECURE_PASSWORD, so applying it would
have overwritten a working upsd/upsmon auth pair with a placeholder and
restarted NUT, leaving the hypervisor's UPS unable to trigger a clean
shutdown on mains loss. The Kuma assert was the only thing stopping that,
so removing it without a replacement would have armed the gun: there is now
an explicit assert that refuses to run on the placeholder. The real
password is in the vault and `Configure upsd users` check-diffs clean.
Templates reconciled with the live files first, so applying 34_nut is close to
a no-op: added `maxretry = 3` to ups.conf and OFFDURATION / RBWARNTIME /
NOCOMMWARNTIME / FINALDELAY plus quoted POWERDOWNFLAG to upsmon.conf. The only
substantive additions left are the NOTIFYMSG/NOTIFYFLAG syslog lines.
/usr/local/bin/ups-heartbeat.sh on the box is an orphan - mode 0644, not
executable, referenced by no unit and no cron entry - but its push token
belongs to a monitor that still exists and answers, so that monitor has had no
heartbeat since January. Harvested as healthcheck_push_urls.ups; applying this
play is what will finally feed it.
Finish the host_vars migration:
infra/nodito/nodito_vars.yml was byte-identical to host_vars/nodito/main.yml.
Deleted it, moved nodito_secrets.yml to host_vars/nodito/vault.yml, and
stripped the dead vars_files entries from all three playbooks.
Extract the 13 inline `content: |` blocks to infra/nodito/templates/, pulled via
a YAML load rather than retyped. 681+582 lines become 313+344 plus templates.
Ownership parity against HEAD checked mechanically: no owner/group/mode drift on
any surviving task.
Neither playbook has been applied. ZFS play 2 check-runs failed=0 with two
changes: the script rewrite (the deployed one is a hand-edited DEBUG VERSION
with `set -x`) and the Environment line.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 19:07:50 +02:00
|
|
|
src: templates/upsd.users.j2
|
2026-01-11 22:43:27 +01:00
|
|
|
owner: root
|
|
|
|
|
group: nut
|
|
|
|
|
mode: "0640"
|
|
|
|
|
notify: Restart NUT services
|
|
|
|
|
|
|
|
|
|
- name: Configure upsmon
|
nodito: de-Uptime-Kuma the ZFS and NUT playbooks, extract templates
These two are host-specific by design - nodito is a pet, not cattle - so they
stay playbooks rather than becoming roles. But they had rotted.
De-Kuma, following the pattern of the six service roles:
Both plays opened with an assert on uptime_kuma_username/password, which were
removed from the vault, so both failed before doing anything. Dropped that,
the two embedded Python monitor-creation scripts, and their /tmp cleanup.
Kept every check, threshold and systemd timer - those are the durable part.
Reporting is now generic: `healthcheck_push_url` goes into the unit as
Environment=HEALTHCHECK_PUSH_URL and the script reads ${HEALTHCHECK_PUSH_URL:-},
treating empty as normal rather than an error. The exit code is the real
answer; systemd keeps it. Both scripts now also report status=down on failure
instead of only going silent. Live push URLs harvested into the vault so
nothing observable changes for ZFS.
Three live bugs found while check-diffing:
1. 32_zfs would have DE-REGISTERED the Proxmox storage. `pvesm remove` was
gated on the storage existing and `pvesm add` on it NOT existing - mutually
exclusive - so a real run removed the proxmox-tank-1 entry backing every VM
and never put it back. It would also have dropped `mountpoint /var/lib/vz`,
which the live entry has and `pvesm add` does not set. Registration is now
add-only.
2. zfs_disk_1 named ata-...WX11TN0Z, a disk no longer in the machine. The live
mirror is WX120LHQ + WX11TN2P; a leg was replaced and the repo never caught
up. Inert behind `when: zfs_pool_exists.rc != 0`, but wrong on any
disaster-recovery run. This is the seventh instance of an identifier
written down once whose hardware later moved.
3. 34_nut has NEVER been applied to nodito - no /etc/nut file carries the
"Managed by Ansible" marker; they were written by hand in January 2026. The
vault held the literal CHANGE_ME_TO_SECURE_PASSWORD, so applying it would
have overwritten a working upsd/upsmon auth pair with a placeholder and
restarted NUT, leaving the hypervisor's UPS unable to trigger a clean
shutdown on mains loss. The Kuma assert was the only thing stopping that,
so removing it without a replacement would have armed the gun: there is now
an explicit assert that refuses to run on the placeholder. The real
password is in the vault and `Configure upsd users` check-diffs clean.
Templates reconciled with the live files first, so applying 34_nut is close to
a no-op: added `maxretry = 3` to ups.conf and OFFDURATION / RBWARNTIME /
NOCOMMWARNTIME / FINALDELAY plus quoted POWERDOWNFLAG to upsmon.conf. The only
substantive additions left are the NOTIFYMSG/NOTIFYFLAG syslog lines.
/usr/local/bin/ups-heartbeat.sh on the box is an orphan - mode 0644, not
executable, referenced by no unit and no cron entry - but its push token
belongs to a monitor that still exists and answers, so that monitor has had no
heartbeat since January. Harvested as healthcheck_push_urls.ups; applying this
play is what will finally feed it.
Finish the host_vars migration:
infra/nodito/nodito_vars.yml was byte-identical to host_vars/nodito/main.yml.
Deleted it, moved nodito_secrets.yml to host_vars/nodito/vault.yml, and
stripped the dead vars_files entries from all three playbooks.
Extract the 13 inline `content: |` blocks to infra/nodito/templates/, pulled via
a YAML load rather than retyped. 681+582 lines become 313+344 plus templates.
Ownership parity against HEAD checked mechanically: no owner/group/mode drift on
any surviving task.
Neither playbook has been applied. ZFS play 2 check-runs failed=0 with two
changes: the script rewrite (the deployed one is a hand-edited DEBUG VERSION
with `set -x`) and the Environment line.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 19:07:50 +02:00
|
|
|
template:
|
2026-01-11 22:43:27 +01:00
|
|
|
dest: /etc/nut/upsmon.conf
|
nodito: de-Uptime-Kuma the ZFS and NUT playbooks, extract templates
These two are host-specific by design - nodito is a pet, not cattle - so they
stay playbooks rather than becoming roles. But they had rotted.
De-Kuma, following the pattern of the six service roles:
Both plays opened with an assert on uptime_kuma_username/password, which were
removed from the vault, so both failed before doing anything. Dropped that,
the two embedded Python monitor-creation scripts, and their /tmp cleanup.
Kept every check, threshold and systemd timer - those are the durable part.
Reporting is now generic: `healthcheck_push_url` goes into the unit as
Environment=HEALTHCHECK_PUSH_URL and the script reads ${HEALTHCHECK_PUSH_URL:-},
treating empty as normal rather than an error. The exit code is the real
answer; systemd keeps it. Both scripts now also report status=down on failure
instead of only going silent. Live push URLs harvested into the vault so
nothing observable changes for ZFS.
Three live bugs found while check-diffing:
1. 32_zfs would have DE-REGISTERED the Proxmox storage. `pvesm remove` was
gated on the storage existing and `pvesm add` on it NOT existing - mutually
exclusive - so a real run removed the proxmox-tank-1 entry backing every VM
and never put it back. It would also have dropped `mountpoint /var/lib/vz`,
which the live entry has and `pvesm add` does not set. Registration is now
add-only.
2. zfs_disk_1 named ata-...WX11TN0Z, a disk no longer in the machine. The live
mirror is WX120LHQ + WX11TN2P; a leg was replaced and the repo never caught
up. Inert behind `when: zfs_pool_exists.rc != 0`, but wrong on any
disaster-recovery run. This is the seventh instance of an identifier
written down once whose hardware later moved.
3. 34_nut has NEVER been applied to nodito - no /etc/nut file carries the
"Managed by Ansible" marker; they were written by hand in January 2026. The
vault held the literal CHANGE_ME_TO_SECURE_PASSWORD, so applying it would
have overwritten a working upsd/upsmon auth pair with a placeholder and
restarted NUT, leaving the hypervisor's UPS unable to trigger a clean
shutdown on mains loss. The Kuma assert was the only thing stopping that,
so removing it without a replacement would have armed the gun: there is now
an explicit assert that refuses to run on the placeholder. The real
password is in the vault and `Configure upsd users` check-diffs clean.
Templates reconciled with the live files first, so applying 34_nut is close to
a no-op: added `maxretry = 3` to ups.conf and OFFDURATION / RBWARNTIME /
NOCOMMWARNTIME / FINALDELAY plus quoted POWERDOWNFLAG to upsmon.conf. The only
substantive additions left are the NOTIFYMSG/NOTIFYFLAG syslog lines.
/usr/local/bin/ups-heartbeat.sh on the box is an orphan - mode 0644, not
executable, referenced by no unit and no cron entry - but its push token
belongs to a monitor that still exists and answers, so that monitor has had no
heartbeat since January. Harvested as healthcheck_push_urls.ups; applying this
play is what will finally feed it.
Finish the host_vars migration:
infra/nodito/nodito_vars.yml was byte-identical to host_vars/nodito/main.yml.
Deleted it, moved nodito_secrets.yml to host_vars/nodito/vault.yml, and
stripped the dead vars_files entries from all three playbooks.
Extract the 13 inline `content: |` blocks to infra/nodito/templates/, pulled via
a YAML load rather than retyped. 681+582 lines become 313+344 plus templates.
Ownership parity against HEAD checked mechanically: no owner/group/mode drift on
any surviving task.
Neither playbook has been applied. ZFS play 2 check-runs failed=0 with two
changes: the script rewrite (the deployed one is a hand-edited DEBUG VERSION
with `set -x`) and the Environment line.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 19:07:50 +02:00
|
|
|
src: templates/upsmon.conf.j2
|
2026-01-11 22:43:27 +01:00
|
|
|
owner: root
|
|
|
|
|
group: nut
|
|
|
|
|
mode: "0640"
|
|
|
|
|
notify: Restart NUT services
|
|
|
|
|
|
|
|
|
|
# ------------------------------------------------------------------
|
|
|
|
|
# Verify late-stage shutdown script
|
|
|
|
|
# ------------------------------------------------------------------
|
|
|
|
|
- name: Verify nutshutdown script exists
|
|
|
|
|
stat:
|
|
|
|
|
path: /lib/systemd/system-shutdown/nutshutdown
|
|
|
|
|
register: nutshutdown_script
|
|
|
|
|
|
|
|
|
|
- name: Warn if nutshutdown script is missing
|
|
|
|
|
debug:
|
|
|
|
|
msg: "WARNING: /lib/systemd/system-shutdown/nutshutdown not found. UPS may not cut power after shutdown."
|
|
|
|
|
when: not nutshutdown_script.stat.exists
|
|
|
|
|
|
|
|
|
|
# ------------------------------------------------------------------
|
|
|
|
|
# Services
|
|
|
|
|
# ------------------------------------------------------------------
|
|
|
|
|
- name: Enable and start NUT driver enumerator
|
|
|
|
|
systemd:
|
|
|
|
|
name: nut-driver-enumerator
|
|
|
|
|
enabled: true
|
|
|
|
|
state: started
|
|
|
|
|
|
|
|
|
|
- name: Enable and start NUT server
|
|
|
|
|
systemd:
|
|
|
|
|
name: nut-server
|
|
|
|
|
enabled: true
|
|
|
|
|
state: started
|
|
|
|
|
|
|
|
|
|
- name: Enable and start NUT monitor
|
|
|
|
|
systemd:
|
|
|
|
|
name: nut-monitor
|
|
|
|
|
enabled: true
|
|
|
|
|
state: started
|
|
|
|
|
|
|
|
|
|
# ------------------------------------------------------------------
|
|
|
|
|
# Verification
|
|
|
|
|
# ------------------------------------------------------------------
|
|
|
|
|
- name: Wait for NUT services to stabilize
|
|
|
|
|
pause:
|
|
|
|
|
seconds: 3
|
|
|
|
|
|
|
|
|
|
- name: Verify NUT can communicate with UPS
|
|
|
|
|
command: upsc {{ ups_name }}@localhost
|
|
|
|
|
register: upsc_output
|
|
|
|
|
changed_when: false
|
|
|
|
|
failed_when: upsc_output.rc != 0
|
|
|
|
|
|
|
|
|
|
- name: Display UPS status
|
|
|
|
|
debug:
|
|
|
|
|
msg: "{{ upsc_output.stdout_lines }}"
|
|
|
|
|
|
|
|
|
|
- name: Get UPS status summary
|
|
|
|
|
shell: |
|
|
|
|
|
echo "Status: $(upsc {{ ups_name }}@localhost ups.status 2>/dev/null)"
|
|
|
|
|
echo "Battery: $(upsc {{ ups_name }}@localhost battery.charge 2>/dev/null)%"
|
|
|
|
|
echo "Runtime: $(upsc {{ ups_name }}@localhost battery.runtime 2>/dev/null)s"
|
|
|
|
|
echo "Load: $(upsc {{ ups_name }}@localhost ups.load 2>/dev/null)%"
|
|
|
|
|
register: ups_summary
|
|
|
|
|
changed_when: false
|
|
|
|
|
|
|
|
|
|
- name: Display UPS summary
|
|
|
|
|
debug:
|
|
|
|
|
msg: "{{ ups_summary.stdout_lines }}"
|
|
|
|
|
|
|
|
|
|
- name: Verify low battery thresholds
|
|
|
|
|
shell: |
|
|
|
|
|
echo "Runtime threshold: $(upsc {{ ups_name }}@localhost battery.runtime.low 2>/dev/null)s"
|
|
|
|
|
echo "Charge threshold: $(upsc {{ ups_name }}@localhost battery.charge.low 2>/dev/null)%"
|
|
|
|
|
register: thresholds
|
|
|
|
|
changed_when: false
|
|
|
|
|
|
|
|
|
|
- name: Display low battery thresholds
|
|
|
|
|
debug:
|
|
|
|
|
msg: "{{ thresholds.stdout_lines }}"
|
|
|
|
|
|
|
|
|
|
handlers:
|
|
|
|
|
- name: Restart NUT services
|
|
|
|
|
systemd:
|
|
|
|
|
name: "{{ item }}"
|
|
|
|
|
state: restarted
|
|
|
|
|
loop:
|
|
|
|
|
- nut-driver-enumerator
|
|
|
|
|
- nut-server
|
|
|
|
|
- nut-monitor
|
|
|
|
|
|
|
|
|
|
|
nodito: de-Uptime-Kuma the ZFS and NUT playbooks, extract templates
These two are host-specific by design - nodito is a pet, not cattle - so they
stay playbooks rather than becoming roles. But they had rotted.
De-Kuma, following the pattern of the six service roles:
Both plays opened with an assert on uptime_kuma_username/password, which were
removed from the vault, so both failed before doing anything. Dropped that,
the two embedded Python monitor-creation scripts, and their /tmp cleanup.
Kept every check, threshold and systemd timer - those are the durable part.
Reporting is now generic: `healthcheck_push_url` goes into the unit as
Environment=HEALTHCHECK_PUSH_URL and the script reads ${HEALTHCHECK_PUSH_URL:-},
treating empty as normal rather than an error. The exit code is the real
answer; systemd keeps it. Both scripts now also report status=down on failure
instead of only going silent. Live push URLs harvested into the vault so
nothing observable changes for ZFS.
Three live bugs found while check-diffing:
1. 32_zfs would have DE-REGISTERED the Proxmox storage. `pvesm remove` was
gated on the storage existing and `pvesm add` on it NOT existing - mutually
exclusive - so a real run removed the proxmox-tank-1 entry backing every VM
and never put it back. It would also have dropped `mountpoint /var/lib/vz`,
which the live entry has and `pvesm add` does not set. Registration is now
add-only.
2. zfs_disk_1 named ata-...WX11TN0Z, a disk no longer in the machine. The live
mirror is WX120LHQ + WX11TN2P; a leg was replaced and the repo never caught
up. Inert behind `when: zfs_pool_exists.rc != 0`, but wrong on any
disaster-recovery run. This is the seventh instance of an identifier
written down once whose hardware later moved.
3. 34_nut has NEVER been applied to nodito - no /etc/nut file carries the
"Managed by Ansible" marker; they were written by hand in January 2026. The
vault held the literal CHANGE_ME_TO_SECURE_PASSWORD, so applying it would
have overwritten a working upsd/upsmon auth pair with a placeholder and
restarted NUT, leaving the hypervisor's UPS unable to trigger a clean
shutdown on mains loss. The Kuma assert was the only thing stopping that,
so removing it without a replacement would have armed the gun: there is now
an explicit assert that refuses to run on the placeholder. The real
password is in the vault and `Configure upsd users` check-diffs clean.
Templates reconciled with the live files first, so applying 34_nut is close to
a no-op: added `maxretry = 3` to ups.conf and OFFDURATION / RBWARNTIME /
NOCOMMWARNTIME / FINALDELAY plus quoted POWERDOWNFLAG to upsmon.conf. The only
substantive additions left are the NOTIFYMSG/NOTIFYFLAG syslog lines.
/usr/local/bin/ups-heartbeat.sh on the box is an orphan - mode 0644, not
executable, referenced by no unit and no cron entry - but its push token
belongs to a monitor that still exists and answers, so that monitor has had no
heartbeat since January. Harvested as healthcheck_push_urls.ups; applying this
play is what will finally feed it.
Finish the host_vars migration:
infra/nodito/nodito_vars.yml was byte-identical to host_vars/nodito/main.yml.
Deleted it, moved nodito_secrets.yml to host_vars/nodito/vault.yml, and
stripped the dead vars_files entries from all three playbooks.
Extract the 13 inline `content: |` blocks to infra/nodito/templates/, pulled via
a YAML load rather than retyped. 681+582 lines become 313+344 plus templates.
Ownership parity against HEAD checked mechanically: no owner/group/mode drift on
any surviving task.
Neither playbook has been applied. ZFS play 2 check-runs failed=0 with two
changes: the script rewrite (the deployed one is a hand-edited DEBUG VERSION
with `set -x`) and the Environment line.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 19:07:50 +02:00
|
|
|
# ─────────────────────────────────────────────────────────────────────────────
|
|
|
|
|
# UPS heartbeat monitoring.
|
2026-09-11 22:43:55 +02:00
|
|
|
#
|
nodito: de-Uptime-Kuma the ZFS and NUT playbooks, extract templates
These two are host-specific by design - nodito is a pet, not cattle - so they
stay playbooks rather than becoming roles. But they had rotted.
De-Kuma, following the pattern of the six service roles:
Both plays opened with an assert on uptime_kuma_username/password, which were
removed from the vault, so both failed before doing anything. Dropped that,
the two embedded Python monitor-creation scripts, and their /tmp cleanup.
Kept every check, threshold and systemd timer - those are the durable part.
Reporting is now generic: `healthcheck_push_url` goes into the unit as
Environment=HEALTHCHECK_PUSH_URL and the script reads ${HEALTHCHECK_PUSH_URL:-},
treating empty as normal rather than an error. The exit code is the real
answer; systemd keeps it. Both scripts now also report status=down on failure
instead of only going silent. Live push URLs harvested into the vault so
nothing observable changes for ZFS.
Three live bugs found while check-diffing:
1. 32_zfs would have DE-REGISTERED the Proxmox storage. `pvesm remove` was
gated on the storage existing and `pvesm add` on it NOT existing - mutually
exclusive - so a real run removed the proxmox-tank-1 entry backing every VM
and never put it back. It would also have dropped `mountpoint /var/lib/vz`,
which the live entry has and `pvesm add` does not set. Registration is now
add-only.
2. zfs_disk_1 named ata-...WX11TN0Z, a disk no longer in the machine. The live
mirror is WX120LHQ + WX11TN2P; a leg was replaced and the repo never caught
up. Inert behind `when: zfs_pool_exists.rc != 0`, but wrong on any
disaster-recovery run. This is the seventh instance of an identifier
written down once whose hardware later moved.
3. 34_nut has NEVER been applied to nodito - no /etc/nut file carries the
"Managed by Ansible" marker; they were written by hand in January 2026. The
vault held the literal CHANGE_ME_TO_SECURE_PASSWORD, so applying it would
have overwritten a working upsd/upsmon auth pair with a placeholder and
restarted NUT, leaving the hypervisor's UPS unable to trigger a clean
shutdown on mains loss. The Kuma assert was the only thing stopping that,
so removing it without a replacement would have armed the gun: there is now
an explicit assert that refuses to run on the placeholder. The real
password is in the vault and `Configure upsd users` check-diffs clean.
Templates reconciled with the live files first, so applying 34_nut is close to
a no-op: added `maxretry = 3` to ups.conf and OFFDURATION / RBWARNTIME /
NOCOMMWARNTIME / FINALDELAY plus quoted POWERDOWNFLAG to upsmon.conf. The only
substantive additions left are the NOTIFYMSG/NOTIFYFLAG syslog lines.
/usr/local/bin/ups-heartbeat.sh on the box is an orphan - mode 0644, not
executable, referenced by no unit and no cron entry - but its push token
belongs to a monitor that still exists and answers, so that monitor has had no
heartbeat since January. Harvested as healthcheck_push_urls.ups; applying this
play is what will finally feed it.
Finish the host_vars migration:
infra/nodito/nodito_vars.yml was byte-identical to host_vars/nodito/main.yml.
Deleted it, moved nodito_secrets.yml to host_vars/nodito/vault.yml, and
stripped the dead vars_files entries from all three playbooks.
Extract the 13 inline `content: |` blocks to infra/nodito/templates/, pulled via
a YAML load rather than retyped. 681+582 lines become 313+344 plus templates.
Ownership parity against HEAD checked mechanically: no owner/group/mode drift on
any surviving task.
Neither playbook has been applied. ZFS play 2 check-runs failed=0 with two
changes: the script rewrite (the deployed one is a hand-edited DEBUG VERSION
with `set -x`) and the Environment line.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 19:07:50 +02:00
|
|
|
# The script decides on-mains/on-battery and says so in its exit code, which
|
|
|
|
|
# systemd keeps: systemctl is-failed ups-heartbeat.service
|
2026-09-11 22:43:55 +02:00
|
|
|
#
|
nodito: de-Uptime-Kuma the ZFS and NUT playbooks, extract templates
These two are host-specific by design - nodito is a pet, not cattle - so they
stay playbooks rather than becoming roles. But they had rotted.
De-Kuma, following the pattern of the six service roles:
Both plays opened with an assert on uptime_kuma_username/password, which were
removed from the vault, so both failed before doing anything. Dropped that,
the two embedded Python monitor-creation scripts, and their /tmp cleanup.
Kept every check, threshold and systemd timer - those are the durable part.
Reporting is now generic: `healthcheck_push_url` goes into the unit as
Environment=HEALTHCHECK_PUSH_URL and the script reads ${HEALTHCHECK_PUSH_URL:-},
treating empty as normal rather than an error. The exit code is the real
answer; systemd keeps it. Both scripts now also report status=down on failure
instead of only going silent. Live push URLs harvested into the vault so
nothing observable changes for ZFS.
Three live bugs found while check-diffing:
1. 32_zfs would have DE-REGISTERED the Proxmox storage. `pvesm remove` was
gated on the storage existing and `pvesm add` on it NOT existing - mutually
exclusive - so a real run removed the proxmox-tank-1 entry backing every VM
and never put it back. It would also have dropped `mountpoint /var/lib/vz`,
which the live entry has and `pvesm add` does not set. Registration is now
add-only.
2. zfs_disk_1 named ata-...WX11TN0Z, a disk no longer in the machine. The live
mirror is WX120LHQ + WX11TN2P; a leg was replaced and the repo never caught
up. Inert behind `when: zfs_pool_exists.rc != 0`, but wrong on any
disaster-recovery run. This is the seventh instance of an identifier
written down once whose hardware later moved.
3. 34_nut has NEVER been applied to nodito - no /etc/nut file carries the
"Managed by Ansible" marker; they were written by hand in January 2026. The
vault held the literal CHANGE_ME_TO_SECURE_PASSWORD, so applying it would
have overwritten a working upsd/upsmon auth pair with a placeholder and
restarted NUT, leaving the hypervisor's UPS unable to trigger a clean
shutdown on mains loss. The Kuma assert was the only thing stopping that,
so removing it without a replacement would have armed the gun: there is now
an explicit assert that refuses to run on the placeholder. The real
password is in the vault and `Configure upsd users` check-diffs clean.
Templates reconciled with the live files first, so applying 34_nut is close to
a no-op: added `maxretry = 3` to ups.conf and OFFDURATION / RBWARNTIME /
NOCOMMWARNTIME / FINALDELAY plus quoted POWERDOWNFLAG to upsmon.conf. The only
substantive additions left are the NOTIFYMSG/NOTIFYFLAG syslog lines.
/usr/local/bin/ups-heartbeat.sh on the box is an orphan - mode 0644, not
executable, referenced by no unit and no cron entry - but its push token
belongs to a monitor that still exists and answers, so that monitor has had no
heartbeat since January. Harvested as healthcheck_push_urls.ups; applying this
play is what will finally feed it.
Finish the host_vars migration:
infra/nodito/nodito_vars.yml was byte-identical to host_vars/nodito/main.yml.
Deleted it, moved nodito_secrets.yml to host_vars/nodito/vault.yml, and
stripped the dead vars_files entries from all three playbooks.
Extract the 13 inline `content: |` blocks to infra/nodito/templates/, pulled via
a YAML load rather than retyped. 681+582 lines become 313+344 plus templates.
Ownership parity against HEAD checked mechanically: no owner/group/mode drift on
any surviving task.
Neither playbook has been applied. ZFS play 2 check-runs failed=0 with two
changes: the script rewrite (the deployed one is a hand-edited DEBUG VERSION
with `set -x`) and the Environment line.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 19:07:50 +02:00
|
|
|
# Reporting anywhere else is optional. Set `healthcheck_push_url` and the script
|
|
|
|
|
# will also GET it with ?status=up|down; leave it empty and the exit code is
|
|
|
|
|
# still the whole answer. Today that URL points at Uptime Kuma, which is still
|
|
|
|
|
# running on watchtower but is no longer deployed by Ansible - the credentials
|
|
|
|
|
# were retired, the service was not. If it is ever replaced, `healthcheck_push_url`
|
|
|
|
|
# is the only thing that needs to change here.
|
2026-09-11 22:43:55 +02:00
|
|
|
#
|
nodito: de-Uptime-Kuma the ZFS and NUT playbooks, extract templates
These two are host-specific by design - nodito is a pet, not cattle - so they
stay playbooks rather than becoming roles. But they had rotted.
De-Kuma, following the pattern of the six service roles:
Both plays opened with an assert on uptime_kuma_username/password, which were
removed from the vault, so both failed before doing anything. Dropped that,
the two embedded Python monitor-creation scripts, and their /tmp cleanup.
Kept every check, threshold and systemd timer - those are the durable part.
Reporting is now generic: `healthcheck_push_url` goes into the unit as
Environment=HEALTHCHECK_PUSH_URL and the script reads ${HEALTHCHECK_PUSH_URL:-},
treating empty as normal rather than an error. The exit code is the real
answer; systemd keeps it. Both scripts now also report status=down on failure
instead of only going silent. Live push URLs harvested into the vault so
nothing observable changes for ZFS.
Three live bugs found while check-diffing:
1. 32_zfs would have DE-REGISTERED the Proxmox storage. `pvesm remove` was
gated on the storage existing and `pvesm add` on it NOT existing - mutually
exclusive - so a real run removed the proxmox-tank-1 entry backing every VM
and never put it back. It would also have dropped `mountpoint /var/lib/vz`,
which the live entry has and `pvesm add` does not set. Registration is now
add-only.
2. zfs_disk_1 named ata-...WX11TN0Z, a disk no longer in the machine. The live
mirror is WX120LHQ + WX11TN2P; a leg was replaced and the repo never caught
up. Inert behind `when: zfs_pool_exists.rc != 0`, but wrong on any
disaster-recovery run. This is the seventh instance of an identifier
written down once whose hardware later moved.
3. 34_nut has NEVER been applied to nodito - no /etc/nut file carries the
"Managed by Ansible" marker; they were written by hand in January 2026. The
vault held the literal CHANGE_ME_TO_SECURE_PASSWORD, so applying it would
have overwritten a working upsd/upsmon auth pair with a placeholder and
restarted NUT, leaving the hypervisor's UPS unable to trigger a clean
shutdown on mains loss. The Kuma assert was the only thing stopping that,
so removing it without a replacement would have armed the gun: there is now
an explicit assert that refuses to run on the placeholder. The real
password is in the vault and `Configure upsd users` check-diffs clean.
Templates reconciled with the live files first, so applying 34_nut is close to
a no-op: added `maxretry = 3` to ups.conf and OFFDURATION / RBWARNTIME /
NOCOMMWARNTIME / FINALDELAY plus quoted POWERDOWNFLAG to upsmon.conf. The only
substantive additions left are the NOTIFYMSG/NOTIFYFLAG syslog lines.
/usr/local/bin/ups-heartbeat.sh on the box is an orphan - mode 0644, not
executable, referenced by no unit and no cron entry - but its push token
belongs to a monitor that still exists and answers, so that monitor has had no
heartbeat since January. Harvested as healthcheck_push_urls.ups; applying this
play is what will finally feed it.
Finish the host_vars migration:
infra/nodito/nodito_vars.yml was byte-identical to host_vars/nodito/main.yml.
Deleted it, moved nodito_secrets.yml to host_vars/nodito/vault.yml, and
stripped the dead vars_files entries from all three playbooks.
Extract the 13 inline `content: |` blocks to infra/nodito/templates/, pulled via
a YAML load rather than retyped. 681+582 lines become 313+344 plus templates.
Ownership parity against HEAD checked mechanically: no owner/group/mode drift on
any surviving task.
Neither playbook has been applied. ZFS play 2 check-runs failed=0 with two
changes: the script rewrite (the deployed one is a hand-edited DEBUG VERSION
with `set -x`) and the Environment line.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 19:07:50 +02:00
|
|
|
# NOTE: this play has never actually been applied to nodito. /opt/ups-monitoring
|
|
|
|
|
# does not exist and there is no ups-heartbeat timer. What is on the box is a
|
|
|
|
|
# hand-written /usr/local/bin/ups-heartbeat.sh - mode 0644, not executable, and
|
|
|
|
|
# referenced by no unit and no cron entry, so nothing has ever run it. Its push
|
|
|
|
|
# token (uLmCPkLLO4) belongs to a monitor that DOES still exist and answer, so
|
|
|
|
|
# that monitor has had no heartbeat since January 2026. It is now
|
|
|
|
|
# healthcheck_push_urls.ups in the vault, and running this play is what will
|
|
|
|
|
# finally start feeding it.
|
|
|
|
|
# ─────────────────────────────────────────────────────────────────────────────
|
|
|
|
|
- name: Setup UPS Heartbeat Monitoring
|
2026-09-11 21:56:24 +02:00
|
|
|
hosts: hypervisor
|
2026-01-11 22:43:27 +01:00
|
|
|
become: true
|
|
|
|
|
|
|
|
|
|
vars:
|
|
|
|
|
ups_heartbeat_interval_seconds: 60
|
|
|
|
|
ups_heartbeat_timeout_seconds: 120
|
|
|
|
|
ups_heartbeat_retries: 1
|
|
|
|
|
ups_monitoring_script_dir: /opt/ups-monitoring
|
|
|
|
|
ups_monitoring_script_path: "{{ ups_monitoring_script_dir }}/ups_heartbeat.sh"
|
|
|
|
|
ups_log_file: "{{ ups_monitoring_script_dir }}/ups_heartbeat.log"
|
|
|
|
|
ups_systemd_service_name: ups-heartbeat
|
nodito: de-Uptime-Kuma the ZFS and NUT playbooks, extract templates
These two are host-specific by design - nodito is a pet, not cattle - so they
stay playbooks rather than becoming roles. But they had rotted.
De-Kuma, following the pattern of the six service roles:
Both plays opened with an assert on uptime_kuma_username/password, which were
removed from the vault, so both failed before doing anything. Dropped that,
the two embedded Python monitor-creation scripts, and their /tmp cleanup.
Kept every check, threshold and systemd timer - those are the durable part.
Reporting is now generic: `healthcheck_push_url` goes into the unit as
Environment=HEALTHCHECK_PUSH_URL and the script reads ${HEALTHCHECK_PUSH_URL:-},
treating empty as normal rather than an error. The exit code is the real
answer; systemd keeps it. Both scripts now also report status=down on failure
instead of only going silent. Live push URLs harvested into the vault so
nothing observable changes for ZFS.
Three live bugs found while check-diffing:
1. 32_zfs would have DE-REGISTERED the Proxmox storage. `pvesm remove` was
gated on the storage existing and `pvesm add` on it NOT existing - mutually
exclusive - so a real run removed the proxmox-tank-1 entry backing every VM
and never put it back. It would also have dropped `mountpoint /var/lib/vz`,
which the live entry has and `pvesm add` does not set. Registration is now
add-only.
2. zfs_disk_1 named ata-...WX11TN0Z, a disk no longer in the machine. The live
mirror is WX120LHQ + WX11TN2P; a leg was replaced and the repo never caught
up. Inert behind `when: zfs_pool_exists.rc != 0`, but wrong on any
disaster-recovery run. This is the seventh instance of an identifier
written down once whose hardware later moved.
3. 34_nut has NEVER been applied to nodito - no /etc/nut file carries the
"Managed by Ansible" marker; they were written by hand in January 2026. The
vault held the literal CHANGE_ME_TO_SECURE_PASSWORD, so applying it would
have overwritten a working upsd/upsmon auth pair with a placeholder and
restarted NUT, leaving the hypervisor's UPS unable to trigger a clean
shutdown on mains loss. The Kuma assert was the only thing stopping that,
so removing it without a replacement would have armed the gun: there is now
an explicit assert that refuses to run on the placeholder. The real
password is in the vault and `Configure upsd users` check-diffs clean.
Templates reconciled with the live files first, so applying 34_nut is close to
a no-op: added `maxretry = 3` to ups.conf and OFFDURATION / RBWARNTIME /
NOCOMMWARNTIME / FINALDELAY plus quoted POWERDOWNFLAG to upsmon.conf. The only
substantive additions left are the NOTIFYMSG/NOTIFYFLAG syslog lines.
/usr/local/bin/ups-heartbeat.sh on the box is an orphan - mode 0644, not
executable, referenced by no unit and no cron entry - but its push token
belongs to a monitor that still exists and answers, so that monitor has had no
heartbeat since January. Harvested as healthcheck_push_urls.ups; applying this
play is what will finally feed it.
Finish the host_vars migration:
infra/nodito/nodito_vars.yml was byte-identical to host_vars/nodito/main.yml.
Deleted it, moved nodito_secrets.yml to host_vars/nodito/vault.yml, and
stripped the dead vars_files entries from all three playbooks.
Extract the 13 inline `content: |` blocks to infra/nodito/templates/, pulled via
a YAML load rather than retyped. 681+582 lines become 313+344 plus templates.
Ownership parity against HEAD checked mechanically: no owner/group/mode drift on
any surviving task.
Neither playbook has been applied. ZFS play 2 check-runs failed=0 with two
changes: the script rewrite (the deployed one is a hand-edited DEBUG VERSION
with `set -x`) and the Environment line.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 19:07:50 +02:00
|
|
|
# Optional. Empty is fine and is not an error - see the banner above.
|
|
|
|
|
healthcheck_push_url: "{{ healthcheck_push_urls.ups | default('') }}"
|
2026-01-11 22:43:27 +01:00
|
|
|
|
|
|
|
|
tasks:
|
|
|
|
|
- name: Install required packages for UPS monitoring
|
|
|
|
|
package:
|
|
|
|
|
name:
|
|
|
|
|
- curl
|
|
|
|
|
state: present
|
|
|
|
|
|
|
|
|
|
- name: Create monitoring script directory
|
|
|
|
|
file:
|
|
|
|
|
path: "{{ ups_monitoring_script_dir }}"
|
|
|
|
|
state: directory
|
|
|
|
|
owner: root
|
|
|
|
|
group: root
|
|
|
|
|
mode: '0755'
|
|
|
|
|
|
|
|
|
|
- name: Create UPS heartbeat monitoring script
|
nodito: de-Uptime-Kuma the ZFS and NUT playbooks, extract templates
These two are host-specific by design - nodito is a pet, not cattle - so they
stay playbooks rather than becoming roles. But they had rotted.
De-Kuma, following the pattern of the six service roles:
Both plays opened with an assert on uptime_kuma_username/password, which were
removed from the vault, so both failed before doing anything. Dropped that,
the two embedded Python monitor-creation scripts, and their /tmp cleanup.
Kept every check, threshold and systemd timer - those are the durable part.
Reporting is now generic: `healthcheck_push_url` goes into the unit as
Environment=HEALTHCHECK_PUSH_URL and the script reads ${HEALTHCHECK_PUSH_URL:-},
treating empty as normal rather than an error. The exit code is the real
answer; systemd keeps it. Both scripts now also report status=down on failure
instead of only going silent. Live push URLs harvested into the vault so
nothing observable changes for ZFS.
Three live bugs found while check-diffing:
1. 32_zfs would have DE-REGISTERED the Proxmox storage. `pvesm remove` was
gated on the storage existing and `pvesm add` on it NOT existing - mutually
exclusive - so a real run removed the proxmox-tank-1 entry backing every VM
and never put it back. It would also have dropped `mountpoint /var/lib/vz`,
which the live entry has and `pvesm add` does not set. Registration is now
add-only.
2. zfs_disk_1 named ata-...WX11TN0Z, a disk no longer in the machine. The live
mirror is WX120LHQ + WX11TN2P; a leg was replaced and the repo never caught
up. Inert behind `when: zfs_pool_exists.rc != 0`, but wrong on any
disaster-recovery run. This is the seventh instance of an identifier
written down once whose hardware later moved.
3. 34_nut has NEVER been applied to nodito - no /etc/nut file carries the
"Managed by Ansible" marker; they were written by hand in January 2026. The
vault held the literal CHANGE_ME_TO_SECURE_PASSWORD, so applying it would
have overwritten a working upsd/upsmon auth pair with a placeholder and
restarted NUT, leaving the hypervisor's UPS unable to trigger a clean
shutdown on mains loss. The Kuma assert was the only thing stopping that,
so removing it without a replacement would have armed the gun: there is now
an explicit assert that refuses to run on the placeholder. The real
password is in the vault and `Configure upsd users` check-diffs clean.
Templates reconciled with the live files first, so applying 34_nut is close to
a no-op: added `maxretry = 3` to ups.conf and OFFDURATION / RBWARNTIME /
NOCOMMWARNTIME / FINALDELAY plus quoted POWERDOWNFLAG to upsmon.conf. The only
substantive additions left are the NOTIFYMSG/NOTIFYFLAG syslog lines.
/usr/local/bin/ups-heartbeat.sh on the box is an orphan - mode 0644, not
executable, referenced by no unit and no cron entry - but its push token
belongs to a monitor that still exists and answers, so that monitor has had no
heartbeat since January. Harvested as healthcheck_push_urls.ups; applying this
play is what will finally feed it.
Finish the host_vars migration:
infra/nodito/nodito_vars.yml was byte-identical to host_vars/nodito/main.yml.
Deleted it, moved nodito_secrets.yml to host_vars/nodito/vault.yml, and
stripped the dead vars_files entries from all three playbooks.
Extract the 13 inline `content: |` blocks to infra/nodito/templates/, pulled via
a YAML load rather than retyped. 681+582 lines become 313+344 plus templates.
Ownership parity against HEAD checked mechanically: no owner/group/mode drift on
any surviving task.
Neither playbook has been applied. ZFS play 2 check-runs failed=0 with two
changes: the script rewrite (the deployed one is a hand-edited DEBUG VERSION
with `set -x`) and the Environment line.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 19:07:50 +02:00
|
|
|
template:
|
2026-01-11 22:43:27 +01:00
|
|
|
dest: "{{ ups_monitoring_script_path }}"
|
nodito: de-Uptime-Kuma the ZFS and NUT playbooks, extract templates
These two are host-specific by design - nodito is a pet, not cattle - so they
stay playbooks rather than becoming roles. But they had rotted.
De-Kuma, following the pattern of the six service roles:
Both plays opened with an assert on uptime_kuma_username/password, which were
removed from the vault, so both failed before doing anything. Dropped that,
the two embedded Python monitor-creation scripts, and their /tmp cleanup.
Kept every check, threshold and systemd timer - those are the durable part.
Reporting is now generic: `healthcheck_push_url` goes into the unit as
Environment=HEALTHCHECK_PUSH_URL and the script reads ${HEALTHCHECK_PUSH_URL:-},
treating empty as normal rather than an error. The exit code is the real
answer; systemd keeps it. Both scripts now also report status=down on failure
instead of only going silent. Live push URLs harvested into the vault so
nothing observable changes for ZFS.
Three live bugs found while check-diffing:
1. 32_zfs would have DE-REGISTERED the Proxmox storage. `pvesm remove` was
gated on the storage existing and `pvesm add` on it NOT existing - mutually
exclusive - so a real run removed the proxmox-tank-1 entry backing every VM
and never put it back. It would also have dropped `mountpoint /var/lib/vz`,
which the live entry has and `pvesm add` does not set. Registration is now
add-only.
2. zfs_disk_1 named ata-...WX11TN0Z, a disk no longer in the machine. The live
mirror is WX120LHQ + WX11TN2P; a leg was replaced and the repo never caught
up. Inert behind `when: zfs_pool_exists.rc != 0`, but wrong on any
disaster-recovery run. This is the seventh instance of an identifier
written down once whose hardware later moved.
3. 34_nut has NEVER been applied to nodito - no /etc/nut file carries the
"Managed by Ansible" marker; they were written by hand in January 2026. The
vault held the literal CHANGE_ME_TO_SECURE_PASSWORD, so applying it would
have overwritten a working upsd/upsmon auth pair with a placeholder and
restarted NUT, leaving the hypervisor's UPS unable to trigger a clean
shutdown on mains loss. The Kuma assert was the only thing stopping that,
so removing it without a replacement would have armed the gun: there is now
an explicit assert that refuses to run on the placeholder. The real
password is in the vault and `Configure upsd users` check-diffs clean.
Templates reconciled with the live files first, so applying 34_nut is close to
a no-op: added `maxretry = 3` to ups.conf and OFFDURATION / RBWARNTIME /
NOCOMMWARNTIME / FINALDELAY plus quoted POWERDOWNFLAG to upsmon.conf. The only
substantive additions left are the NOTIFYMSG/NOTIFYFLAG syslog lines.
/usr/local/bin/ups-heartbeat.sh on the box is an orphan - mode 0644, not
executable, referenced by no unit and no cron entry - but its push token
belongs to a monitor that still exists and answers, so that monitor has had no
heartbeat since January. Harvested as healthcheck_push_urls.ups; applying this
play is what will finally feed it.
Finish the host_vars migration:
infra/nodito/nodito_vars.yml was byte-identical to host_vars/nodito/main.yml.
Deleted it, moved nodito_secrets.yml to host_vars/nodito/vault.yml, and
stripped the dead vars_files entries from all three playbooks.
Extract the 13 inline `content: |` blocks to infra/nodito/templates/, pulled via
a YAML load rather than retyped. 681+582 lines become 313+344 plus templates.
Ownership parity against HEAD checked mechanically: no owner/group/mode drift on
any surviving task.
Neither playbook has been applied. ZFS play 2 check-runs failed=0 with two
changes: the script rewrite (the deployed one is a hand-edited DEBUG VERSION
with `set -x`) and the Environment line.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 19:07:50 +02:00
|
|
|
src: templates/ups_heartbeat.sh.j2
|
2026-01-11 22:43:27 +01:00
|
|
|
owner: root
|
|
|
|
|
group: root
|
|
|
|
|
mode: '0755'
|
|
|
|
|
|
|
|
|
|
- name: Create systemd service for UPS heartbeat
|
nodito: de-Uptime-Kuma the ZFS and NUT playbooks, extract templates
These two are host-specific by design - nodito is a pet, not cattle - so they
stay playbooks rather than becoming roles. But they had rotted.
De-Kuma, following the pattern of the six service roles:
Both plays opened with an assert on uptime_kuma_username/password, which were
removed from the vault, so both failed before doing anything. Dropped that,
the two embedded Python monitor-creation scripts, and their /tmp cleanup.
Kept every check, threshold and systemd timer - those are the durable part.
Reporting is now generic: `healthcheck_push_url` goes into the unit as
Environment=HEALTHCHECK_PUSH_URL and the script reads ${HEALTHCHECK_PUSH_URL:-},
treating empty as normal rather than an error. The exit code is the real
answer; systemd keeps it. Both scripts now also report status=down on failure
instead of only going silent. Live push URLs harvested into the vault so
nothing observable changes for ZFS.
Three live bugs found while check-diffing:
1. 32_zfs would have DE-REGISTERED the Proxmox storage. `pvesm remove` was
gated on the storage existing and `pvesm add` on it NOT existing - mutually
exclusive - so a real run removed the proxmox-tank-1 entry backing every VM
and never put it back. It would also have dropped `mountpoint /var/lib/vz`,
which the live entry has and `pvesm add` does not set. Registration is now
add-only.
2. zfs_disk_1 named ata-...WX11TN0Z, a disk no longer in the machine. The live
mirror is WX120LHQ + WX11TN2P; a leg was replaced and the repo never caught
up. Inert behind `when: zfs_pool_exists.rc != 0`, but wrong on any
disaster-recovery run. This is the seventh instance of an identifier
written down once whose hardware later moved.
3. 34_nut has NEVER been applied to nodito - no /etc/nut file carries the
"Managed by Ansible" marker; they were written by hand in January 2026. The
vault held the literal CHANGE_ME_TO_SECURE_PASSWORD, so applying it would
have overwritten a working upsd/upsmon auth pair with a placeholder and
restarted NUT, leaving the hypervisor's UPS unable to trigger a clean
shutdown on mains loss. The Kuma assert was the only thing stopping that,
so removing it without a replacement would have armed the gun: there is now
an explicit assert that refuses to run on the placeholder. The real
password is in the vault and `Configure upsd users` check-diffs clean.
Templates reconciled with the live files first, so applying 34_nut is close to
a no-op: added `maxretry = 3` to ups.conf and OFFDURATION / RBWARNTIME /
NOCOMMWARNTIME / FINALDELAY plus quoted POWERDOWNFLAG to upsmon.conf. The only
substantive additions left are the NOTIFYMSG/NOTIFYFLAG syslog lines.
/usr/local/bin/ups-heartbeat.sh on the box is an orphan - mode 0644, not
executable, referenced by no unit and no cron entry - but its push token
belongs to a monitor that still exists and answers, so that monitor has had no
heartbeat since January. Harvested as healthcheck_push_urls.ups; applying this
play is what will finally feed it.
Finish the host_vars migration:
infra/nodito/nodito_vars.yml was byte-identical to host_vars/nodito/main.yml.
Deleted it, moved nodito_secrets.yml to host_vars/nodito/vault.yml, and
stripped the dead vars_files entries from all three playbooks.
Extract the 13 inline `content: |` blocks to infra/nodito/templates/, pulled via
a YAML load rather than retyped. 681+582 lines become 313+344 plus templates.
Ownership parity against HEAD checked mechanically: no owner/group/mode drift on
any surviving task.
Neither playbook has been applied. ZFS play 2 check-runs failed=0 with two
changes: the script rewrite (the deployed one is a hand-edited DEBUG VERSION
with `set -x`) and the Environment line.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 19:07:50 +02:00
|
|
|
template:
|
2026-01-11 22:43:27 +01:00
|
|
|
dest: "/etc/systemd/system/{{ ups_systemd_service_name }}.service"
|
nodito: de-Uptime-Kuma the ZFS and NUT playbooks, extract templates
These two are host-specific by design - nodito is a pet, not cattle - so they
stay playbooks rather than becoming roles. But they had rotted.
De-Kuma, following the pattern of the six service roles:
Both plays opened with an assert on uptime_kuma_username/password, which were
removed from the vault, so both failed before doing anything. Dropped that,
the two embedded Python monitor-creation scripts, and their /tmp cleanup.
Kept every check, threshold and systemd timer - those are the durable part.
Reporting is now generic: `healthcheck_push_url` goes into the unit as
Environment=HEALTHCHECK_PUSH_URL and the script reads ${HEALTHCHECK_PUSH_URL:-},
treating empty as normal rather than an error. The exit code is the real
answer; systemd keeps it. Both scripts now also report status=down on failure
instead of only going silent. Live push URLs harvested into the vault so
nothing observable changes for ZFS.
Three live bugs found while check-diffing:
1. 32_zfs would have DE-REGISTERED the Proxmox storage. `pvesm remove` was
gated on the storage existing and `pvesm add` on it NOT existing - mutually
exclusive - so a real run removed the proxmox-tank-1 entry backing every VM
and never put it back. It would also have dropped `mountpoint /var/lib/vz`,
which the live entry has and `pvesm add` does not set. Registration is now
add-only.
2. zfs_disk_1 named ata-...WX11TN0Z, a disk no longer in the machine. The live
mirror is WX120LHQ + WX11TN2P; a leg was replaced and the repo never caught
up. Inert behind `when: zfs_pool_exists.rc != 0`, but wrong on any
disaster-recovery run. This is the seventh instance of an identifier
written down once whose hardware later moved.
3. 34_nut has NEVER been applied to nodito - no /etc/nut file carries the
"Managed by Ansible" marker; they were written by hand in January 2026. The
vault held the literal CHANGE_ME_TO_SECURE_PASSWORD, so applying it would
have overwritten a working upsd/upsmon auth pair with a placeholder and
restarted NUT, leaving the hypervisor's UPS unable to trigger a clean
shutdown on mains loss. The Kuma assert was the only thing stopping that,
so removing it without a replacement would have armed the gun: there is now
an explicit assert that refuses to run on the placeholder. The real
password is in the vault and `Configure upsd users` check-diffs clean.
Templates reconciled with the live files first, so applying 34_nut is close to
a no-op: added `maxretry = 3` to ups.conf and OFFDURATION / RBWARNTIME /
NOCOMMWARNTIME / FINALDELAY plus quoted POWERDOWNFLAG to upsmon.conf. The only
substantive additions left are the NOTIFYMSG/NOTIFYFLAG syslog lines.
/usr/local/bin/ups-heartbeat.sh on the box is an orphan - mode 0644, not
executable, referenced by no unit and no cron entry - but its push token
belongs to a monitor that still exists and answers, so that monitor has had no
heartbeat since January. Harvested as healthcheck_push_urls.ups; applying this
play is what will finally feed it.
Finish the host_vars migration:
infra/nodito/nodito_vars.yml was byte-identical to host_vars/nodito/main.yml.
Deleted it, moved nodito_secrets.yml to host_vars/nodito/vault.yml, and
stripped the dead vars_files entries from all three playbooks.
Extract the 13 inline `content: |` blocks to infra/nodito/templates/, pulled via
a YAML load rather than retyped. 681+582 lines become 313+344 plus templates.
Ownership parity against HEAD checked mechanically: no owner/group/mode drift on
any surviving task.
Neither playbook has been applied. ZFS play 2 check-runs failed=0 with two
changes: the script rewrite (the deployed one is a hand-edited DEBUG VERSION
with `set -x`) and the Environment line.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 19:07:50 +02:00
|
|
|
src: templates/ups-heartbeat.service.j2
|
2026-01-11 22:43:27 +01:00
|
|
|
owner: root
|
|
|
|
|
group: root
|
|
|
|
|
mode: '0644'
|
|
|
|
|
|
|
|
|
|
- name: Create systemd timer for UPS heartbeat
|
nodito: de-Uptime-Kuma the ZFS and NUT playbooks, extract templates
These two are host-specific by design - nodito is a pet, not cattle - so they
stay playbooks rather than becoming roles. But they had rotted.
De-Kuma, following the pattern of the six service roles:
Both plays opened with an assert on uptime_kuma_username/password, which were
removed from the vault, so both failed before doing anything. Dropped that,
the two embedded Python monitor-creation scripts, and their /tmp cleanup.
Kept every check, threshold and systemd timer - those are the durable part.
Reporting is now generic: `healthcheck_push_url` goes into the unit as
Environment=HEALTHCHECK_PUSH_URL and the script reads ${HEALTHCHECK_PUSH_URL:-},
treating empty as normal rather than an error. The exit code is the real
answer; systemd keeps it. Both scripts now also report status=down on failure
instead of only going silent. Live push URLs harvested into the vault so
nothing observable changes for ZFS.
Three live bugs found while check-diffing:
1. 32_zfs would have DE-REGISTERED the Proxmox storage. `pvesm remove` was
gated on the storage existing and `pvesm add` on it NOT existing - mutually
exclusive - so a real run removed the proxmox-tank-1 entry backing every VM
and never put it back. It would also have dropped `mountpoint /var/lib/vz`,
which the live entry has and `pvesm add` does not set. Registration is now
add-only.
2. zfs_disk_1 named ata-...WX11TN0Z, a disk no longer in the machine. The live
mirror is WX120LHQ + WX11TN2P; a leg was replaced and the repo never caught
up. Inert behind `when: zfs_pool_exists.rc != 0`, but wrong on any
disaster-recovery run. This is the seventh instance of an identifier
written down once whose hardware later moved.
3. 34_nut has NEVER been applied to nodito - no /etc/nut file carries the
"Managed by Ansible" marker; they were written by hand in January 2026. The
vault held the literal CHANGE_ME_TO_SECURE_PASSWORD, so applying it would
have overwritten a working upsd/upsmon auth pair with a placeholder and
restarted NUT, leaving the hypervisor's UPS unable to trigger a clean
shutdown on mains loss. The Kuma assert was the only thing stopping that,
so removing it without a replacement would have armed the gun: there is now
an explicit assert that refuses to run on the placeholder. The real
password is in the vault and `Configure upsd users` check-diffs clean.
Templates reconciled with the live files first, so applying 34_nut is close to
a no-op: added `maxretry = 3` to ups.conf and OFFDURATION / RBWARNTIME /
NOCOMMWARNTIME / FINALDELAY plus quoted POWERDOWNFLAG to upsmon.conf. The only
substantive additions left are the NOTIFYMSG/NOTIFYFLAG syslog lines.
/usr/local/bin/ups-heartbeat.sh on the box is an orphan - mode 0644, not
executable, referenced by no unit and no cron entry - but its push token
belongs to a monitor that still exists and answers, so that monitor has had no
heartbeat since January. Harvested as healthcheck_push_urls.ups; applying this
play is what will finally feed it.
Finish the host_vars migration:
infra/nodito/nodito_vars.yml was byte-identical to host_vars/nodito/main.yml.
Deleted it, moved nodito_secrets.yml to host_vars/nodito/vault.yml, and
stripped the dead vars_files entries from all three playbooks.
Extract the 13 inline `content: |` blocks to infra/nodito/templates/, pulled via
a YAML load rather than retyped. 681+582 lines become 313+344 plus templates.
Ownership parity against HEAD checked mechanically: no owner/group/mode drift on
any surviving task.
Neither playbook has been applied. ZFS play 2 check-runs failed=0 with two
changes: the script rewrite (the deployed one is a hand-edited DEBUG VERSION
with `set -x`) and the Environment line.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 19:07:50 +02:00
|
|
|
template:
|
2026-01-11 22:43:27 +01:00
|
|
|
dest: "/etc/systemd/system/{{ ups_systemd_service_name }}.timer"
|
nodito: de-Uptime-Kuma the ZFS and NUT playbooks, extract templates
These two are host-specific by design - nodito is a pet, not cattle - so they
stay playbooks rather than becoming roles. But they had rotted.
De-Kuma, following the pattern of the six service roles:
Both plays opened with an assert on uptime_kuma_username/password, which were
removed from the vault, so both failed before doing anything. Dropped that,
the two embedded Python monitor-creation scripts, and their /tmp cleanup.
Kept every check, threshold and systemd timer - those are the durable part.
Reporting is now generic: `healthcheck_push_url` goes into the unit as
Environment=HEALTHCHECK_PUSH_URL and the script reads ${HEALTHCHECK_PUSH_URL:-},
treating empty as normal rather than an error. The exit code is the real
answer; systemd keeps it. Both scripts now also report status=down on failure
instead of only going silent. Live push URLs harvested into the vault so
nothing observable changes for ZFS.
Three live bugs found while check-diffing:
1. 32_zfs would have DE-REGISTERED the Proxmox storage. `pvesm remove` was
gated on the storage existing and `pvesm add` on it NOT existing - mutually
exclusive - so a real run removed the proxmox-tank-1 entry backing every VM
and never put it back. It would also have dropped `mountpoint /var/lib/vz`,
which the live entry has and `pvesm add` does not set. Registration is now
add-only.
2. zfs_disk_1 named ata-...WX11TN0Z, a disk no longer in the machine. The live
mirror is WX120LHQ + WX11TN2P; a leg was replaced and the repo never caught
up. Inert behind `when: zfs_pool_exists.rc != 0`, but wrong on any
disaster-recovery run. This is the seventh instance of an identifier
written down once whose hardware later moved.
3. 34_nut has NEVER been applied to nodito - no /etc/nut file carries the
"Managed by Ansible" marker; they were written by hand in January 2026. The
vault held the literal CHANGE_ME_TO_SECURE_PASSWORD, so applying it would
have overwritten a working upsd/upsmon auth pair with a placeholder and
restarted NUT, leaving the hypervisor's UPS unable to trigger a clean
shutdown on mains loss. The Kuma assert was the only thing stopping that,
so removing it without a replacement would have armed the gun: there is now
an explicit assert that refuses to run on the placeholder. The real
password is in the vault and `Configure upsd users` check-diffs clean.
Templates reconciled with the live files first, so applying 34_nut is close to
a no-op: added `maxretry = 3` to ups.conf and OFFDURATION / RBWARNTIME /
NOCOMMWARNTIME / FINALDELAY plus quoted POWERDOWNFLAG to upsmon.conf. The only
substantive additions left are the NOTIFYMSG/NOTIFYFLAG syslog lines.
/usr/local/bin/ups-heartbeat.sh on the box is an orphan - mode 0644, not
executable, referenced by no unit and no cron entry - but its push token
belongs to a monitor that still exists and answers, so that monitor has had no
heartbeat since January. Harvested as healthcheck_push_urls.ups; applying this
play is what will finally feed it.
Finish the host_vars migration:
infra/nodito/nodito_vars.yml was byte-identical to host_vars/nodito/main.yml.
Deleted it, moved nodito_secrets.yml to host_vars/nodito/vault.yml, and
stripped the dead vars_files entries from all three playbooks.
Extract the 13 inline `content: |` blocks to infra/nodito/templates/, pulled via
a YAML load rather than retyped. 681+582 lines become 313+344 plus templates.
Ownership parity against HEAD checked mechanically: no owner/group/mode drift on
any surviving task.
Neither playbook has been applied. ZFS play 2 check-runs failed=0 with two
changes: the script rewrite (the deployed one is a hand-edited DEBUG VERSION
with `set -x`) and the Environment line.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 19:07:50 +02:00
|
|
|
src: templates/ups-heartbeat.timer.j2
|
2026-01-11 22:43:27 +01:00
|
|
|
owner: root
|
|
|
|
|
group: root
|
|
|
|
|
mode: '0644'
|
|
|
|
|
|
|
|
|
|
- name: Reload systemd daemon
|
|
|
|
|
systemd:
|
|
|
|
|
daemon_reload: yes
|
|
|
|
|
|
|
|
|
|
- name: Enable and start UPS heartbeat timer
|
|
|
|
|
systemd:
|
|
|
|
|
name: "{{ ups_systemd_service_name }}.timer"
|
|
|
|
|
enabled: yes
|
|
|
|
|
state: started
|
|
|
|
|
|
|
|
|
|
- name: Test UPS heartbeat script
|
|
|
|
|
command: "{{ ups_monitoring_script_path }}"
|
|
|
|
|
register: script_test
|
|
|
|
|
changed_when: false
|
|
|
|
|
|
|
|
|
|
- name: Verify script execution
|
|
|
|
|
assert:
|
|
|
|
|
that:
|
|
|
|
|
- script_test.rc == 0
|
|
|
|
|
fail_msg: "UPS heartbeat script failed - check UPS status and communication"
|
|
|
|
|
|
|
|
|
|
- name: Display monitoring configuration
|
|
|
|
|
debug:
|
|
|
|
|
msg:
|
|
|
|
|
- "UPS Monitoring configured successfully"
|
|
|
|
|
- ""
|
|
|
|
|
- "NUT Configuration:"
|
|
|
|
|
- " UPS Name: {{ ups_name }}"
|
|
|
|
|
- " UPS Description: {{ ups_desc }}"
|
|
|
|
|
- " Off Delay: {{ ups_offdelay }}s (time after shutdown before UPS cuts power)"
|
|
|
|
|
- " On Delay: {{ ups_ondelay }}s (time after mains returns before UPS restores power)"
|
|
|
|
|
- ""
|
nodito: de-Uptime-Kuma the ZFS and NUT playbooks, extract templates
These two are host-specific by design - nodito is a pet, not cattle - so they
stay playbooks rather than becoming roles. But they had rotted.
De-Kuma, following the pattern of the six service roles:
Both plays opened with an assert on uptime_kuma_username/password, which were
removed from the vault, so both failed before doing anything. Dropped that,
the two embedded Python monitor-creation scripts, and their /tmp cleanup.
Kept every check, threshold and systemd timer - those are the durable part.
Reporting is now generic: `healthcheck_push_url` goes into the unit as
Environment=HEALTHCHECK_PUSH_URL and the script reads ${HEALTHCHECK_PUSH_URL:-},
treating empty as normal rather than an error. The exit code is the real
answer; systemd keeps it. Both scripts now also report status=down on failure
instead of only going silent. Live push URLs harvested into the vault so
nothing observable changes for ZFS.
Three live bugs found while check-diffing:
1. 32_zfs would have DE-REGISTERED the Proxmox storage. `pvesm remove` was
gated on the storage existing and `pvesm add` on it NOT existing - mutually
exclusive - so a real run removed the proxmox-tank-1 entry backing every VM
and never put it back. It would also have dropped `mountpoint /var/lib/vz`,
which the live entry has and `pvesm add` does not set. Registration is now
add-only.
2. zfs_disk_1 named ata-...WX11TN0Z, a disk no longer in the machine. The live
mirror is WX120LHQ + WX11TN2P; a leg was replaced and the repo never caught
up. Inert behind `when: zfs_pool_exists.rc != 0`, but wrong on any
disaster-recovery run. This is the seventh instance of an identifier
written down once whose hardware later moved.
3. 34_nut has NEVER been applied to nodito - no /etc/nut file carries the
"Managed by Ansible" marker; they were written by hand in January 2026. The
vault held the literal CHANGE_ME_TO_SECURE_PASSWORD, so applying it would
have overwritten a working upsd/upsmon auth pair with a placeholder and
restarted NUT, leaving the hypervisor's UPS unable to trigger a clean
shutdown on mains loss. The Kuma assert was the only thing stopping that,
so removing it without a replacement would have armed the gun: there is now
an explicit assert that refuses to run on the placeholder. The real
password is in the vault and `Configure upsd users` check-diffs clean.
Templates reconciled with the live files first, so applying 34_nut is close to
a no-op: added `maxretry = 3` to ups.conf and OFFDURATION / RBWARNTIME /
NOCOMMWARNTIME / FINALDELAY plus quoted POWERDOWNFLAG to upsmon.conf. The only
substantive additions left are the NOTIFYMSG/NOTIFYFLAG syslog lines.
/usr/local/bin/ups-heartbeat.sh on the box is an orphan - mode 0644, not
executable, referenced by no unit and no cron entry - but its push token
belongs to a monitor that still exists and answers, so that monitor has had no
heartbeat since January. Harvested as healthcheck_push_urls.ups; applying this
play is what will finally feed it.
Finish the host_vars migration:
infra/nodito/nodito_vars.yml was byte-identical to host_vars/nodito/main.yml.
Deleted it, moved nodito_secrets.yml to host_vars/nodito/vault.yml, and
stripped the dead vars_files entries from all three playbooks.
Extract the 13 inline `content: |` blocks to infra/nodito/templates/, pulled via
a YAML load rather than retyped. 681+582 lines become 313+344 plus templates.
Ownership parity against HEAD checked mechanically: no owner/group/mode drift on
any surviving task.
Neither playbook has been applied. ZFS play 2 check-runs failed=0 with two
changes: the script rewrite (the deployed one is a hand-edited DEBUG VERSION
with `set -x`) and the Environment line.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 19:07:50 +02:00
|
|
|
- "Health reporting:"
|
|
|
|
|
- " Check interval: {{ ups_heartbeat_interval_seconds }}s"
|
|
|
|
|
- " Push URL: {{ healthcheck_push_url | default('', true) | ternary('set', 'not set - exit code only') }}"
|
2026-01-11 22:43:27 +01:00
|
|
|
- ""
|
|
|
|
|
- "Scripts and Services:"
|
|
|
|
|
- " Script: {{ ups_monitoring_script_path }}"
|
|
|
|
|
- " Log: {{ ups_log_file }}"
|
|
|
|
|
- " Service: {{ ups_systemd_service_name }}.service"
|
|
|
|
|
- " Timer: {{ ups_systemd_service_name }}.timer"
|