personal_infra/ansible/infra/nodito/34_nut_ups_setup_playbook.yml
counterweight 954b683c71
ansible: delete the duplicated vars files, move globals to group_vars/all
Three files existed only as second copies of things group_vars/all already
auto-loads, and 34 playbooks named them in vars_files: - which outranks
group_vars, so the copies won. The day someone edited one and not the other,
those plays would silently keep the stale value. infra_vars.yml was already
drifting: group_vars/all/main.yml had grown age_backup_recipient and
backup_pull_public_key that it lacked.

  infra_vars.yml          - a strict subset of group_vars/all/main.yml
  infra_secrets.yml       - decrypts byte-identical to group_vars/all/vault.yml
  infra_secrets.yml.example - documented Uptime Kuma credentials as the reason
                              the file exists, which stopped being true

Deleted, along with 62 vars_files entries across 34 playbooks (12 of which
named ../../group_vars/all/main.yml directly - same defect, a vars_files entry
duplicating an auto-loaded file at higher precedence than the file itself).

Checked before touching anything: infra_secrets.yml was listed LAST in 10 plays,
after services_config.yml, so removal would flip precedence if the two shared a
key. They share none, and neither does services_config.yml with
group_vars/all/main.yml, so the removal is provably inert.

services_config.yml was the last one standing. It held four unrelated things:

  caddy_sites_dir            - an identical copy of roles/caddy_site/defaults/.
                               Deleted; the role default is now the only one.
  *.tailscale_hostname (x3)  - a THIRD copy of each box's identity, which
                               inventory.ini already holds as ansible_host.
                               Deleted. Edge plays now read
                               hostvars['<host>'].ansible_host - verified an
                               edge play resolves that with nothing loaded and
                               the other host in no play. Three copies of one
                               name is how bitcoin_rpc_host ended up labelled
                               "knots_box" while pointing at fulcrum-box.
  subdomains, ntfy topic,    - genuinely global: their readers span managed,
  headscale namespace          monitoring, vpn_control and edge, so no single
                               group covers them. Moved to group_vars/all/main.yml
                               where they auto-load. The ntfy_topic and
                               headscale_namespace indirection through
                               service_settings collapses to the global name.
  the four cross-host ports  - the only entries with a real justification.
                               Left in place; they move in the next commit.

Also dead, all Uptime Kuma residue or duplication:
  phoenixd_monitor_name, forgejo_runner healthcheck_timeout_seconds/retries,
  fulcrum_tailscale_hostname, and bitcoin_knots_version - the last being a
  v-prefixed copy of bitcoin_knots_version_short that nothing read, two
  hand-maintained copies of one version string.

Corrected a false comment: services_config.yml claimed the uptime_kuma subdomain
"no longer resolves to anything". It resolves to 164.92.239.72 and answers HTTP
302, and 11 playbooks still template it. Same wrong premise as PLAN_3.

Verification: all 37 playbooks' --list-tasks output is byte-identical before and
after. A probe resolving all 22 values services_config.yml used to supply returns
21 identical and one intended deletion (caddy_sites_dir, now role-only - confirmed
the role still resolves it: "Ensure Caddy sites-enabled directory exists" comes
back ok against the real path). memos check-diff identical before and after.
Syntax passes on every playbook.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 20:58:46 +02:00

345 lines
12 KiB
YAML

- name: Setup NUT (Network UPS Tools) for CyberPower UPS
hosts: hypervisor
become: true
tasks:
# ------------------------------------------------------------------
# Safety catch
#
# /etc/nut/upsd.users and /etc/nut/upsmon.conf on nodito were written by
# hand in January 2026 and carry a working password. host_vars/nodito/vault.yml
# (formerly infra/nodito/nodito_secrets.yml) still holds the literal string
# CHANGE_ME_TO_SECURE_PASSWORD, so running this play would overwrite that
# working pair with a placeholder and restart NUT - leaving the hypervisor's
# UPS unmonitored and unable to trigger a clean shutdown on mains loss.
#
# Until the real password is put in the vault, stop here.
# ansible-vault edit host_vars/nodito/vault.yml
# ------------------------------------------------------------------
- name: Refuse to run with a placeholder UPS password
assert:
that:
- ups_password is defined
- ups_password | length > 0
- ups_password != "CHANGE_ME_TO_SECURE_PASSWORD"
fail_msg: >-
ups_password is unset or still the placeholder. Applying this play would
overwrite the working /etc/nut/upsd.users and /etc/nut/upsmon.conf on
nodito and restart NUT. Put the real password in the vault first:
ansible-vault edit host_vars/nodito/vault.yml
# ------------------------------------------------------------------
# Installation
# ------------------------------------------------------------------
- name: Install NUT packages
apt:
name:
- nut
- nut-client
- nut-server
state: present
update_cache: true
# ------------------------------------------------------------------
# Verify UPS is detected
# ------------------------------------------------------------------
- name: Check if UPS is detected via USB
shell: lsusb | grep -i cyber
register: lsusb_output
changed_when: false
failed_when: false
- name: Display USB detection result
debug:
msg: "{{ lsusb_output.stdout | default('UPS not detected via USB - ensure it is plugged in') }}"
- name: Fail if UPS not detected
fail:
msg: "CyberPower UPS not detected via USB. Ensure the USB cable is connected."
when: lsusb_output.rc != 0
- name: Reload udev rules for USB permissions
shell: |
udevadm control --reload-rules
udevadm trigger --subsystem-match=usb --action=add
changed_when: true
- name: Verify USB device has nut group permissions
shell: |
BUS_DEV=$(lsusb | grep -i cyber | grep -oP 'Bus \K\d+|Device \K\d+' | tr '\n' '/' | sed 's/\/$//')
if [ -n "$BUS_DEV" ]; then
BUS=$(echo $BUS_DEV | cut -d'/' -f1)
DEV=$(echo $BUS_DEV | cut -d'/' -f2)
ls -la /dev/bus/usb/$BUS/$DEV
else
echo "UPS device not found"
exit 1
fi
register: usb_permissions
changed_when: false
- name: Display USB permissions
debug:
msg: "{{ usb_permissions.stdout }} (should show 'root nut', not 'root root')"
- name: Scan for UPS with nut-scanner
command: nut-scanner -U
register: nut_scanner_output
changed_when: false
failed_when: false
- name: Display nut-scanner result
debug:
msg: "{{ nut_scanner_output.stdout_lines }}"
# ------------------------------------------------------------------
# Configuration files
# ------------------------------------------------------------------
- name: Configure NUT mode (standalone)
template:
dest: /etc/nut/nut.conf
src: templates/nut.conf.j2
owner: root
group: nut
mode: "0640"
notify: Restart NUT services
- name: Configure UPS device
template:
dest: /etc/nut/ups.conf
src: templates/ups.conf.j2
owner: root
group: nut
mode: "0640"
notify: Restart NUT services
- name: Configure upsd to listen on localhost
template:
dest: /etc/nut/upsd.conf
src: templates/upsd.conf.j2
owner: root
group: nut
mode: "0640"
notify: Restart NUT services
- name: Configure upsd users
template:
dest: /etc/nut/upsd.users
src: templates/upsd.users.j2
owner: root
group: nut
mode: "0640"
notify: Restart NUT services
- name: Configure upsmon
template:
dest: /etc/nut/upsmon.conf
src: templates/upsmon.conf.j2
owner: root
group: nut
mode: "0640"
notify: Restart NUT services
# ------------------------------------------------------------------
# Verify late-stage shutdown script
# ------------------------------------------------------------------
- name: Verify nutshutdown script exists
stat:
path: /lib/systemd/system-shutdown/nutshutdown
register: nutshutdown_script
- name: Warn if nutshutdown script is missing
debug:
msg: "WARNING: /lib/systemd/system-shutdown/nutshutdown not found. UPS may not cut power after shutdown."
when: not nutshutdown_script.stat.exists
# ------------------------------------------------------------------
# Services
# ------------------------------------------------------------------
- name: Enable and start NUT driver enumerator
systemd:
name: nut-driver-enumerator
enabled: true
state: started
- name: Enable and start NUT server
systemd:
name: nut-server
enabled: true
state: started
- name: Enable and start NUT monitor
systemd:
name: nut-monitor
enabled: true
state: started
# ------------------------------------------------------------------
# Verification
# ------------------------------------------------------------------
- name: Wait for NUT services to stabilize
pause:
seconds: 3
- name: Verify NUT can communicate with UPS
command: upsc {{ ups_name }}@localhost
register: upsc_output
changed_when: false
failed_when: upsc_output.rc != 0
- name: Display UPS status
debug:
msg: "{{ upsc_output.stdout_lines }}"
- name: Get UPS status summary
shell: |
echo "Status: $(upsc {{ ups_name }}@localhost ups.status 2>/dev/null)"
echo "Battery: $(upsc {{ ups_name }}@localhost battery.charge 2>/dev/null)%"
echo "Runtime: $(upsc {{ ups_name }}@localhost battery.runtime 2>/dev/null)s"
echo "Load: $(upsc {{ ups_name }}@localhost ups.load 2>/dev/null)%"
register: ups_summary
changed_when: false
- name: Display UPS summary
debug:
msg: "{{ ups_summary.stdout_lines }}"
- name: Verify low battery thresholds
shell: |
echo "Runtime threshold: $(upsc {{ ups_name }}@localhost battery.runtime.low 2>/dev/null)s"
echo "Charge threshold: $(upsc {{ ups_name }}@localhost battery.charge.low 2>/dev/null)%"
register: thresholds
changed_when: false
- name: Display low battery thresholds
debug:
msg: "{{ thresholds.stdout_lines }}"
handlers:
- name: Restart NUT services
systemd:
name: "{{ item }}"
state: restarted
loop:
- nut-driver-enumerator
- nut-server
- nut-monitor
# ─────────────────────────────────────────────────────────────────────────────
# UPS heartbeat monitoring.
#
# The script decides on-mains/on-battery and says so in its exit code, which
# systemd keeps: systemctl is-failed ups-heartbeat.service
#
# Reporting anywhere else is optional. Set `healthcheck_push_url` and the script
# will also GET it with ?status=up|down; leave it empty and the exit code is
# still the whole answer. Today that URL points at Uptime Kuma, which is still
# running on watchtower but is no longer deployed by Ansible - the credentials
# were retired, the service was not. If it is ever replaced, `healthcheck_push_url`
# is the only thing that needs to change here.
#
# NOTE: this play has never actually been applied to nodito. /opt/ups-monitoring
# does not exist and there is no ups-heartbeat timer. What is on the box is a
# hand-written /usr/local/bin/ups-heartbeat.sh - mode 0644, not executable, and
# referenced by no unit and no cron entry, so nothing has ever run it. Its push
# token (uLmCPkLLO4) belongs to a monitor that DOES still exist and answer, so
# that monitor has had no heartbeat since January 2026. It is now
# healthcheck_push_urls.ups in the vault, and running this play is what will
# finally start feeding it.
# ─────────────────────────────────────────────────────────────────────────────
- name: Setup UPS Heartbeat Monitoring
hosts: hypervisor
become: true
vars:
ups_heartbeat_interval_seconds: 60
ups_heartbeat_timeout_seconds: 120
ups_heartbeat_retries: 1
ups_monitoring_script_dir: /opt/ups-monitoring
ups_monitoring_script_path: "{{ ups_monitoring_script_dir }}/ups_heartbeat.sh"
ups_log_file: "{{ ups_monitoring_script_dir }}/ups_heartbeat.log"
ups_systemd_service_name: ups-heartbeat
# Optional. Empty is fine and is not an error - see the banner above.
healthcheck_push_url: "{{ healthcheck_push_urls.ups | default('') }}"
tasks:
- name: Install required packages for UPS monitoring
package:
name:
- curl
state: present
- name: Create monitoring script directory
file:
path: "{{ ups_monitoring_script_dir }}"
state: directory
owner: root
group: root
mode: '0755'
- name: Create UPS heartbeat monitoring script
template:
dest: "{{ ups_monitoring_script_path }}"
src: templates/ups_heartbeat.sh.j2
owner: root
group: root
mode: '0755'
- name: Create systemd service for UPS heartbeat
template:
dest: "/etc/systemd/system/{{ ups_systemd_service_name }}.service"
src: templates/ups-heartbeat.service.j2
owner: root
group: root
mode: '0644'
- name: Create systemd timer for UPS heartbeat
template:
dest: "/etc/systemd/system/{{ ups_systemd_service_name }}.timer"
src: templates/ups-heartbeat.timer.j2
owner: root
group: root
mode: '0644'
- name: Reload systemd daemon
systemd:
daemon_reload: yes
- name: Enable and start UPS heartbeat timer
systemd:
name: "{{ ups_systemd_service_name }}.timer"
enabled: yes
state: started
- name: Test UPS heartbeat script
command: "{{ ups_monitoring_script_path }}"
register: script_test
changed_when: false
- name: Verify script execution
assert:
that:
- script_test.rc == 0
fail_msg: "UPS heartbeat script failed - check UPS status and communication"
- name: Display monitoring configuration
debug:
msg:
- "UPS Monitoring configured successfully"
- ""
- "NUT Configuration:"
- " UPS Name: {{ ups_name }}"
- " UPS Description: {{ ups_desc }}"
- " Off Delay: {{ ups_offdelay }}s (time after shutdown before UPS cuts power)"
- " On Delay: {{ ups_ondelay }}s (time after mains returns before UPS restores power)"
- ""
- "Health reporting:"
- " Check interval: {{ ups_heartbeat_interval_seconds }}s"
- " Push URL: {{ healthcheck_push_url | default('', true) | ternary('set', 'not set - exit code only') }}"
- ""
- "Scripts and Services:"
- " Script: {{ ups_monitoring_script_path }}"
- " Log: {{ ups_log_file }}"
- " Service: {{ ups_systemd_service_name }}.service"
- " Timer: {{ ups_systemd_service_name }}.timer"