phoenixd: convert to a role, de-Uptime-Kuma the health check

552-line playbook becomes 18 lines plus a 411-line role
(install/service/healthcheck phases, four templates, two handlers).
phoenixd_vars.yml is deleted; its content is the role's defaults.

Task-list diff vs the old playbook shows ONLY the eight Uptime Kuma tasks
removed - everything else identical and in the same order.

phoenixd holds a Lightning node, so the run was checked against a pre-flight:

  before: channel 6c25fa83..., balanceSat 1550723, capacitySat 3114830,
          blockHeight 966692, active since 2026-09-02
  after:  identical, and still active since 2026-09-02 - it did NOT restart

`Create phoenixd systemd service` came back unchanged, which is what proves the
template reproduces the live unit byte-for-byte. changed=3 was the health check
script, its unit (Environment rename), and the timer restart. Second run:
changed=0.

Two things the conversion fixed, both symptoms of the deprecation banner having
been applied to contiguous blocks rather than to individual tasks:

- The health check logged "ERROR: UPTIME_KUMA_PUSH_URL not set" on every fire -
  about 1,400 times a day - because its Environment= was emptied at
  decommissioning. The exit code was still correct so nothing was broken, but it
  is exactly the kind of noise that trains you to ignore a log. An unset push
  URL is now normal and silent.
- `Enable and start phoenixd health check timer` was guarded by
  uptime_kuma_enabled and so had not run since the decommissioning, while the
  timer itself was still live on the host from before. Ansible had quietly
  stopped managing something that was still running. Ungated.

Noted, not changed: seed.dat is mode 0644 on the host. That is phoenixd's own
doing, but it is a Lightning seed and worth tightening.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
counterweight 2026-09-12 18:21:24 +02:00
parent 73340d5fbe
commit 6c1bcbed95
Signed by: counterweight
GPG key ID: 883EDBAA726BD96C
12 changed files with 436 additions and 552 deletions

View file

@ -0,0 +1,68 @@
---
# Everything here answers "is phoenixd healthy" and records the answer. The
# Uptime Kuma specifics that used to follow it — an embedded Python script
# creating monitors over the API, a /tmp credentials file, a push-URL file read
# back and parsed, and a systemd Environment= rewrite — are gone. What reports
# where is now one variable, healthcheck_push_url. See the role README.
- name: Create phoenixd health check script
ansible.builtin.template:
src: healthcheck.sh.j2
dest: "{{ phoenixd_healthcheck_script_path }}"
owner: root
group: root
mode: "0755"
validate: "bash -n %s"
- name: Create phoenixd health check systemd service
ansible.builtin.template:
src: healthcheck.service.j2
dest: "/etc/systemd/system/{{ phoenixd_healthcheck_service_name }}.service"
owner: root
group: root
mode: "0644"
notify: Restart phoenixd health check timer
- name: Create phoenixd health check systemd timer
ansible.builtin.template:
src: healthcheck.timer.j2
dest: "/etc/systemd/system/{{ phoenixd_healthcheck_service_name }}.timer"
owner: root
group: root
mode: "0644"
notify: Restart phoenixd health check timer
- name: Reload systemd daemon after health check units
systemd:
daemon_reload: yes
# Ungated on purpose. This was guarded by `uptime_kuma_enabled`, but enabling a
# timer is deployment, not monitoring — the deprecation banner swept it up with
# the push plumbing. The timer is in fact running on the host, from before the
# decommissioning, so the guard meant Ansible had stopped managing something
# that was still live.
- name: Enable and start phoenixd health check timer
systemd:
name: "{{ phoenixd_healthcheck_service_name }}.timer"
enabled: yes
state: started
- name: Display post-install information
debug:
msg: |
✓ phoenixd {{ phoenixd_version }} deployed
Status: systemctl status phoenixd
Logs: journalctl -u phoenixd -f
CLI: sudo PHOENIX_DATADIR={{ phoenixd_data_dir }} phoenix-cli --http-bind-port {{ phoenixd_http_bind_port }} getinfo
HTTP API: http://{{ phoenixd_http_bind_ip }}:{{ phoenixd_http_bind_port }} (loopback only)
Data dir: {{ phoenixd_data_dir }}
Health: systemctl is-failed {{ phoenixd_healthcheck_service_name }}.service
API password (needed to wire LNBits up to this node):
sudo grep '^http-password=' {{ phoenixd_data_dir }}/phoenix.conf
⚠️ BACK UP THE SEED: {{ phoenixd_data_dir }}/seed.dat
Losing it means losing the funds. phoenixd is deliberately excluded
from the automated backups (Plan 5, Model C) because the seed is 12
fixed words — write them down offline, once:
sudo cat {{ phoenixd_data_dir }}/seed.dat