personal_infra/ansible/roles/phoenixd/templates/healthcheck.sh.j2
counterweight 6c1bcbed95
phoenixd: convert to a role, de-Uptime-Kuma the health check
552-line playbook becomes 18 lines plus a 411-line role
(install/service/healthcheck phases, four templates, two handlers).
phoenixd_vars.yml is deleted; its content is the role's defaults.

Task-list diff vs the old playbook shows ONLY the eight Uptime Kuma tasks
removed - everything else identical and in the same order.

phoenixd holds a Lightning node, so the run was checked against a pre-flight:

  before: channel 6c25fa83..., balanceSat 1550723, capacitySat 3114830,
          blockHeight 966692, active since 2026-09-02
  after:  identical, and still active since 2026-09-02 - it did NOT restart

`Create phoenixd systemd service` came back unchanged, which is what proves the
template reproduces the live unit byte-for-byte. changed=3 was the health check
script, its unit (Environment rename), and the timer restart. Second run:
changed=0.

Two things the conversion fixed, both symptoms of the deprecation banner having
been applied to contiguous blocks rather than to individual tasks:

- The health check logged "ERROR: UPTIME_KUMA_PUSH_URL not set" on every fire -
  about 1,400 times a day - because its Environment= was emptied at
  decommissioning. The exit code was still correct so nothing was broken, but it
  is exactly the kind of noise that trains you to ignore a log. An unset push
  URL is now normal and silent.
- `Enable and start phoenixd health check timer` was guarded by
  uptime_kuma_enabled and so had not run since the decommissioning, while the
  timer itself was still live on the host from before. Ansible had quietly
  stopped managing something that was still running. Ungated.

Noted, not changed: seed.dat is mode 0644 on the host. That is phoenixd's own
doing, but it is a Lightning seed and worth tightening.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-12 18:21:24 +02:00

39 lines
1.4 KiB
Django/Jinja

#!/bin/bash
# phoenixd health check — managed by Ansible (roles/phoenixd)
#
# Asks the node whether it is healthy and records the answer in the exit code,
# which systemd keeps:
# systemctl is-failed {{ phoenixd_healthcheck_service_name }}.service
# That is a complete answer on its own. Reporting anywhere else is optional.
PUSH_URL="${HEALTHCHECK_PUSH_URL:-}"
export PHOENIX_DATADIR="{{ phoenixd_data_dir }}"
check_phoenixd() {
# Service must be active and the node must answer getinfo.
# phoenix-cli reads the api password from $PHOENIX_DATADIR/phoenix.conf,
# but not the bind address, so pass it explicitly.
systemctl is-active --quiet phoenixd && \
{{ phoenixd_bin_dir }}/phoenix-cli \
--http-bind-ip {{ phoenixd_http_bind_ip }} \
--http-bind-port {{ phoenixd_http_bind_port }} \
getinfo 2>/dev/null | grep -q '"nodeId"'
}
report() {
local status=$1 msg=$2
# No push URL configured is NORMAL, not an error: the exit code below still
# answers the question. The previous version logged ERROR here on every
# single fire, once a minute, which is noise that trains you to ignore it.
[ -n "$PUSH_URL" ] || return 0
curl -s --max-time 10 --retry 2 -o /dev/null \
"${PUSH_URL}?status=${status}&msg=${msg// /%20}&ping=" || true
}
if check_phoenixd; then
report "up" "OK"
exit 0
else
echo "phoenixd is not responding"
report "down" "phoenixd not responding"
exit 1
fi