personal_infra/ansible/roles/phoenixd/README.md

57 lines
2.3 KiB
Markdown
Raw Normal View History

phoenixd: convert to a role, de-Uptime-Kuma the health check 552-line playbook becomes 18 lines plus a 411-line role (install/service/healthcheck phases, four templates, two handlers). phoenixd_vars.yml is deleted; its content is the role's defaults. Task-list diff vs the old playbook shows ONLY the eight Uptime Kuma tasks removed - everything else identical and in the same order. phoenixd holds a Lightning node, so the run was checked against a pre-flight: before: channel 6c25fa83..., balanceSat 1550723, capacitySat 3114830, blockHeight 966692, active since 2026-09-02 after: identical, and still active since 2026-09-02 - it did NOT restart `Create phoenixd systemd service` came back unchanged, which is what proves the template reproduces the live unit byte-for-byte. changed=3 was the health check script, its unit (Environment rename), and the timer restart. Second run: changed=0. Two things the conversion fixed, both symptoms of the deprecation banner having been applied to contiguous blocks rather than to individual tasks: - The health check logged "ERROR: UPTIME_KUMA_PUSH_URL not set" on every fire - about 1,400 times a day - because its Environment= was emptied at decommissioning. The exit code was still correct so nothing was broken, but it is exactly the kind of noise that trains you to ignore a log. An unset push URL is now normal and silent. - `Enable and start phoenixd health check timer` was guarded by uptime_kuma_enabled and so had not run since the decommissioning, while the timer itself was still live on the host from before. Ansible had quietly stopped managing something that was still running. Ungated. Noted, not changed: seed.dat is mode 0644 on the host. That is phoenixd's own doing, but it is a Lightning seed and worth tightening. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-12 18:21:24 +02:00
# `phoenixd`
Deploys and runs [phoenixd](https://phoenix.acinq.co/server), an ACINQ Lightning
node, on the edge host. LNBits uses it as a wallet backend. The HTTP API stays on
loopback — phoenixd is never published through Caddy.
Converted from `deploy_phoenixd_playbook.yml` (552 lines) under Plan 6. The
playbook is now 18 lines.
## Phases
| | |
|---|---|
| `install.yml` | packages, system user, directories, versioned download and install |
| `service.yml` | systemd unit, start, then first-boot checks (config written, seed created) |
| `healthcheck.yml` | check script, unit, timer |
## The seed
`{{ phoenixd_data_dir }}/seed.dat` **is** the funds. phoenixd is deliberately
excluded from the automated backups (Plan 5, Model C): the seed is twelve fixed
words that never change, so an automated job would only manufacture more copies
of a static secret on more machines. Write them down offline, once.
Note the live file is mode `0644`. That is phoenixd's own doing, not this role's,
and it is worth tightening.
## Monitoring: one variable, no product knowledge
The check asks the node itself — the service must be active **and**
`phoenix-cli getinfo` must return a `nodeId` — and records the answer in its exit
code, which systemd keeps:
```bash
systemctl is-failed phoenixd-healthcheck.service
```
That is a complete answer with no monitoring system involved. To report
elsewhere, set `healthcheck_push_url` to anything accepting an HTTP ping. Gone
from this role: the embedded Python that created monitors over the Uptime Kuma
API, the `/tmp` credentials file, the push-URL file written and parsed back, and
the systemd `Environment=` rewrite.
### Two things the conversion fixed
**The check used to log an error once a minute.** Its `Environment=` push URL had
been empty since the decommissioning, and the script printed
`ERROR: UPTIME_KUMA_PUSH_URL not set` on every fire — roughly 1,400 times a day.
The exit code was still correct, so nothing was broken; it was pure noise, and
noise that trains you to ignore the log. An unset push URL is now normal and
silent.
**`Enable and start phoenixd health check timer` was guarded by
`uptime_kuma_enabled`** and so had not run since the decommissioning — while the
timer itself was still live on the host from before. Ansible had quietly stopped
managing something that was still running. Ungated.