personal_infra/ansible/roles/phoenixd
counterweight 954b683c71
ansible: delete the duplicated vars files, move globals to group_vars/all
Three files existed only as second copies of things group_vars/all already
auto-loads, and 34 playbooks named them in vars_files: - which outranks
group_vars, so the copies won. The day someone edited one and not the other,
those plays would silently keep the stale value. infra_vars.yml was already
drifting: group_vars/all/main.yml had grown age_backup_recipient and
backup_pull_public_key that it lacked.

  infra_vars.yml          - a strict subset of group_vars/all/main.yml
  infra_secrets.yml       - decrypts byte-identical to group_vars/all/vault.yml
  infra_secrets.yml.example - documented Uptime Kuma credentials as the reason
                              the file exists, which stopped being true

Deleted, along with 62 vars_files entries across 34 playbooks (12 of which
named ../../group_vars/all/main.yml directly - same defect, a vars_files entry
duplicating an auto-loaded file at higher precedence than the file itself).

Checked before touching anything: infra_secrets.yml was listed LAST in 10 plays,
after services_config.yml, so removal would flip precedence if the two shared a
key. They share none, and neither does services_config.yml with
group_vars/all/main.yml, so the removal is provably inert.

services_config.yml was the last one standing. It held four unrelated things:

  caddy_sites_dir            - an identical copy of roles/caddy_site/defaults/.
                               Deleted; the role default is now the only one.
  *.tailscale_hostname (x3)  - a THIRD copy of each box's identity, which
                               inventory.ini already holds as ansible_host.
                               Deleted. Edge plays now read
                               hostvars['<host>'].ansible_host - verified an
                               edge play resolves that with nothing loaded and
                               the other host in no play. Three copies of one
                               name is how bitcoin_rpc_host ended up labelled
                               "knots_box" while pointing at fulcrum-box.
  subdomains, ntfy topic,    - genuinely global: their readers span managed,
  headscale namespace          monitoring, vpn_control and edge, so no single
                               group covers them. Moved to group_vars/all/main.yml
                               where they auto-load. The ntfy_topic and
                               headscale_namespace indirection through
                               service_settings collapses to the global name.
  the four cross-host ports  - the only entries with a real justification.
                               Left in place; they move in the next commit.

Also dead, all Uptime Kuma residue or duplication:
  phoenixd_monitor_name, forgejo_runner healthcheck_timeout_seconds/retries,
  fulcrum_tailscale_hostname, and bitcoin_knots_version - the last being a
  v-prefixed copy of bitcoin_knots_version_short that nothing read, two
  hand-maintained copies of one version string.

Corrected a false comment: services_config.yml claimed the uptime_kuma subdomain
"no longer resolves to anything". It resolves to 164.92.239.72 and answers HTTP
302, and 11 playbooks still template it. Same wrong premise as PLAN_3.

Verification: all 37 playbooks' --list-tasks output is byte-identical before and
after. A probe resolving all 22 values services_config.yml used to supply returns
21 identical and one intended deletion (caddy_sites_dir, now role-only - confirmed
the role still resolves it: "Ensure Caddy sites-enabled directory exists" comes
back ok against the real path). memos check-diff identical before and after.
Syntax passes on every playbook.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 20:58:46 +02:00
..
defaults ansible: delete the duplicated vars files, move globals to group_vars/all 2026-09-13 20:58:46 +02:00
handlers phoenixd: convert to a role, de-Uptime-Kuma the health check 2026-09-12 18:21:24 +02:00
tasks phoenixd: convert to a role, de-Uptime-Kuma the health check 2026-09-12 18:21:24 +02:00
templates phoenixd: convert to a role, de-Uptime-Kuma the health check 2026-09-12 18:21:24 +02:00
README.md phoenixd: convert to a role, de-Uptime-Kuma the health check 2026-09-12 18:21:24 +02:00

phoenixd

Deploys and runs phoenixd, an ACINQ Lightning node, on the edge host. LNBits uses it as a wallet backend. The HTTP API stays on loopback — phoenixd is never published through Caddy.

Converted from deploy_phoenixd_playbook.yml (552 lines) under Plan 6. The playbook is now 18 lines.

Phases

install.yml packages, system user, directories, versioned download and install
service.yml systemd unit, start, then first-boot checks (config written, seed created)
healthcheck.yml check script, unit, timer

The seed

{{ phoenixd_data_dir }}/seed.dat is the funds. phoenixd is deliberately excluded from the automated backups (Plan 5, Model C): the seed is twelve fixed words that never change, so an automated job would only manufacture more copies of a static secret on more machines. Write them down offline, once.

Note the live file is mode 0644. That is phoenixd's own doing, not this role's, and it is worth tightening.

Monitoring: one variable, no product knowledge

The check asks the node itself — the service must be active and phoenix-cli getinfo must return a nodeId — and records the answer in its exit code, which systemd keeps:

systemctl is-failed phoenixd-healthcheck.service

That is a complete answer with no monitoring system involved. To report elsewhere, set healthcheck_push_url to anything accepting an HTTP ping. Gone from this role: the embedded Python that created monitors over the Uptime Kuma API, the /tmp credentials file, the push-URL file written and parsed back, and the systemd Environment= rewrite.

Two things the conversion fixed

The check used to log an error once a minute. Its Environment= push URL had been empty since the decommissioning, and the script printed ERROR: UPTIME_KUMA_PUSH_URL not set on every fire — roughly 1,400 times a day. The exit code was still correct, so nothing was broken; it was pure noise, and noise that trains you to ignore the log. An unset push URL is now normal and silent.

Enable and start phoenixd health check timer was guarded by uptime_kuma_enabled and so had not run since the decommissioning — while the timer itself was still live on the host from before. Ansible had quietly stopped managing something that was still running. Ungated.