The four ports were the only entries in services_config.yml with a real
justification: each is read twice, by the role that deploys the service on its
own box AND by a socket-proxy or Caddy play that runs on the EDGE host and
publishes it. A role default is invisible to that second play.
But the shape was wrong in two ways. The file had to be named in vars_files: by
30 plays - opt-in configuration that someone will eventually forget - and five
role defaults silently interpolated service_settings.*, so bitcoin_knots,
fulcrum, datum_gateway and mempool were not self-contained: using any of them
without that one vars_file entry broke it.
Each port now lives in host_vars/<owning box>/main.yml:
host_vars/knots_box_local/main.yml bitcoin_p2p_port, datum_gateway_api_port,
datum_gateway_stratum_port
host_vars/fulcrum_box_local/main.yml fulcrum_ssl_port
host_vars/mempool_box_local/main.yml mempool_frontend_port
host_vars auto-loads and outranks role defaults, so the owning role picks the
value up with no vars_files at all, and the edge play reads the same single
definition as hostvars['<host>'].<name>. The role defaults keep the protocol
standard (8333, 50002, ...) so each role still works standalone, with the live
deployment's value in host_vars winning.
Also fixed a fourth copy of an inventory identity: the mempool Caddy play had
"mempool-box:{{ ... }}" hardcoded in the upstream. It now derives the host from
hostvars['mempool_box_local'].ansible_host, so inventory is the only place any
box's name is written down.
services_config.yml is deleted, with 25 more vars_files entries across 19
playbooks. Between this and the previous commit, 87 vars_files entries are gone
and every variable in the repo now comes from group_vars/all, host_vars,
inventory, a role default, or that service's own *_vars.yml.
Verification: an edge-host probe resolves all eight ports and hostnames to
byte-identical values to the ones services_config.yml used to supply. Each
owning host resolves its own port through host_vars. All 37 playbooks'
--list-tasks output is unchanged. The four edge plays that consume these values
all check-diff changed=0 - the socket-proxy and Caddy units on vipy are
byte-identical, which is the direct proof the rewiring landed on the same
values. fulcrum and datum-gateway check-diff exactly as before (ok=28/changed=1
and ok=15/changed=1, both the known timer re-arm).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Three files existed only as second copies of things group_vars/all already
auto-loads, and 34 playbooks named them in vars_files: - which outranks
group_vars, so the copies won. The day someone edited one and not the other,
those plays would silently keep the stale value. infra_vars.yml was already
drifting: group_vars/all/main.yml had grown age_backup_recipient and
backup_pull_public_key that it lacked.
infra_vars.yml - a strict subset of group_vars/all/main.yml
infra_secrets.yml - decrypts byte-identical to group_vars/all/vault.yml
infra_secrets.yml.example - documented Uptime Kuma credentials as the reason
the file exists, which stopped being true
Deleted, along with 62 vars_files entries across 34 playbooks (12 of which
named ../../group_vars/all/main.yml directly - same defect, a vars_files entry
duplicating an auto-loaded file at higher precedence than the file itself).
Checked before touching anything: infra_secrets.yml was listed LAST in 10 plays,
after services_config.yml, so removal would flip precedence if the two shared a
key. They share none, and neither does services_config.yml with
group_vars/all/main.yml, so the removal is provably inert.
services_config.yml was the last one standing. It held four unrelated things:
caddy_sites_dir - an identical copy of roles/caddy_site/defaults/.
Deleted; the role default is now the only one.
*.tailscale_hostname (x3) - a THIRD copy of each box's identity, which
inventory.ini already holds as ansible_host.
Deleted. Edge plays now read
hostvars['<host>'].ansible_host - verified an
edge play resolves that with nothing loaded and
the other host in no play. Three copies of one
name is how bitcoin_rpc_host ended up labelled
"knots_box" while pointing at fulcrum-box.
subdomains, ntfy topic, - genuinely global: their readers span managed,
headscale namespace monitoring, vpn_control and edge, so no single
group covers them. Moved to group_vars/all/main.yml
where they auto-load. The ntfy_topic and
headscale_namespace indirection through
service_settings collapses to the global name.
the four cross-host ports - the only entries with a real justification.
Left in place; they move in the next commit.
Also dead, all Uptime Kuma residue or duplication:
phoenixd_monitor_name, forgejo_runner healthcheck_timeout_seconds/retries,
fulcrum_tailscale_hostname, and bitcoin_knots_version - the last being a
v-prefixed copy of bitcoin_knots_version_short that nothing read, two
hand-maintained copies of one version string.
Corrected a false comment: services_config.yml claimed the uptime_kuma subdomain
"no longer resolves to anything". It resolves to 164.92.239.72 and answers HTTP
302, and 11 playbooks still template it. Same wrong premise as PLAN_3.
Verification: all 37 playbooks' --list-tasks output is byte-identical before and
after. A probe resolving all 22 values services_config.yml used to supply returns
21 identical and one intended deletion (caddy_sites_dir, now role-only - confirmed
the role still resolves it: "Ensure Caddy sites-enabled directory exists" comes
back ok against the real path). memos check-diff identical before and after.
Syntax passes on every playbook.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
These two are host-specific by design - nodito is a pet, not cattle - so they
stay playbooks rather than becoming roles. But they had rotted.
De-Kuma, following the pattern of the six service roles:
Both plays opened with an assert on uptime_kuma_username/password, which were
removed from the vault, so both failed before doing anything. Dropped that,
the two embedded Python monitor-creation scripts, and their /tmp cleanup.
Kept every check, threshold and systemd timer - those are the durable part.
Reporting is now generic: `healthcheck_push_url` goes into the unit as
Environment=HEALTHCHECK_PUSH_URL and the script reads ${HEALTHCHECK_PUSH_URL:-},
treating empty as normal rather than an error. The exit code is the real
answer; systemd keeps it. Both scripts now also report status=down on failure
instead of only going silent. Live push URLs harvested into the vault so
nothing observable changes for ZFS.
Three live bugs found while check-diffing:
1. 32_zfs would have DE-REGISTERED the Proxmox storage. `pvesm remove` was
gated on the storage existing and `pvesm add` on it NOT existing - mutually
exclusive - so a real run removed the proxmox-tank-1 entry backing every VM
and never put it back. It would also have dropped `mountpoint /var/lib/vz`,
which the live entry has and `pvesm add` does not set. Registration is now
add-only.
2. zfs_disk_1 named ata-...WX11TN0Z, a disk no longer in the machine. The live
mirror is WX120LHQ + WX11TN2P; a leg was replaced and the repo never caught
up. Inert behind `when: zfs_pool_exists.rc != 0`, but wrong on any
disaster-recovery run. This is the seventh instance of an identifier
written down once whose hardware later moved.
3. 34_nut has NEVER been applied to nodito - no /etc/nut file carries the
"Managed by Ansible" marker; they were written by hand in January 2026. The
vault held the literal CHANGE_ME_TO_SECURE_PASSWORD, so applying it would
have overwritten a working upsd/upsmon auth pair with a placeholder and
restarted NUT, leaving the hypervisor's UPS unable to trigger a clean
shutdown on mains loss. The Kuma assert was the only thing stopping that,
so removing it without a replacement would have armed the gun: there is now
an explicit assert that refuses to run on the placeholder. The real
password is in the vault and `Configure upsd users` check-diffs clean.
Templates reconciled with the live files first, so applying 34_nut is close to
a no-op: added `maxretry = 3` to ups.conf and OFFDURATION / RBWARNTIME /
NOCOMMWARNTIME / FINALDELAY plus quoted POWERDOWNFLAG to upsmon.conf. The only
substantive additions left are the NOTIFYMSG/NOTIFYFLAG syslog lines.
/usr/local/bin/ups-heartbeat.sh on the box is an orphan - mode 0644, not
executable, referenced by no unit and no cron entry - but its push token
belongs to a monitor that still exists and answers, so that monitor has had no
heartbeat since January. Harvested as healthcheck_push_urls.ups; applying this
play is what will finally feed it.
Finish the host_vars migration:
infra/nodito/nodito_vars.yml was byte-identical to host_vars/nodito/main.yml.
Deleted it, moved nodito_secrets.yml to host_vars/nodito/vault.yml, and
stripped the dead vars_files entries from all three playbooks.
Extract the 13 inline `content: |` blocks to infra/nodito/templates/, pulled via
a YAML load rather than retyped. 681+582 lines become 313+344 plus templates.
Ownership parity against HEAD checked mechanically: no owner/group/mode drift on
any surviving task.
Neither playbook has been applied. ZFS play 2 check-runs failed=0 with two
changes: the script rewrite (the deployed one is a hand-edited DEBUG VERSION
with `set -x`) and the Environment line.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
These five plus the ntfy notification playbook assert on the credentials, so
they now fail immediately instead of running — deliberately, before anything is
installed. The banner says so and points at archive/uptime_kuma/.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Drops uptime_kuma_username/password from infra_secrets.yml and its group_vars
copy and example, plus a dead push token in nodito_secrets.yml that nothing
referenced. They remain in git history — rotation is what actually retires them.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>