Three files existed only as second copies of things group_vars/all already
auto-loads, and 34 playbooks named them in vars_files: - which outranks
group_vars, so the copies won. The day someone edited one and not the other,
those plays would silently keep the stale value. infra_vars.yml was already
drifting: group_vars/all/main.yml had grown age_backup_recipient and
backup_pull_public_key that it lacked.
infra_vars.yml - a strict subset of group_vars/all/main.yml
infra_secrets.yml - decrypts byte-identical to group_vars/all/vault.yml
infra_secrets.yml.example - documented Uptime Kuma credentials as the reason
the file exists, which stopped being true
Deleted, along with 62 vars_files entries across 34 playbooks (12 of which
named ../../group_vars/all/main.yml directly - same defect, a vars_files entry
duplicating an auto-loaded file at higher precedence than the file itself).
Checked before touching anything: infra_secrets.yml was listed LAST in 10 plays,
after services_config.yml, so removal would flip precedence if the two shared a
key. They share none, and neither does services_config.yml with
group_vars/all/main.yml, so the removal is provably inert.
services_config.yml was the last one standing. It held four unrelated things:
caddy_sites_dir - an identical copy of roles/caddy_site/defaults/.
Deleted; the role default is now the only one.
*.tailscale_hostname (x3) - a THIRD copy of each box's identity, which
inventory.ini already holds as ansible_host.
Deleted. Edge plays now read
hostvars['<host>'].ansible_host - verified an
edge play resolves that with nothing loaded and
the other host in no play. Three copies of one
name is how bitcoin_rpc_host ended up labelled
"knots_box" while pointing at fulcrum-box.
subdomains, ntfy topic, - genuinely global: their readers span managed,
headscale namespace monitoring, vpn_control and edge, so no single
group covers them. Moved to group_vars/all/main.yml
where they auto-load. The ntfy_topic and
headscale_namespace indirection through
service_settings collapses to the global name.
the four cross-host ports - the only entries with a real justification.
Left in place; they move in the next commit.
Also dead, all Uptime Kuma residue or duplication:
phoenixd_monitor_name, forgejo_runner healthcheck_timeout_seconds/retries,
fulcrum_tailscale_hostname, and bitcoin_knots_version - the last being a
v-prefixed copy of bitcoin_knots_version_short that nothing read, two
hand-maintained copies of one version string.
Corrected a false comment: services_config.yml claimed the uptime_kuma subdomain
"no longer resolves to anything". It resolves to 164.92.239.72 and answers HTTP
302, and 11 playbooks still template it. Same wrong premise as PLAN_3.
Verification: all 37 playbooks' --list-tasks output is byte-identical before and
after. A probe resolving all 22 values services_config.yml used to supply returns
21 identical and one intended deletion (caddy_sites_dir, now role-only - confirmed
the role still resolves it: "Ensure Caddy sites-enabled directory exists" comes
back ok against the real path). memos check-diff identical before and after.
Syntax passes on every playbook.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|---|---|---|
| .. | ||
| defaults | ||
| handlers | ||
| tasks | ||
| templates | ||
| README.md | ||
fulcrum
Deploys Fulcrum, an Electrum server
indexing the Bitcoin Knots node, on fulcrum-box. The second play in the
calling playbook publishes its SSL port from the edge host via socket_proxy.
Converted from deploy_fulcrum_playbook.yml (685 lines) under Plan 6. The
playbook is now 33 lines.
The index is the expensive thing
{{ fulcrum_db_dir }} is ~192 GB and takes days to rebuild. Nothing in this role
touches it beyond state: directory with the ownership it already has
(fulcrum:fulcrum 0755). Restarting Fulcrum re-opens the database; it does not
reindex.
Three things this conversion fixed, all pre-existing
bitcoind pointed at the wrong machine. The vars file carried
bitcoin_rpc_host: "192.168.1.140" commented "IP of knots_box_local", but .140
is fulcrum-box itself — knots-box is .135. The DHCP leases had reshuffled.
The live config had already been hand-corrected to knots-box; running the
playbook would have reverted it and broken indexing. Now addressed by Tailscale
name, like everything else in this repo.
The restart handler was inert. It carried
when: uptime_kuma_enabled | default(false), so the three tasks that notify it
(SSL certificate, fulcrum.conf, systemd unit) could not restart anything. A
configuration change would write to disk, report success, and silently never take
effect. Ungated.
db_mem was about to quadruple. The role computes a share of RAM; on this
5931 MB host 75% is 4448 MB, leaving ~1.4 GB for the OS and Fulcrum's non-cache
memory. The live value had been hand-tuned to 2048. fulcrum_db_mem_mb_override
pins it. Note set_fact outranks role defaults, so the calculation has to
honour the override — pinning it in defaults/ alone is silently ignored.
The health check timer, and how to read it
The timer is OnBootSec + OnUnitActiveSec with no OnCalendar. That
combination has a failure mode worth knowing: OnBootSec is monotonic and
elapses once; OnUnitActiveSec schedules relative to the service last being
active. If the service does not run in a given boot, there is no reference to
schedule from and the timer sits active and enabled doing nothing. That is
exactly what had happened here — last trigger 2026-02-17, seven months of no
health check, with every surface-level indicator green.
Restarting the timer does not supply that reference; running the service does. So the role runs the check once after enabling the timer, which is both the fix and a smoke test.
Diagnosing this is easy to get wrong: NextElapseUSecRealtime is always
empty for a monotonic timer, so it looks broken even when it is fine. Read
NextElapseUSecMonotonic, or just use systemctl list-timers.
The timer also no longer carries Requires=fulcrum.service. On a timer that
means "stop watching when the watched thing stops", which is backwards for a
health check.
Monitoring: one variable, no product knowledge
The check tests the Electrum TCP port and records the answer in its exit code,
which systemd keeps: systemctl is-failed fulcrum-healthcheck.service. To report
elsewhere set healthcheck_push_url to any endpoint accepting an HTTP ping. The
Uptime Kuma API calls, monitor creation and token handling are gone.