personal_infra/ansible/roles/fulcrum
counterweight 3711421af5
ansible: move the cross-host ports to host_vars, delete services_config.yml
The four ports were the only entries in services_config.yml with a real
justification: each is read twice, by the role that deploys the service on its
own box AND by a socket-proxy or Caddy play that runs on the EDGE host and
publishes it. A role default is invisible to that second play.

But the shape was wrong in two ways. The file had to be named in vars_files: by
30 plays - opt-in configuration that someone will eventually forget - and five
role defaults silently interpolated service_settings.*, so bitcoin_knots,
fulcrum, datum_gateway and mempool were not self-contained: using any of them
without that one vars_file entry broke it.

Each port now lives in host_vars/<owning box>/main.yml:

  host_vars/knots_box_local/main.yml    bitcoin_p2p_port, datum_gateway_api_port,
                                        datum_gateway_stratum_port
  host_vars/fulcrum_box_local/main.yml  fulcrum_ssl_port
  host_vars/mempool_box_local/main.yml  mempool_frontend_port

host_vars auto-loads and outranks role defaults, so the owning role picks the
value up with no vars_files at all, and the edge play reads the same single
definition as hostvars['<host>'].<name>. The role defaults keep the protocol
standard (8333, 50002, ...) so each role still works standalone, with the live
deployment's value in host_vars winning.

Also fixed a fourth copy of an inventory identity: the mempool Caddy play had
"mempool-box:{{ ... }}" hardcoded in the upstream. It now derives the host from
hostvars['mempool_box_local'].ansible_host, so inventory is the only place any
box's name is written down.

services_config.yml is deleted, with 25 more vars_files entries across 19
playbooks. Between this and the previous commit, 87 vars_files entries are gone
and every variable in the repo now comes from group_vars/all, host_vars,
inventory, a role default, or that service's own *_vars.yml.

Verification: an edge-host probe resolves all eight ports and hostnames to
byte-identical values to the ones services_config.yml used to supply. Each
owning host resolves its own port through host_vars. All 37 playbooks'
--list-tasks output is unchanged. The four edge plays that consume these values
all check-diff changed=0 - the socket-proxy and Caddy units on vipy are
byte-identical, which is the direct proof the rewiring landed on the same
values. fulcrum and datum-gateway check-diff exactly as before (ok=28/changed=1
and ok=15/changed=1, both the known timer re-arm).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 21:02:57 +02:00
..
defaults ansible: move the cross-host ports to host_vars, delete services_config.yml 2026-09-13 21:02:57 +02:00
handlers fulcrum: convert to a role, de-Uptime-Kuma the health check 2026-09-13 18:02:00 +02:00
tasks fulcrum: convert to a role, de-Uptime-Kuma the health check 2026-09-13 18:02:00 +02:00
templates fulcrum: convert to a role, de-Uptime-Kuma the health check 2026-09-13 18:02:00 +02:00
README.md fulcrum: convert to a role, de-Uptime-Kuma the health check 2026-09-13 18:02:00 +02:00

fulcrum

Deploys Fulcrum, an Electrum server indexing the Bitcoin Knots node, on fulcrum-box. The second play in the calling playbook publishes its SSL port from the edge host via socket_proxy.

Converted from deploy_fulcrum_playbook.yml (685 lines) under Plan 6. The playbook is now 33 lines.

The index is the expensive thing

{{ fulcrum_db_dir }} is ~192 GB and takes days to rebuild. Nothing in this role touches it beyond state: directory with the ownership it already has (fulcrum:fulcrum 0755). Restarting Fulcrum re-opens the database; it does not reindex.

Three things this conversion fixed, all pre-existing

bitcoind pointed at the wrong machine. The vars file carried bitcoin_rpc_host: "192.168.1.140" commented "IP of knots_box_local", but .140 is fulcrum-box itself — knots-box is .135. The DHCP leases had reshuffled. The live config had already been hand-corrected to knots-box; running the playbook would have reverted it and broken indexing. Now addressed by Tailscale name, like everything else in this repo.

The restart handler was inert. It carried when: uptime_kuma_enabled | default(false), so the three tasks that notify it (SSL certificate, fulcrum.conf, systemd unit) could not restart anything. A configuration change would write to disk, report success, and silently never take effect. Ungated.

db_mem was about to quadruple. The role computes a share of RAM; on this 5931 MB host 75% is 4448 MB, leaving ~1.4 GB for the OS and Fulcrum's non-cache memory. The live value had been hand-tuned to 2048. fulcrum_db_mem_mb_override pins it. Note set_fact outranks role defaults, so the calculation has to honour the override — pinning it in defaults/ alone is silently ignored.

The health check timer, and how to read it

The timer is OnBootSec + OnUnitActiveSec with no OnCalendar. That combination has a failure mode worth knowing: OnBootSec is monotonic and elapses once; OnUnitActiveSec schedules relative to the service last being active. If the service does not run in a given boot, there is no reference to schedule from and the timer sits active and enabled doing nothing. That is exactly what had happened here — last trigger 2026-02-17, seven months of no health check, with every surface-level indicator green.

Restarting the timer does not supply that reference; running the service does. So the role runs the check once after enabling the timer, which is both the fix and a smoke test.

Diagnosing this is easy to get wrong: NextElapseUSecRealtime is always empty for a monotonic timer, so it looks broken even when it is fine. Read NextElapseUSecMonotonic, or just use systemctl list-timers.

The timer also no longer carries Requires=fulcrum.service. On a timer that means "stop watching when the watched thing stops", which is backwards for a health check.

Monitoring: one variable, no product knowledge

The check tests the Electrum TCP port and records the answer in its exit code, which systemd keeps: systemctl is-failed fulcrum-healthcheck.service. To report elsewhere set healthcheck_push_url to any endpoint accepting an HTTP ping. The Uptime Kuma API calls, monitor creation and token handling are gone.