personal_infra/ansible/roles/fulcrum/README.md
counterweight e83191c029
fulcrum: convert to a role, de-Uptime-Kuma the health check
685-line playbook becomes 33 lines plus a 426-line role
(install/service/healthcheck phases, six templates, one handler).
fulcrum_vars.yml is deleted; its content is the role's defaults.

Verified: fulcrum untouched - active since 2026-07-29 (no restart), 192G datadir,
height 966842, bitcoind and db_mem unchanged on disk. Second run changed=1 (the
arming run, changed_when: false aside). vipy changed=0.

THREE PRE-EXISTING LANDMINES the check-mode diff caught, any of which a faithful
extraction would have detonated:

- bitcoin_rpc_host was "192.168.1.140", commented "IP of knots_box_local". But
  .140 is fulcrum-box ITSELF; knots-box is .135. The DHCP leases had reshuffled -
  the fifth instance of this same disease in this estate. The live config had
  been hand-corrected to knots-box; running the playbook would have reverted it
  and pointed Fulcrum at itself. Now addressed by Tailscale name.

- The `Restart fulcrum` handler was guarded by uptime_kuma_enabled, so the three
  tasks that notify it (SSL cert, fulcrum.conf, systemd unit) could not restart
  anything. A config change applied to disk, reported success, and silently never
  took effect. That is worse than the other banner casualties: it makes the
  deployment itself lie. Ungated.

- db_mem was about to go 2048 -> 4448 (75% of 5931MB RAM), leaving ~1.4GB for the
  OS and Fulcrum's non-cache memory. The live value had been hand-tuned down.
  fulcrum_db_mem_mb_override pins it. Note set_fact outranks role defaults, so
  the calculation itself has to honour the override.

MY OWN ERROR, third instance: retyping `copy:` as `template:` lost `owner:` on
the banner and on fulcrum.conf. Rather than keep catching these by eye, every
managed path's owner/group/mode is now compared against `git show HEAD:`
mechanically - 12/12 match.

The health check timer had not fired since 2026-02-17 while reporting `active`
and `enabled`. It is OnBootSec + OnUnitActiveSec with no OnCalendar: OnBootSec
elapses once, and OnUnitActiveSec needs the SERVICE to have run this boot to have
anything to schedule from. Restarting the timer does not supply that; running the
service does, so the role now runs the check once after enabling. Also dropped
`Requires=fulcrum.service` from the timer - on a timer that means "stop watching
when the watched thing stops".

Diagnostic note: NextElapseUSecRealtime is always empty for a monotonic timer, so
it reads as broken even when healthy. I misread it once and wrongly called the
timer dead. Use NextElapseUSecMonotonic or systemctl list-timers.

fulcrum_ssl_port and fulcrum_tailscale_hostname moved to services_config.yml -
the socket-proxy play on the edge host needs them and a role default cannot reach
a second play.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 18:02:00 +02:00

65 lines
3.2 KiB
Markdown

# `fulcrum`
Deploys [Fulcrum](https://github.com/cculianu/Fulcrum), an Electrum server
indexing the Bitcoin Knots node, on `fulcrum-box`. The second play in the
calling playbook publishes its SSL port from the edge host via `socket_proxy`.
Converted from `deploy_fulcrum_playbook.yml` (685 lines) under Plan 6. The
playbook is now 33 lines.
## The index is the expensive thing
`{{ fulcrum_db_dir }}` is ~192 GB and takes days to rebuild. Nothing in this role
touches it beyond `state: directory` with the ownership it already has
(`fulcrum:fulcrum 0755`). Restarting Fulcrum re-opens the database; it does not
reindex.
## Three things this conversion fixed, all pre-existing
**`bitcoind` pointed at the wrong machine.** The vars file carried
`bitcoin_rpc_host: "192.168.1.140"` commented "IP of knots_box_local", but `.140`
is **fulcrum-box itself** — knots-box is `.135`. The DHCP leases had reshuffled.
The live config had already been hand-corrected to `knots-box`; running the
playbook would have reverted it and broken indexing. Now addressed by Tailscale
name, like everything else in this repo.
**The restart handler was inert.** It carried
`when: uptime_kuma_enabled | default(false)`, so the three tasks that notify it
(SSL certificate, `fulcrum.conf`, systemd unit) could not restart anything. A
configuration change would write to disk, report success, and silently never take
effect. Ungated.
**`db_mem` was about to quadruple.** The role computes a share of RAM; on this
5931 MB host 75% is 4448 MB, leaving ~1.4 GB for the OS and Fulcrum's non-cache
memory. The live value had been hand-tuned to 2048. `fulcrum_db_mem_mb_override`
pins it. Note `set_fact` outranks role defaults, so the *calculation* has to
honour the override — pinning it in `defaults/` alone is silently ignored.
## The health check timer, and how to read it
The timer is `OnBootSec` + `OnUnitActiveSec` with no `OnCalendar`. That
combination has a failure mode worth knowing: `OnBootSec` is monotonic and
elapses once; `OnUnitActiveSec` schedules relative to the **service** last being
active. If the service does not run in a given boot, there is no reference to
schedule from and the timer sits `active` and `enabled` doing nothing. That is
exactly what had happened here — last trigger **2026-02-17**, seven months of no
health check, with every surface-level indicator green.
Restarting the timer does not supply that reference; running the service does.
So the role runs the check once after enabling the timer, which is both the fix
and a smoke test.
**Diagnosing this is easy to get wrong**: `NextElapseUSecRealtime` is always
empty for a monotonic timer, so it looks broken even when it is fine. Read
`NextElapseUSecMonotonic`, or just use `systemctl list-timers`.
The timer also no longer carries `Requires=fulcrum.service`. On a timer that
means "stop watching when the watched thing stops", which is backwards for a
health check.
## Monitoring: one variable, no product knowledge
The check tests the Electrum TCP port and records the answer in its exit code,
which systemd keeps: `systemctl is-failed fulcrum-healthcheck.service`. To report
elsewhere set `healthcheck_push_url` to any endpoint accepting an HTTP ping. The
Uptime Kuma API calls, monitor creation and token handling are gone.