66 lines
3.2 KiB
Markdown
66 lines
3.2 KiB
Markdown
|
|
# `fulcrum`
|
||
|
|
|
||
|
|
Deploys [Fulcrum](https://github.com/cculianu/Fulcrum), an Electrum server
|
||
|
|
indexing the Bitcoin Knots node, on `fulcrum-box`. The second play in the
|
||
|
|
calling playbook publishes its SSL port from the edge host via `socket_proxy`.
|
||
|
|
|
||
|
|
Converted from `deploy_fulcrum_playbook.yml` (685 lines) under Plan 6. The
|
||
|
|
playbook is now 33 lines.
|
||
|
|
|
||
|
|
## The index is the expensive thing
|
||
|
|
|
||
|
|
`{{ fulcrum_db_dir }}` is ~192 GB and takes days to rebuild. Nothing in this role
|
||
|
|
touches it beyond `state: directory` with the ownership it already has
|
||
|
|
(`fulcrum:fulcrum 0755`). Restarting Fulcrum re-opens the database; it does not
|
||
|
|
reindex.
|
||
|
|
|
||
|
|
## Three things this conversion fixed, all pre-existing
|
||
|
|
|
||
|
|
**`bitcoind` pointed at the wrong machine.** The vars file carried
|
||
|
|
`bitcoin_rpc_host: "192.168.1.140"` commented "IP of knots_box_local", but `.140`
|
||
|
|
is **fulcrum-box itself** — knots-box is `.135`. The DHCP leases had reshuffled.
|
||
|
|
The live config had already been hand-corrected to `knots-box`; running the
|
||
|
|
playbook would have reverted it and broken indexing. Now addressed by Tailscale
|
||
|
|
name, like everything else in this repo.
|
||
|
|
|
||
|
|
**The restart handler was inert.** It carried
|
||
|
|
`when: uptime_kuma_enabled | default(false)`, so the three tasks that notify it
|
||
|
|
(SSL certificate, `fulcrum.conf`, systemd unit) could not restart anything. A
|
||
|
|
configuration change would write to disk, report success, and silently never take
|
||
|
|
effect. Ungated.
|
||
|
|
|
||
|
|
**`db_mem` was about to quadruple.** The role computes a share of RAM; on this
|
||
|
|
5931 MB host 75% is 4448 MB, leaving ~1.4 GB for the OS and Fulcrum's non-cache
|
||
|
|
memory. The live value had been hand-tuned to 2048. `fulcrum_db_mem_mb_override`
|
||
|
|
pins it. Note `set_fact` outranks role defaults, so the *calculation* has to
|
||
|
|
honour the override — pinning it in `defaults/` alone is silently ignored.
|
||
|
|
|
||
|
|
## The health check timer, and how to read it
|
||
|
|
|
||
|
|
The timer is `OnBootSec` + `OnUnitActiveSec` with no `OnCalendar`. That
|
||
|
|
combination has a failure mode worth knowing: `OnBootSec` is monotonic and
|
||
|
|
elapses once; `OnUnitActiveSec` schedules relative to the **service** last being
|
||
|
|
active. If the service does not run in a given boot, there is no reference to
|
||
|
|
schedule from and the timer sits `active` and `enabled` doing nothing. That is
|
||
|
|
exactly what had happened here — last trigger **2026-02-17**, seven months of no
|
||
|
|
health check, with every surface-level indicator green.
|
||
|
|
|
||
|
|
Restarting the timer does not supply that reference; running the service does.
|
||
|
|
So the role runs the check once after enabling the timer, which is both the fix
|
||
|
|
and a smoke test.
|
||
|
|
|
||
|
|
**Diagnosing this is easy to get wrong**: `NextElapseUSecRealtime` is always
|
||
|
|
empty for a monotonic timer, so it looks broken even when it is fine. Read
|
||
|
|
`NextElapseUSecMonotonic`, or just use `systemctl list-timers`.
|
||
|
|
|
||
|
|
The timer also no longer carries `Requires=fulcrum.service`. On a timer that
|
||
|
|
means "stop watching when the watched thing stops", which is backwards for a
|
||
|
|
health check.
|
||
|
|
|
||
|
|
## Monitoring: one variable, no product knowledge
|
||
|
|
|
||
|
|
The check tests the Electrum TCP port and records the answer in its exit code,
|
||
|
|
which systemd keeps: `systemctl is-failed fulcrum-healthcheck.service`. To report
|
||
|
|
elsewhere set `healthcheck_push_url` to any endpoint accepting an HTTP ping. The
|
||
|
|
Uptime Kuma API calls, monitor creation and token handling are gone.
|