bitcoin-knots: convert to a role, de-Uptime-Kuma the health check

892-line playbook becomes 40 lines plus a role with install/build/configure/
service/healthcheck phases, five templates and one handler. bitcoin_knots_vars.yml
is deleted; its content is the role's defaults.

Verified after a real run: bitcoind still active since 2026-08-19 (NO restart),
chain at 966844 blocks / 875 GB, DATUM config intact, dbcache still 200, health
check timer firing again. changed=3, all health-check. vipy changed=0.

⚠ THE BIG ONE: the playbook would have deleted the mining integration.

bitcoin.conf on the node carries a section that was hand-added and was missing
from the template entirely:

    blockmaxsize=3985000
    blockmaxweight=3985000
    blocknotify=killall -USR1 datum_gateway
    maxmempool=1000
    blockreconstructionextratxn=1000000

blocknotify is how datum_gateway learns a new block landed. Running the old
playbook would have stripped all of it and solo mining would have carried on
against a stale template - a silent failure that costs money rather than raising
an error. Also dbcache 200 -> 3528 (hand-tuned down; the calculation wants 90% of
RAM) and logging moved off the file. All now reconciled, dbcache behind
bitcoin_dbcache_mb_override.

bitcoin-knots and datum-gateway are ONE SYSTEM. Noted in the README.

AND MY OWN FIX MADE IT MORE DANGEROUS. The `Restart bitcoind` handler was guarded
by uptime_kuma_enabled, so it had been inert: bitcoin.conf and the systemd unit
both notify it and neither could restart anything - a config change applied to
disk, reported success, and never took effect. Ungating that is right, but it
converts "wrong config sitting inertly on disk" into "node restarted onto a
config that breaks mining". The ungating had to land WITH the template
reconciliation, not before it.

It also raises the bar permanently: any residual template/live difference now
restarts a Bitcoin node on every run. Four rounds of --check --diff to reach
changed=0 - the DATUM section, an explanatory comment that was rendering into the
deployed config (now a {# #} Jinja comment), a "# Pruning (optional)" comment the
live file had, and one trailing blank line.

The build path is 32 tasks all guarded by `not bitcoind_binary_exists.stat.exists`,
so a converged host skips the 30-60 minute compile and both `state: absent`
deletions. Those target /opt/bitcoin-knots/{source,bitcoin-<version>}; the chain
is in /mnt/knots_data and is never touched. Signature-verification tasks copied
verbatim.

The health check timer had last fired 2026-08-09 while reporting active/enabled -
same OnBootSec + OnUnitActiveSec dead chain as fulcrum. The role runs the check
once after enabling to supply the reference the timer schedules from.

Ownership parity checked mechanically against `git show HEAD:`, keyed by TASK
NAME rather than path - keying by path gave a false positive, because
bitcoin_knots_source_dir is created with ownership and later removed with
state: absent, so whichever task comes last wins and that differs between one
file and five. 13/13 match.

bitcoin_p2p_port and the tailscale hostname moved to services_config.yml for the
socket-proxy play on the edge host.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
counterweight 2026-09-13 18:13:40 +02:00
parent e83191c029
commit d26dc78b3c
Signed by: counterweight
GPG key ID: 883EDBAA726BD96C
19 changed files with 1163 additions and 1250 deletions

View file

@ -0,0 +1,85 @@
# `bitcoin_knots`
Builds Bitcoin Knots from source with PGP + SHA256 verification of the release
tarball, runs it as a full node on `knots-box`, and keeps a health check on a
systemd timer. The second play in the calling playbook publishes the P2P port
from the edge host via `socket_proxy`.
Converted from `deploy_bitcoin_knots_playbook.yml` (892 lines) under Plan 6. The
playbook is now 40 lines.
## The build is guarded; the chain is never touched
`build.yml` is 32 tasks, every one carrying
`when: not bitcoind_binary_exists.stat.exists`. On a host that already has the
binary the whole download / verify / 30-60 minute compile skips — **including the
two `state: absent` deletions**, which target `/opt/bitcoin-knots/source` and the
extracted build directory.
The chain lives elsewhere and nothing here touches it:
| | |
|---|---|
| `bitcoin_knots_dir` | `/opt/bitcoin-knots` — build tree, safe to delete |
| `bitcoin_data_dir` | `/var/lib/bitcoin` — config, logs, wallets |
| `bitcoin_large_data_dir` | `/mnt/knots_data`**~875 GB of blockchain** |
The signature-verification tasks are the security control of this role. They are
copied verbatim; do not "simplify" them.
## ⚠ This node is half of the mining setup
`bitcoin.conf` carries a DATUM Gateway section that was hand-added on the node
and was **missing from the playbook's template**:
```ini
blockmaxsize=3985000
blockmaxweight=3985000
blocknotify=killall -USR1 datum_gateway
maxmempool=1000
blockreconstructionextratxn=1000000
```
`blocknotify` is how `datum_gateway` learns a new block landed. Running the old
playbook would have deleted all of it, and solo mining would have carried on
grinding against a stale template — a silent failure that costs money rather
than raising an error. The template now carries it behind
`bitcoin_datum_gateway_enabled`.
**bitcoin-knots and datum-gateway are one system, not two services.** Changing
either config means thinking about both.
## The restart handler, and why exactness matters now
The hand-written `Restart bitcoind` handler carried
`when: uptime_kuma_enabled | default(false)`, so it had been inert since the
decommissioning: `bitcoin.conf` and the systemd unit both notify it and neither
could restart anything. A config change applied to disk, reported success, and
never took effect.
It is ungated here — which raises the bar for the template. **Any** residual
difference between the template and the live file, down to a trailing newline,
means the task reports `changed` and restarts a Bitcoin node on every run. It
took four rounds of `--check --diff` to reach `changed=0`: the DATUM section, an
explanatory comment that was rendering into the deployed file (now a `{# #}`
Jinja comment), a `# Pruning (optional)` comment the live file had, and one
trailing blank line.
## `dbcache`
Computed as 90% of RAM unless `bitcoin_dbcache_mb_override` is set. The live node
was hand-tuned to **200 MB**; the calculation produces 3528. As with fulcrum,
`set_fact` outranks role defaults, so the *calculation* honours the override — a
value pinned only in `defaults/` is silently ignored.
## Monitoring: one variable, no product knowledge
The check tests bitcoind's RPC and records the answer in its exit code, which
systemd keeps: `systemctl is-failed bitcoin-knots-healthcheck.service`. Set
`healthcheck_push_url` to report anywhere that accepts an HTTP ping.
The timer had last fired **2026-08-09** while still reporting `active` and
`enabled` — the same `OnBootSec` + `OnUnitActiveSec` dead chain as fulcrum, where
nothing re-arms it if the service does not run in a given boot. The role runs the
check once after enabling, which both smoke-tests it and supplies the reference
the timer schedules from.