personal_infra/ansible/roles/bitcoin_knots
counterweight 3711421af5
ansible: move the cross-host ports to host_vars, delete services_config.yml
The four ports were the only entries in services_config.yml with a real
justification: each is read twice, by the role that deploys the service on its
own box AND by a socket-proxy or Caddy play that runs on the EDGE host and
publishes it. A role default is invisible to that second play.

But the shape was wrong in two ways. The file had to be named in vars_files: by
30 plays - opt-in configuration that someone will eventually forget - and five
role defaults silently interpolated service_settings.*, so bitcoin_knots,
fulcrum, datum_gateway and mempool were not self-contained: using any of them
without that one vars_file entry broke it.

Each port now lives in host_vars/<owning box>/main.yml:

  host_vars/knots_box_local/main.yml    bitcoin_p2p_port, datum_gateway_api_port,
                                        datum_gateway_stratum_port
  host_vars/fulcrum_box_local/main.yml  fulcrum_ssl_port
  host_vars/mempool_box_local/main.yml  mempool_frontend_port

host_vars auto-loads and outranks role defaults, so the owning role picks the
value up with no vars_files at all, and the edge play reads the same single
definition as hostvars['<host>'].<name>. The role defaults keep the protocol
standard (8333, 50002, ...) so each role still works standalone, with the live
deployment's value in host_vars winning.

Also fixed a fourth copy of an inventory identity: the mempool Caddy play had
"mempool-box:{{ ... }}" hardcoded in the upstream. It now derives the host from
hostvars['mempool_box_local'].ansible_host, so inventory is the only place any
box's name is written down.

services_config.yml is deleted, with 25 more vars_files entries across 19
playbooks. Between this and the previous commit, 87 vars_files entries are gone
and every variable in the repo now comes from group_vars/all, host_vars,
inventory, a role default, or that service's own *_vars.yml.

Verification: an edge-host probe resolves all eight ports and hostnames to
byte-identical values to the ones services_config.yml used to supply. Each
owning host resolves its own port through host_vars. All 37 playbooks'
--list-tasks output is unchanged. The four edge plays that consume these values
all check-diff changed=0 - the socket-proxy and Caddy units on vipy are
byte-identical, which is the direct proof the rewiring landed on the same
values. fulcrum and datum-gateway check-diff exactly as before (ok=28/changed=1
and ok=15/changed=1, both the known timer re-arm).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 21:02:57 +02:00
..
defaults ansible: move the cross-host ports to host_vars, delete services_config.yml 2026-09-13 21:02:57 +02:00
handlers bitcoin-knots: convert to a role, de-Uptime-Kuma the health check 2026-09-13 18:13:40 +02:00
tasks bitcoin-knots: convert to a role, de-Uptime-Kuma the health check 2026-09-13 18:13:40 +02:00
templates bitcoin-knots: convert to a role, de-Uptime-Kuma the health check 2026-09-13 18:13:40 +02:00
README.md bitcoin-knots: convert to a role, de-Uptime-Kuma the health check 2026-09-13 18:13:40 +02:00

bitcoin_knots

Builds Bitcoin Knots from source with PGP + SHA256 verification of the release tarball, runs it as a full node on knots-box, and keeps a health check on a systemd timer. The second play in the calling playbook publishes the P2P port from the edge host via socket_proxy.

Converted from deploy_bitcoin_knots_playbook.yml (892 lines) under Plan 6. The playbook is now 40 lines.

The build is guarded; the chain is never touched

build.yml is 32 tasks, every one carrying when: not bitcoind_binary_exists.stat.exists. On a host that already has the binary the whole download / verify / 30-60 minute compile skips — including the two state: absent deletions, which target /opt/bitcoin-knots/source and the extracted build directory.

The chain lives elsewhere and nothing here touches it:

bitcoin_knots_dir /opt/bitcoin-knots — build tree, safe to delete
bitcoin_data_dir /var/lib/bitcoin — config, logs, wallets
bitcoin_large_data_dir /mnt/knots_data~875 GB of blockchain

The signature-verification tasks are the security control of this role. They are copied verbatim; do not "simplify" them.

⚠ This node is half of the mining setup

bitcoin.conf carries a DATUM Gateway section that was hand-added on the node and was missing from the playbook's template:

blockmaxsize=3985000
blockmaxweight=3985000
blocknotify=killall -USR1 datum_gateway
maxmempool=1000
blockreconstructionextratxn=1000000

blocknotify is how datum_gateway learns a new block landed. Running the old playbook would have deleted all of it, and solo mining would have carried on grinding against a stale template — a silent failure that costs money rather than raising an error. The template now carries it behind bitcoin_datum_gateway_enabled.

bitcoin-knots and datum-gateway are one system, not two services. Changing either config means thinking about both.

The restart handler, and why exactness matters now

The hand-written Restart bitcoind handler carried when: uptime_kuma_enabled | default(false), so it had been inert since the decommissioning: bitcoin.conf and the systemd unit both notify it and neither could restart anything. A config change applied to disk, reported success, and never took effect.

It is ungated here — which raises the bar for the template. Any residual difference between the template and the live file, down to a trailing newline, means the task reports changed and restarts a Bitcoin node on every run. It took four rounds of --check --diff to reach changed=0: the DATUM section, an explanatory comment that was rendering into the deployed file (now a {# #} Jinja comment), a # Pruning (optional) comment the live file had, and one trailing blank line.

dbcache

Computed as 90% of RAM unless bitcoin_dbcache_mb_override is set. The live node was hand-tuned to 200 MB; the calculation produces 3528. As with fulcrum, set_fact outranks role defaults, so the calculation honours the override — a value pinned only in defaults/ is silently ignored.

Monitoring: one variable, no product knowledge

The check tests bitcoind's RPC and records the answer in its exit code, which systemd keeps: systemctl is-failed bitcoin-knots-healthcheck.service. Set healthcheck_push_url to report anywhere that accepts an HTTP ping.

The timer had last fired 2026-08-09 while still reporting active and enabled — the same OnBootSec + OnUnitActiveSec dead chain as fulcrum, where nothing re-arms it if the service does not run in a given boot. The role runs the check once after enabling, which both smoke-tests it and supplies the reference the timer schedules from.