personal_infra/ansible/roles/bitcoin_knots
counterweight 954b683c71
ansible: delete the duplicated vars files, move globals to group_vars/all
Three files existed only as second copies of things group_vars/all already
auto-loads, and 34 playbooks named them in vars_files: - which outranks
group_vars, so the copies won. The day someone edited one and not the other,
those plays would silently keep the stale value. infra_vars.yml was already
drifting: group_vars/all/main.yml had grown age_backup_recipient and
backup_pull_public_key that it lacked.

  infra_vars.yml          - a strict subset of group_vars/all/main.yml
  infra_secrets.yml       - decrypts byte-identical to group_vars/all/vault.yml
  infra_secrets.yml.example - documented Uptime Kuma credentials as the reason
                              the file exists, which stopped being true

Deleted, along with 62 vars_files entries across 34 playbooks (12 of which
named ../../group_vars/all/main.yml directly - same defect, a vars_files entry
duplicating an auto-loaded file at higher precedence than the file itself).

Checked before touching anything: infra_secrets.yml was listed LAST in 10 plays,
after services_config.yml, so removal would flip precedence if the two shared a
key. They share none, and neither does services_config.yml with
group_vars/all/main.yml, so the removal is provably inert.

services_config.yml was the last one standing. It held four unrelated things:

  caddy_sites_dir            - an identical copy of roles/caddy_site/defaults/.
                               Deleted; the role default is now the only one.
  *.tailscale_hostname (x3)  - a THIRD copy of each box's identity, which
                               inventory.ini already holds as ansible_host.
                               Deleted. Edge plays now read
                               hostvars['<host>'].ansible_host - verified an
                               edge play resolves that with nothing loaded and
                               the other host in no play. Three copies of one
                               name is how bitcoin_rpc_host ended up labelled
                               "knots_box" while pointing at fulcrum-box.
  subdomains, ntfy topic,    - genuinely global: their readers span managed,
  headscale namespace          monitoring, vpn_control and edge, so no single
                               group covers them. Moved to group_vars/all/main.yml
                               where they auto-load. The ntfy_topic and
                               headscale_namespace indirection through
                               service_settings collapses to the global name.
  the four cross-host ports  - the only entries with a real justification.
                               Left in place; they move in the next commit.

Also dead, all Uptime Kuma residue or duplication:
  phoenixd_monitor_name, forgejo_runner healthcheck_timeout_seconds/retries,
  fulcrum_tailscale_hostname, and bitcoin_knots_version - the last being a
  v-prefixed copy of bitcoin_knots_version_short that nothing read, two
  hand-maintained copies of one version string.

Corrected a false comment: services_config.yml claimed the uptime_kuma subdomain
"no longer resolves to anything". It resolves to 164.92.239.72 and answers HTTP
302, and 11 playbooks still template it. Same wrong premise as PLAN_3.

Verification: all 37 playbooks' --list-tasks output is byte-identical before and
after. A probe resolving all 22 values services_config.yml used to supply returns
21 identical and one intended deletion (caddy_sites_dir, now role-only - confirmed
the role still resolves it: "Ensure Caddy sites-enabled directory exists" comes
back ok against the real path). memos check-diff identical before and after.
Syntax passes on every playbook.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 20:58:46 +02:00
..
defaults ansible: delete the duplicated vars files, move globals to group_vars/all 2026-09-13 20:58:46 +02:00
handlers bitcoin-knots: convert to a role, de-Uptime-Kuma the health check 2026-09-13 18:13:40 +02:00
tasks bitcoin-knots: convert to a role, de-Uptime-Kuma the health check 2026-09-13 18:13:40 +02:00
templates bitcoin-knots: convert to a role, de-Uptime-Kuma the health check 2026-09-13 18:13:40 +02:00
README.md bitcoin-knots: convert to a role, de-Uptime-Kuma the health check 2026-09-13 18:13:40 +02:00

bitcoin_knots

Builds Bitcoin Knots from source with PGP + SHA256 verification of the release tarball, runs it as a full node on knots-box, and keeps a health check on a systemd timer. The second play in the calling playbook publishes the P2P port from the edge host via socket_proxy.

Converted from deploy_bitcoin_knots_playbook.yml (892 lines) under Plan 6. The playbook is now 40 lines.

The build is guarded; the chain is never touched

build.yml is 32 tasks, every one carrying when: not bitcoind_binary_exists.stat.exists. On a host that already has the binary the whole download / verify / 30-60 minute compile skips — including the two state: absent deletions, which target /opt/bitcoin-knots/source and the extracted build directory.

The chain lives elsewhere and nothing here touches it:

bitcoin_knots_dir /opt/bitcoin-knots — build tree, safe to delete
bitcoin_data_dir /var/lib/bitcoin — config, logs, wallets
bitcoin_large_data_dir /mnt/knots_data~875 GB of blockchain

The signature-verification tasks are the security control of this role. They are copied verbatim; do not "simplify" them.

⚠ This node is half of the mining setup

bitcoin.conf carries a DATUM Gateway section that was hand-added on the node and was missing from the playbook's template:

blockmaxsize=3985000
blockmaxweight=3985000
blocknotify=killall -USR1 datum_gateway
maxmempool=1000
blockreconstructionextratxn=1000000

blocknotify is how datum_gateway learns a new block landed. Running the old playbook would have deleted all of it, and solo mining would have carried on grinding against a stale template — a silent failure that costs money rather than raising an error. The template now carries it behind bitcoin_datum_gateway_enabled.

bitcoin-knots and datum-gateway are one system, not two services. Changing either config means thinking about both.

The restart handler, and why exactness matters now

The hand-written Restart bitcoind handler carried when: uptime_kuma_enabled | default(false), so it had been inert since the decommissioning: bitcoin.conf and the systemd unit both notify it and neither could restart anything. A config change applied to disk, reported success, and never took effect.

It is ungated here — which raises the bar for the template. Any residual difference between the template and the live file, down to a trailing newline, means the task reports changed and restarts a Bitcoin node on every run. It took four rounds of --check --diff to reach changed=0: the DATUM section, an explanatory comment that was rendering into the deployed file (now a {# #} Jinja comment), a # Pruning (optional) comment the live file had, and one trailing blank line.

dbcache

Computed as 90% of RAM unless bitcoin_dbcache_mb_override is set. The live node was hand-tuned to 200 MB; the calculation produces 3528. As with fulcrum, set_fact outranks role defaults, so the calculation honours the override — a value pinned only in defaults/ is silently ignored.

Monitoring: one variable, no product knowledge

The check tests bitcoind's RPC and records the answer in its exit code, which systemd keeps: systemctl is-failed bitcoin-knots-healthcheck.service. Set healthcheck_push_url to report anywhere that accepts an HTTP ping.

The timer had last fired 2026-08-09 while still reporting active and enabled — the same OnBootSec + OnUnitActiveSec dead chain as fulcrum, where nothing re-arms it if the service does not run in a given boot. The role runs the check once after enabling, which both smoke-tests it and supplies the reference the timer schedules from.