personal_infra/ansible/roles/bitcoin_knots/defaults/main.yml

76 lines
3.2 KiB
YAML
Raw Normal View History

bitcoin-knots: convert to a role, de-Uptime-Kuma the health check 892-line playbook becomes 40 lines plus a role with install/build/configure/ service/healthcheck phases, five templates and one handler. bitcoin_knots_vars.yml is deleted; its content is the role's defaults. Verified after a real run: bitcoind still active since 2026-08-19 (NO restart), chain at 966844 blocks / 875 GB, DATUM config intact, dbcache still 200, health check timer firing again. changed=3, all health-check. vipy changed=0. ⚠ THE BIG ONE: the playbook would have deleted the mining integration. bitcoin.conf on the node carries a section that was hand-added and was missing from the template entirely: blockmaxsize=3985000 blockmaxweight=3985000 blocknotify=killall -USR1 datum_gateway maxmempool=1000 blockreconstructionextratxn=1000000 blocknotify is how datum_gateway learns a new block landed. Running the old playbook would have stripped all of it and solo mining would have carried on against a stale template - a silent failure that costs money rather than raising an error. Also dbcache 200 -> 3528 (hand-tuned down; the calculation wants 90% of RAM) and logging moved off the file. All now reconciled, dbcache behind bitcoin_dbcache_mb_override. bitcoin-knots and datum-gateway are ONE SYSTEM. Noted in the README. AND MY OWN FIX MADE IT MORE DANGEROUS. The `Restart bitcoind` handler was guarded by uptime_kuma_enabled, so it had been inert: bitcoin.conf and the systemd unit both notify it and neither could restart anything - a config change applied to disk, reported success, and never took effect. Ungating that is right, but it converts "wrong config sitting inertly on disk" into "node restarted onto a config that breaks mining". The ungating had to land WITH the template reconciliation, not before it. It also raises the bar permanently: any residual template/live difference now restarts a Bitcoin node on every run. Four rounds of --check --diff to reach changed=0 - the DATUM section, an explanatory comment that was rendering into the deployed config (now a {# #} Jinja comment), a "# Pruning (optional)" comment the live file had, and one trailing blank line. The build path is 32 tasks all guarded by `not bitcoind_binary_exists.stat.exists`, so a converged host skips the 30-60 minute compile and both `state: absent` deletions. Those target /opt/bitcoin-knots/{source,bitcoin-<version>}; the chain is in /mnt/knots_data and is never touched. Signature-verification tasks copied verbatim. The health check timer had last fired 2026-08-09 while reporting active/enabled - same OnBootSec + OnUnitActiveSec dead chain as fulcrum. The role runs the check once after enabling to supply the reference the timer schedules from. Ownership parity checked mechanically against `git show HEAD:`, keyed by TASK NAME rather than path - keying by path gave a false positive, because bitcoin_knots_source_dir is created with ownership and later removed with state: absent, so whichever task comes last wins and that differs between one file and five. 13/13 match. bitcoin_p2p_port and the tailscale hostname moved to services_config.yml for the socket-proxy play on the edge host. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 18:13:40 +02:00
# Bitcoin Knots Configuration Variables
# Version - REQUIRED: Specify exact version/tag to build
ansible: delete the duplicated vars files, move globals to group_vars/all Three files existed only as second copies of things group_vars/all already auto-loads, and 34 playbooks named them in vars_files: - which outranks group_vars, so the copies won. The day someone edited one and not the other, those plays would silently keep the stale value. infra_vars.yml was already drifting: group_vars/all/main.yml had grown age_backup_recipient and backup_pull_public_key that it lacked. infra_vars.yml - a strict subset of group_vars/all/main.yml infra_secrets.yml - decrypts byte-identical to group_vars/all/vault.yml infra_secrets.yml.example - documented Uptime Kuma credentials as the reason the file exists, which stopped being true Deleted, along with 62 vars_files entries across 34 playbooks (12 of which named ../../group_vars/all/main.yml directly - same defect, a vars_files entry duplicating an auto-loaded file at higher precedence than the file itself). Checked before touching anything: infra_secrets.yml was listed LAST in 10 plays, after services_config.yml, so removal would flip precedence if the two shared a key. They share none, and neither does services_config.yml with group_vars/all/main.yml, so the removal is provably inert. services_config.yml was the last one standing. It held four unrelated things: caddy_sites_dir - an identical copy of roles/caddy_site/defaults/. Deleted; the role default is now the only one. *.tailscale_hostname (x3) - a THIRD copy of each box's identity, which inventory.ini already holds as ansible_host. Deleted. Edge plays now read hostvars['<host>'].ansible_host - verified an edge play resolves that with nothing loaded and the other host in no play. Three copies of one name is how bitcoin_rpc_host ended up labelled "knots_box" while pointing at fulcrum-box. subdomains, ntfy topic, - genuinely global: their readers span managed, headscale namespace monitoring, vpn_control and edge, so no single group covers them. Moved to group_vars/all/main.yml where they auto-load. The ntfy_topic and headscale_namespace indirection through service_settings collapses to the global name. the four cross-host ports - the only entries with a real justification. Left in place; they move in the next commit. Also dead, all Uptime Kuma residue or duplication: phoenixd_monitor_name, forgejo_runner healthcheck_timeout_seconds/retries, fulcrum_tailscale_hostname, and bitcoin_knots_version - the last being a v-prefixed copy of bitcoin_knots_version_short that nothing read, two hand-maintained copies of one version string. Corrected a false comment: services_config.yml claimed the uptime_kuma subdomain "no longer resolves to anything". It resolves to 164.92.239.72 and answers HTTP 302, and 11 playbooks still template it. Same wrong premise as PLAN_3. Verification: all 37 playbooks' --list-tasks output is byte-identical before and after. A probe resolving all 22 values services_config.yml used to supply returns 21 identical and one intended deletion (caddy_sites_dir, now role-only - confirmed the role still resolves it: "Ensure Caddy sites-enabled directory exists" comes back ok against the real path). memos check-diff identical before and after. Syntax passes on every playbook. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 20:58:46 +02:00
# The only version string. There used to be a second, v-prefixed copy
# (bitcoin_knots_version) that nothing read - two hand-maintained copies of one
# fact, with nothing keeping them in step.
bitcoin_knots_version_short: "29.2.knots20251110"
bitcoin-knots: convert to a role, de-Uptime-Kuma the health check 892-line playbook becomes 40 lines plus a role with install/build/configure/ service/healthcheck phases, five templates and one handler. bitcoin_knots_vars.yml is deleted; its content is the role's defaults. Verified after a real run: bitcoind still active since 2026-08-19 (NO restart), chain at 966844 blocks / 875 GB, DATUM config intact, dbcache still 200, health check timer firing again. changed=3, all health-check. vipy changed=0. ⚠ THE BIG ONE: the playbook would have deleted the mining integration. bitcoin.conf on the node carries a section that was hand-added and was missing from the template entirely: blockmaxsize=3985000 blockmaxweight=3985000 blocknotify=killall -USR1 datum_gateway maxmempool=1000 blockreconstructionextratxn=1000000 blocknotify is how datum_gateway learns a new block landed. Running the old playbook would have stripped all of it and solo mining would have carried on against a stale template - a silent failure that costs money rather than raising an error. Also dbcache 200 -> 3528 (hand-tuned down; the calculation wants 90% of RAM) and logging moved off the file. All now reconciled, dbcache behind bitcoin_dbcache_mb_override. bitcoin-knots and datum-gateway are ONE SYSTEM. Noted in the README. AND MY OWN FIX MADE IT MORE DANGEROUS. The `Restart bitcoind` handler was guarded by uptime_kuma_enabled, so it had been inert: bitcoin.conf and the systemd unit both notify it and neither could restart anything - a config change applied to disk, reported success, and never took effect. Ungating that is right, but it converts "wrong config sitting inertly on disk" into "node restarted onto a config that breaks mining". The ungating had to land WITH the template reconciliation, not before it. It also raises the bar permanently: any residual template/live difference now restarts a Bitcoin node on every run. Four rounds of --check --diff to reach changed=0 - the DATUM section, an explanatory comment that was rendering into the deployed config (now a {# #} Jinja comment), a "# Pruning (optional)" comment the live file had, and one trailing blank line. The build path is 32 tasks all guarded by `not bitcoind_binary_exists.stat.exists`, so a converged host skips the 30-60 minute compile and both `state: absent` deletions. Those target /opt/bitcoin-knots/{source,bitcoin-<version>}; the chain is in /mnt/knots_data and is never touched. Signature-verification tasks copied verbatim. The health check timer had last fired 2026-08-09 while reporting active/enabled - same OnBootSec + OnUnitActiveSec dead chain as fulcrum. The role runs the check once after enabling to supply the reference the timer schedules from. Ownership parity checked mechanically against `git show HEAD:`, keyed by TASK NAME rather than path - keying by path gave a false positive, because bitcoin_knots_source_dir is created with ownership and later removed with state: absent, so whichever task comes last wins and that differs between one file and five. 13/13 match. bitcoin_p2p_port and the tailscale hostname moved to services_config.yml for the socket-proxy play on the edge host. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 18:13:40 +02:00
# Directories
bitcoin_knots_dir: /opt/bitcoin-knots
bitcoin_knots_source_dir: "{{ bitcoin_knots_dir }}/source"
bitcoin_data_dir: /var/lib/bitcoin # Standard location for config, logs, wallets
bitcoin_large_data_dir: /mnt/knots_data # Custom location for blockchain data (blocks, chainstate)
bitcoin_conf_dir: /etc/bitcoin
# Network
bitcoin_rpc_port: 8332
ansible: move the cross-host ports to host_vars, delete services_config.yml The four ports were the only entries in services_config.yml with a real justification: each is read twice, by the role that deploys the service on its own box AND by a socket-proxy or Caddy play that runs on the EDGE host and publishes it. A role default is invisible to that second play. But the shape was wrong in two ways. The file had to be named in vars_files: by 30 plays - opt-in configuration that someone will eventually forget - and five role defaults silently interpolated service_settings.*, so bitcoin_knots, fulcrum, datum_gateway and mempool were not self-contained: using any of them without that one vars_file entry broke it. Each port now lives in host_vars/<owning box>/main.yml: host_vars/knots_box_local/main.yml bitcoin_p2p_port, datum_gateway_api_port, datum_gateway_stratum_port host_vars/fulcrum_box_local/main.yml fulcrum_ssl_port host_vars/mempool_box_local/main.yml mempool_frontend_port host_vars auto-loads and outranks role defaults, so the owning role picks the value up with no vars_files at all, and the edge play reads the same single definition as hostvars['<host>'].<name>. The role defaults keep the protocol standard (8333, 50002, ...) so each role still works standalone, with the live deployment's value in host_vars winning. Also fixed a fourth copy of an inventory identity: the mempool Caddy play had "mempool-box:{{ ... }}" hardcoded in the upstream. It now derives the host from hostvars['mempool_box_local'].ansible_host, so inventory is the only place any box's name is written down. services_config.yml is deleted, with 25 more vars_files entries across 19 playbooks. Between this and the previous commit, 87 vars_files entries are gone and every variable in the repo now comes from group_vars/all, host_vars, inventory, a role default, or that service's own *_vars.yml. Verification: an edge-host probe resolves all eight ports and hostnames to byte-identical values to the ones services_config.yml used to supply. Each owning host resolves its own port through host_vars. All 37 playbooks' --list-tasks output is unchanged. The four edge plays that consume these values all check-diff changed=0 - the socket-proxy and Caddy units on vipy are byte-identical, which is the direct proof the rewiring landed on the same values. fulcrum and datum-gateway check-diff exactly as before (ok=28/changed=1 and ok=15/changed=1, both the known timer re-arm). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 21:02:57 +02:00
# The edge host's socket-proxy/Caddy play needs this too, and a role default is
# invisible outside this role. The authoritative value for the live deployment is
# in host_vars/knots_box_local/main.yml, which outranks this; the value here is the
# protocol standard, so the role still works standalone.
bitcoin_p2p_port: 8333
bitcoin-knots: convert to a role, de-Uptime-Kuma the health check 892-line playbook becomes 40 lines plus a role with install/build/configure/ service/healthcheck phases, five templates and one handler. bitcoin_knots_vars.yml is deleted; its content is the role's defaults. Verified after a real run: bitcoind still active since 2026-08-19 (NO restart), chain at 966844 blocks / 875 GB, DATUM config intact, dbcache still 200, health check timer firing again. changed=3, all health-check. vipy changed=0. ⚠ THE BIG ONE: the playbook would have deleted the mining integration. bitcoin.conf on the node carries a section that was hand-added and was missing from the template entirely: blockmaxsize=3985000 blockmaxweight=3985000 blocknotify=killall -USR1 datum_gateway maxmempool=1000 blockreconstructionextratxn=1000000 blocknotify is how datum_gateway learns a new block landed. Running the old playbook would have stripped all of it and solo mining would have carried on against a stale template - a silent failure that costs money rather than raising an error. Also dbcache 200 -> 3528 (hand-tuned down; the calculation wants 90% of RAM) and logging moved off the file. All now reconciled, dbcache behind bitcoin_dbcache_mb_override. bitcoin-knots and datum-gateway are ONE SYSTEM. Noted in the README. AND MY OWN FIX MADE IT MORE DANGEROUS. The `Restart bitcoind` handler was guarded by uptime_kuma_enabled, so it had been inert: bitcoin.conf and the systemd unit both notify it and neither could restart anything - a config change applied to disk, reported success, and never took effect. Ungating that is right, but it converts "wrong config sitting inertly on disk" into "node restarted onto a config that breaks mining". The ungating had to land WITH the template reconciliation, not before it. It also raises the bar permanently: any residual template/live difference now restarts a Bitcoin node on every run. Four rounds of --check --diff to reach changed=0 - the DATUM section, an explanatory comment that was rendering into the deployed config (now a {# #} Jinja comment), a "# Pruning (optional)" comment the live file had, and one trailing blank line. The build path is 32 tasks all guarded by `not bitcoind_binary_exists.stat.exists`, so a converged host skips the 30-60 minute compile and both `state: absent` deletions. Those target /opt/bitcoin-knots/{source,bitcoin-<version>}; the chain is in /mnt/knots_data and is never touched. Signature-verification tasks copied verbatim. The health check timer had last fired 2026-08-09 while reporting active/enabled - same OnBootSec + OnUnitActiveSec dead chain as fulcrum. The role runs the check once after enabling to supply the reference the timer schedules from. Ownership parity checked mechanically against `git show HEAD:`, keyed by TASK NAME rather than path - keying by path gave a false positive, because bitcoin_knots_source_dir is created with ownership and later removed with state: absent, so whichever task comes last wins and that differs between one file and five. 13/13 match. bitcoin_p2p_port and the tailscale hostname moved to services_config.yml for the socket-proxy play on the edge host. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 18:13:40 +02:00
bitcoin_rpc_bind: "0.0.0.0"
# Build options
bitcoin_build_jobs: 4 # Parallel build jobs (-j flag), adjust based on CPU cores
bitcoin_build_prefix: /usr/local
# Configuration options
bitcoin_enable_txindex: true # Set to true if transaction index needed (REQUIRED for Electrum servers like Electrs/ElectrumX)
bitcoin_max_connections: 125
# dbcache will be calculated as 90% of host RAM automatically in playbook
# ZMQ Configuration
bitcoin_zmq_enabled: true
bitcoin_zmq_bind: "tcp://0.0.0.0"
bitcoin_zmq_port_rawblock: 28332
bitcoin_zmq_port_rawtx: 28333
bitcoin_zmq_port_hashblock: 28334
bitcoin_zmq_port_hashtx: 28335
# Service user
bitcoin_user: bitcoin
bitcoin_group: bitcoin
# --- Health check ----------------------------------------------------------
# Checks bitcoind RPC and records the answer in its exit code, which systemd
# keeps: `systemctl is-failed bitcoin-knots-healthcheck.service`.
#
# WHERE TO REPORT HEALTH — the one place to plug in monitoring. Empty means
# check, exit honestly, report nowhere. Any endpoint accepting an HTTP ping
# works; nothing here is specific to a monitoring product.
healthcheck_push_url: ""
# --- Logging ----------------------------------------------------------------
# The live node logs to a file. Set to "" to use printtoconsole=1 (journald).
bitcoin_logfile: "{{ bitcoin_data_dir }}/debug.log"
# --- dbcache ----------------------------------------------------------------
# Computed as 90% of RAM unless this is set. The live node was hand-tuned to
# 200 MB; the calculation would have produced 3528. As with fulcrum, note that
# set_fact outranks role defaults, so the CALCULATION has to honour this - a
# value pinned only in defaults/ is silently ignored.
bitcoin_dbcache_mb_override: 200
# --- DATUM Gateway ----------------------------------------------------------
# This node feeds block templates to datum_gateway on knots-box. These settings
# were hand-added to bitcoin.conf and were missing from the template, so a
# playbook run would have removed them and broken the mining setup.
bitcoin_datum_gateway_enabled: true
bitcoin_blockmaxsize: 3985000
bitcoin_blockmaxweight: 3985000
bitcoin_blocknotify: "killall -USR1 datum_gateway"
bitcoin_maxmempool: 1000
bitcoin_blockreconstructionextratxn: 1000000