bitcoin-knots: convert to a role, de-Uptime-Kuma the health check

892-line playbook becomes 40 lines plus a role with install/build/configure/
service/healthcheck phases, five templates and one handler. bitcoin_knots_vars.yml
is deleted; its content is the role's defaults.

Verified after a real run: bitcoind still active since 2026-08-19 (NO restart),
chain at 966844 blocks / 875 GB, DATUM config intact, dbcache still 200, health
check timer firing again. changed=3, all health-check. vipy changed=0.

⚠ THE BIG ONE: the playbook would have deleted the mining integration.

bitcoin.conf on the node carries a section that was hand-added and was missing
from the template entirely:

    blockmaxsize=3985000
    blockmaxweight=3985000
    blocknotify=killall -USR1 datum_gateway
    maxmempool=1000
    blockreconstructionextratxn=1000000

blocknotify is how datum_gateway learns a new block landed. Running the old
playbook would have stripped all of it and solo mining would have carried on
against a stale template - a silent failure that costs money rather than raising
an error. Also dbcache 200 -> 3528 (hand-tuned down; the calculation wants 90% of
RAM) and logging moved off the file. All now reconciled, dbcache behind
bitcoin_dbcache_mb_override.

bitcoin-knots and datum-gateway are ONE SYSTEM. Noted in the README.

AND MY OWN FIX MADE IT MORE DANGEROUS. The `Restart bitcoind` handler was guarded
by uptime_kuma_enabled, so it had been inert: bitcoin.conf and the systemd unit
both notify it and neither could restart anything - a config change applied to
disk, reported success, and never took effect. Ungating that is right, but it
converts "wrong config sitting inertly on disk" into "node restarted onto a
config that breaks mining". The ungating had to land WITH the template
reconciliation, not before it.

It also raises the bar permanently: any residual template/live difference now
restarts a Bitcoin node on every run. Four rounds of --check --diff to reach
changed=0 - the DATUM section, an explanatory comment that was rendering into the
deployed config (now a {# #} Jinja comment), a "# Pruning (optional)" comment the
live file had, and one trailing blank line.

The build path is 32 tasks all guarded by `not bitcoind_binary_exists.stat.exists`,
so a converged host skips the 30-60 minute compile and both `state: absent`
deletions. Those target /opt/bitcoin-knots/{source,bitcoin-<version>}; the chain
is in /mnt/knots_data and is never touched. Signature-verification tasks copied
verbatim.

The health check timer had last fired 2026-08-09 while reporting active/enabled -
same OnBootSec + OnUnitActiveSec dead chain as fulcrum. The role runs the check
once after enabling to supply the reference the timer schedules from.

Ownership parity checked mechanically against `git show HEAD:`, keyed by TASK
NAME rather than path - keying by path gave a false positive, because
bitcoin_knots_source_dir is created with ownership and later removed with
state: absent, so whichever task comes last wins and that differs between one
file and five. 13/13 match.

bitcoin_p2p_port and the tailscale hostname moved to services_config.yml for the
socket-proxy play on the edge host.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
counterweight 2026-09-13 18:13:40 +02:00
parent e83191c029
commit d26dc78b3c
Signed by: counterweight
GPG key ID: 883EDBAA726BD96C
19 changed files with 1163 additions and 1250 deletions

View file

@ -0,0 +1,67 @@
# Bitcoin Knots Configuration
# Generated by Ansible
# Data directory (blockchain storage)
datadir={{ bitcoin_large_data_dir }}
# RPC Configuration
server=1
rpcuser={{ bitcoin_rpc_user }}
rpcpassword={{ bitcoin_rpc_password }}
rpcbind={{ bitcoin_rpc_bind }}
rpcport={{ bitcoin_rpc_port }}
rpcallowip=0.0.0.0/0
# Network Configuration
listen=1
port={{ bitcoin_p2p_port }}
maxconnections={{ bitcoin_max_connections }}
# Performance
dbcache={{ bitcoin_dbcache_mb }}
# Transaction Index (optional)
{% if bitcoin_enable_txindex %}
txindex=1
{% endif %}
{# The live node carries this comment and the template never produced it, so a
run would have silently deleted it. Harmless in itself, but matching it keeps
this task at `ok` - which means any future `changed` here is a real signal
rather than known noise. #}
# Pruning (optional)
# Logging
logtimestamps=1
{% if bitcoin_logfile %}
logfile={{ bitcoin_logfile }}
{% else %}
printtoconsole=1
{% endif %}
# ZMQ Configuration
{% if bitcoin_zmq_enabled | default(false) %}
zmqpubrawblock={{ bitcoin_zmq_bind }}:{{ bitcoin_zmq_port_rawblock }}
zmqpubrawtx={{ bitcoin_zmq_bind }}:{{ bitcoin_zmq_port_rawtx }}
zmqpubhashblock={{ bitcoin_zmq_bind }}:{{ bitcoin_zmq_port_hashblock }}
zmqpubhashtx={{ bitcoin_zmq_bind }}:{{ bitcoin_zmq_port_hashtx }}
{% endif %}
# Security
disablewallet=1
{% if bitcoin_datum_gateway_enabled %}
{# These were hand-added on the node and were NOT in this template, so running
the playbook would have stripped them. blocknotify is how datum_gateway
learns a new block landed; without it solo mining keeps grinding on a stale
template - a silent failure that costs money rather than raising an error.
Kept as a Jinja comment so the explanation stays in the repo and out of the
deployed config. #}
# Specific for DATUM gateway
blockmaxsize={{ bitcoin_blockmaxsize }}
blockmaxweight={{ bitcoin_blockmaxweight }}
blocknotify={{ bitcoin_blocknotify }}
maxmempool={{ bitcoin_maxmempool }}
blockreconstructionextratxn={{ bitcoin_blockreconstructionextratxn }}
{% endif %}

View file

@ -0,0 +1,17 @@
[Unit]
Description=Bitcoin Knots daemon
After=network.target
[Service]
Type=simple
User={{ bitcoin_user }}
Group={{ bitcoin_group }}
ExecStart={{ bitcoin_build_prefix }}/bin/bitcoind -conf={{ bitcoin_conf_dir }}/bitcoin.conf
Restart=always
RestartSec=10
TimeoutStopSec=600
StandardOutput=journal
StandardError=journal
[Install]
WantedBy=multi-user.target

View file

@ -0,0 +1,14 @@
[Unit]
Description=Bitcoin Knots Health Check
After=network.target bitcoind.service
[Service]
Type=oneshot
User=root
ExecStart=/usr/local/bin/bitcoin-knots-healthcheck-push.sh
Environment=HEALTHCHECK_PUSH_URL={{ healthcheck_push_url }}
StandardOutput=journal
StandardError=journal
[Install]
WantedBy=multi-user.target

View file

@ -0,0 +1,62 @@
#!/bin/bash
# Bitcoin Knots health check — managed by Ansible (roles/bitcoin_knots)
#
# The exit code is the answer and systemd keeps it:
# systemctl is-failed bitcoin-knots-healthcheck.service
# Reporting anywhere else is optional and generic.
#
#
RPC_HOST="{{ bitcoin_rpc_bind }}"
RPC_PORT={{ bitcoin_rpc_port }}
RPC_USER="{{ bitcoin_rpc_user }}"
RPC_PASSWORD="{{ bitcoin_rpc_password }}"
PUSH_URL="${HEALTHCHECK_PUSH_URL:-}"
# Check if bitcoind RPC is responding
check_bitcoind() {
local response
response=$(curl -s --max-time 30 \
--user "${RPC_USER}:${RPC_PASSWORD}" \
--data-binary '{"jsonrpc":"1.0","id":"healthcheck","method":"getblockchaininfo","params":[]}' \
--header 'Content-Type: application/json' \
"http://${RPC_HOST}:${RPC_PORT}" 2>&1)
if [ $? -eq 0 ]; then
# Check if response contains a non-null error
# Successful responses have "error": null, failures have "error": {...}
if echo "$response" | grep -q '"error":null\|"error": null'; then
return 0
else
return 1
fi
else
return 1
fi
}
report() {
local status=$1
local msg=$2
# No push URL is normal, not an error: the exit code below is still a
# complete answer for anything reading unit state.
[ -n "$PUSH_URL" ] || return 0
# URL encode spaces in message
local encoded_msg="${msg// /%20}"
if ! curl -s --max-time 10 --retry 2 -o /dev/null \
"${PUSH_URL}?status=${status}&msg=${encoded_msg}&ping="; then
return 1
fi
}
# Main health check
if check_bitcoind; then
report "up" "OK"
exit 0
else
report "down" "bitcoind RPC not responding"
exit 1
fi

View file

@ -0,0 +1,11 @@
[Unit]
Description=Bitcoin Knots Health Check Timer
Requires=bitcoind.service
[Timer]
OnBootSec=1min
OnUnitActiveSec=1min
Persistent=true
[Install]
WantedBy=timers.target