fulcrum: convert to a role, de-Uptime-Kuma the health check
685-line playbook becomes 33 lines plus a 426-line role (install/service/healthcheck phases, six templates, one handler). fulcrum_vars.yml is deleted; its content is the role's defaults. Verified: fulcrum untouched - active since 2026-07-29 (no restart), 192G datadir, height 966842, bitcoind and db_mem unchanged on disk. Second run changed=1 (the arming run, changed_when: false aside). vipy changed=0. THREE PRE-EXISTING LANDMINES the check-mode diff caught, any of which a faithful extraction would have detonated: - bitcoin_rpc_host was "192.168.1.140", commented "IP of knots_box_local". But .140 is fulcrum-box ITSELF; knots-box is .135. The DHCP leases had reshuffled - the fifth instance of this same disease in this estate. The live config had been hand-corrected to knots-box; running the playbook would have reverted it and pointed Fulcrum at itself. Now addressed by Tailscale name. - The `Restart fulcrum` handler was guarded by uptime_kuma_enabled, so the three tasks that notify it (SSL cert, fulcrum.conf, systemd unit) could not restart anything. A config change applied to disk, reported success, and silently never took effect. That is worse than the other banner casualties: it makes the deployment itself lie. Ungated. - db_mem was about to go 2048 -> 4448 (75% of 5931MB RAM), leaving ~1.4GB for the OS and Fulcrum's non-cache memory. The live value had been hand-tuned down. fulcrum_db_mem_mb_override pins it. Note set_fact outranks role defaults, so the calculation itself has to honour the override. MY OWN ERROR, third instance: retyping `copy:` as `template:` lost `owner:` on the banner and on fulcrum.conf. Rather than keep catching these by eye, every managed path's owner/group/mode is now compared against `git show HEAD:` mechanically - 12/12 match. The health check timer had not fired since 2026-02-17 while reporting `active` and `enabled`. It is OnBootSec + OnUnitActiveSec with no OnCalendar: OnBootSec elapses once, and OnUnitActiveSec needs the SERVICE to have run this boot to have anything to schedule from. Restarting the timer does not supply that; running the service does, so the role now runs the check once after enabling. Also dropped `Requires=fulcrum.service` from the timer - on a timer that means "stop watching when the watched thing stops". Diagnostic note: NextElapseUSecRealtime is always empty for a monotonic timer, so it reads as broken even when healthy. I misread it once and wrongly called the timer dead. Use NextElapseUSecMonotonic or systemctl list-timers. fulcrum_ssl_port and fulcrum_tailscale_hostname moved to services_config.yml - the socket-proxy play on the edge host needs them and a role default cannot reach a second play. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
parent
356139290f
commit
e83191c029
15 changed files with 516 additions and 666 deletions
3
ansible/roles/fulcrum/templates/banner.txt.j2
Normal file
3
ansible/roles/fulcrum/templates/banner.txt.j2
Normal file
|
|
@ -0,0 +1,3 @@
|
|||
counterinfra
|
||||
|
||||
PER ASPERA AD ASTRA
|
||||
29
ansible/roles/fulcrum/templates/fulcrum.conf.j2
Normal file
29
ansible/roles/fulcrum/templates/fulcrum.conf.j2
Normal file
|
|
@ -0,0 +1,29 @@
|
|||
# Fulcrum Configuration
|
||||
# Generated by Ansible
|
||||
|
||||
# Bitcoin Core/Knots RPC settings
|
||||
bitcoind = {{ bitcoin_rpc_host }}:{{ bitcoin_rpc_port }}
|
||||
rpcuser = {{ bitcoin_rpc_user }}
|
||||
rpcpassword = {{ bitcoin_rpc_password }}
|
||||
|
||||
# Fulcrum server general settings
|
||||
datadir = {{ fulcrum_db_dir }}
|
||||
tcp = {{ fulcrum_tcp_bind }}:{{ fulcrum_tcp_port }}
|
||||
peering = {{ 'true' if fulcrum_peering else 'false' }}
|
||||
zmq_allow_hashtx = {{ 'true' if fulcrum_zmq_allow_hashtx else 'false' }}
|
||||
|
||||
# SSL/TLS Configuration
|
||||
{% if fulcrum_ssl_enabled | default(false) %}
|
||||
ssl = {{ fulcrum_ssl_bind }}:{{ fulcrum_ssl_port }}
|
||||
cert = {{ fulcrum_ssl_cert_path }}
|
||||
key = {{ fulcrum_ssl_key_path }}
|
||||
{% endif %}
|
||||
|
||||
# Anonymize client IP addresses and TxIDs in logs
|
||||
anon_logs = {{ 'true' if fulcrum_anon_logs else 'false' }}
|
||||
|
||||
# Max RocksDB Memory in MiB
|
||||
db_mem = {{ fulcrum_db_mem_mb }}.0
|
||||
|
||||
# Banner
|
||||
banner = {{ fulcrum_lib_dir }}/fulcrum-banner.txt
|
||||
24
ansible/roles/fulcrum/templates/fulcrum.service.j2
Normal file
24
ansible/roles/fulcrum/templates/fulcrum.service.j2
Normal file
|
|
@ -0,0 +1,24 @@
|
|||
# MiniBolt: systemd unit for Fulcrum
|
||||
# /etc/systemd/system/fulcrum.service
|
||||
|
||||
[Unit]
|
||||
Description=Fulcrum
|
||||
After=network.target
|
||||
|
||||
StartLimitBurst=2
|
||||
StartLimitIntervalSec=20
|
||||
|
||||
[Service]
|
||||
ExecStart={{ fulcrum_binary_path }} {{ fulcrum_config_dir }}/fulcrum.conf
|
||||
|
||||
User={{ fulcrum_user }}
|
||||
Group={{ fulcrum_group }}
|
||||
|
||||
# Process management
|
||||
####################
|
||||
Type=simple
|
||||
KillSignal=SIGINT
|
||||
TimeoutStopSec=300
|
||||
|
||||
[Install]
|
||||
WantedBy=multi-user.target
|
||||
14
ansible/roles/fulcrum/templates/healthcheck.service.j2
Normal file
14
ansible/roles/fulcrum/templates/healthcheck.service.j2
Normal file
|
|
@ -0,0 +1,14 @@
|
|||
[Unit]
|
||||
Description=Fulcrum Health Check
|
||||
After=network.target fulcrum.service
|
||||
|
||||
[Service]
|
||||
Type=oneshot
|
||||
User=root
|
||||
ExecStart=/usr/local/bin/fulcrum-healthcheck-push.sh
|
||||
Environment=HEALTHCHECK_PUSH_URL={{ healthcheck_push_url }}
|
||||
StandardOutput=journal
|
||||
StandardError=journal
|
||||
|
||||
[Install]
|
||||
WantedBy=multi-user.target
|
||||
34
ansible/roles/fulcrum/templates/healthcheck.sh.j2
Normal file
34
ansible/roles/fulcrum/templates/healthcheck.sh.j2
Normal file
|
|
@ -0,0 +1,34 @@
|
|||
#!/bin/bash
|
||||
# Fulcrum health check — managed by Ansible (roles/fulcrum)
|
||||
#
|
||||
# Checks that Fulcrum's Electrum TCP port is accepting connections, and records
|
||||
# the answer in the exit code, which systemd keeps:
|
||||
# systemctl is-failed fulcrum-healthcheck.service
|
||||
# That is a complete answer with no monitoring system involved. Reporting
|
||||
# elsewhere is optional and generic — set healthcheck_push_url.
|
||||
|
||||
FULCRUM_HOST="{{ fulcrum_tcp_bind }}"
|
||||
FULCRUM_PORT={{ fulcrum_tcp_port }}
|
||||
PUSH_URL="${HEALTHCHECK_PUSH_URL:-}"
|
||||
|
||||
check_fulcrum() {
|
||||
timeout 5 bash -c "echo > /dev/tcp/${FULCRUM_HOST}/${FULCRUM_PORT}" 2>/dev/null
|
||||
}
|
||||
|
||||
report() {
|
||||
local status=$1 msg=$2
|
||||
# No push URL is normal, not an error: the exit code below is still a
|
||||
# complete answer for anything reading unit state.
|
||||
[ -n "$PUSH_URL" ] || return 0
|
||||
curl -s --max-time 10 --retry 2 -o /dev/null \
|
||||
"${PUSH_URL}?status=${status}&msg=${msg// /%20}&ping=" || true
|
||||
}
|
||||
|
||||
if check_fulcrum; then
|
||||
report "up" "OK"
|
||||
exit 0
|
||||
else
|
||||
echo "Fulcrum TCP port ${FULCRUM_PORT} not responding"
|
||||
report "down" "Fulcrum TCP port not responding"
|
||||
exit 1
|
||||
fi
|
||||
17
ansible/roles/fulcrum/templates/healthcheck.timer.j2
Normal file
17
ansible/roles/fulcrum/templates/healthcheck.timer.j2
Normal file
|
|
@ -0,0 +1,17 @@
|
|||
[Unit]
|
||||
Description=Fulcrum Health Check Timer
|
||||
# NOTE: this deliberately does NOT carry `Requires=fulcrum.service`, which the
|
||||
# hand-written unit had. Requires on a timer means the timer is stopped when the
|
||||
# required unit stops — i.e. "if the thing I am watching goes down, stop
|
||||
# watching it", which is backwards for a health check and leaves nothing to
|
||||
# re-arm the timer when the service returns. The live timer had been `active`
|
||||
# and `enabled` with NextElapseUSecMonotonic=infinity and a last trigger of
|
||||
# 2026-02-17: seven months with no health check and no outward sign of it.
|
||||
|
||||
[Timer]
|
||||
OnBootSec=1min
|
||||
OnUnitActiveSec=1min
|
||||
Persistent=true
|
||||
|
||||
[Install]
|
||||
WantedBy=timers.target
|
||||
Loading…
Add table
Add a link
Reference in a new issue