fulcrum: convert to a role, de-Uptime-Kuma the health check
685-line playbook becomes 33 lines plus a 426-line role (install/service/healthcheck phases, six templates, one handler). fulcrum_vars.yml is deleted; its content is the role's defaults. Verified: fulcrum untouched - active since 2026-07-29 (no restart), 192G datadir, height 966842, bitcoind and db_mem unchanged on disk. Second run changed=1 (the arming run, changed_when: false aside). vipy changed=0. THREE PRE-EXISTING LANDMINES the check-mode diff caught, any of which a faithful extraction would have detonated: - bitcoin_rpc_host was "192.168.1.140", commented "IP of knots_box_local". But .140 is fulcrum-box ITSELF; knots-box is .135. The DHCP leases had reshuffled - the fifth instance of this same disease in this estate. The live config had been hand-corrected to knots-box; running the playbook would have reverted it and pointed Fulcrum at itself. Now addressed by Tailscale name. - The `Restart fulcrum` handler was guarded by uptime_kuma_enabled, so the three tasks that notify it (SSL cert, fulcrum.conf, systemd unit) could not restart anything. A config change applied to disk, reported success, and silently never took effect. That is worse than the other banner casualties: it makes the deployment itself lie. Ungated. - db_mem was about to go 2048 -> 4448 (75% of 5931MB RAM), leaving ~1.4GB for the OS and Fulcrum's non-cache memory. The live value had been hand-tuned down. fulcrum_db_mem_mb_override pins it. Note set_fact outranks role defaults, so the calculation itself has to honour the override. MY OWN ERROR, third instance: retyping `copy:` as `template:` lost `owner:` on the banner and on fulcrum.conf. Rather than keep catching these by eye, every managed path's owner/group/mode is now compared against `git show HEAD:` mechanically - 12/12 match. The health check timer had not fired since 2026-02-17 while reporting `active` and `enabled`. It is OnBootSec + OnUnitActiveSec with no OnCalendar: OnBootSec elapses once, and OnUnitActiveSec needs the SERVICE to have run this boot to have anything to schedule from. Restarting the timer does not supply that; running the service does, so the role now runs the check once after enabling. Also dropped `Requires=fulcrum.service` from the timer - on a timer that means "stop watching when the watched thing stops". Diagnostic note: NextElapseUSecRealtime is always empty for a monotonic timer, so it reads as broken even when healthy. I misread it once and wrongly called the timer dead. Use NextElapseUSecMonotonic or systemctl list-timers. fulcrum_ssl_port and fulcrum_tailscale_hostname moved to services_config.yml - the socket-proxy play on the edge host needs them and a role default cannot reach a second play. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
parent
356139290f
commit
e83191c029
15 changed files with 516 additions and 666 deletions
78
ansible/roles/fulcrum/defaults/main.yml
Normal file
78
ansible/roles/fulcrum/defaults/main.yml
Normal file
|
|
@ -0,0 +1,78 @@
|
|||
# Fulcrum Configuration Variables
|
||||
|
||||
# Version - Pinned to specific release
|
||||
fulcrum_version: "2.1.0" # Fulcrum version to install
|
||||
|
||||
# Directories
|
||||
fulcrum_db_dir: /mnt/fulcrum_data/fulcrum_db # Database directory (heavy data on special mount)
|
||||
fulcrum_config_dir: /etc/fulcrum # Config file location (standard OS path)
|
||||
fulcrum_lib_dir: /var/lib/fulcrum # Other data files (banner, etc.) on OS disk
|
||||
fulcrum_binary_path: /usr/local/bin/Fulcrum
|
||||
|
||||
# Network - Bitcoin RPC connection
|
||||
# Bitcoin Knots is on a different host (knots_box_local)
|
||||
# Using RPC user/password authentication (credentials from infra_secrets.yml)
|
||||
# Addressed by Tailscale name, never a LAN IP. This was
|
||||
# bitcoin_rpc_host: "192.168.1.140" # IP of knots_box_local
|
||||
# but .140 is fulcrum-box ITSELF - knots-box is .135. The DHCP leases had
|
||||
# reshuffled (the same drift that transposed the inventory), so running this
|
||||
# playbook would have pointed Fulcrum at itself and broken indexing. The live
|
||||
# config had already been hand-corrected to knots-box; this makes the repo
|
||||
# agree with it.
|
||||
bitcoin_rpc_host: "knots-box"
|
||||
bitcoin_rpc_port: 8332 # Bitcoin Knots RPC port
|
||||
# Note: bitcoin_rpc_user and bitcoin_rpc_password are loaded from infra_secrets.yml
|
||||
|
||||
# Network - Fulcrum server
|
||||
fulcrum_tcp_port: 50001
|
||||
# Shared with the socket-proxy play on the edge host, so it lives in
|
||||
# services_config.yml rather than only here.
|
||||
fulcrum_ssl_port: "{{ service_settings.fulcrum.ssl_port }}"
|
||||
# Binding address for Fulcrum TCP/SSL server:
|
||||
# - "127.0.0.1" = localhost only (use when Caddy is on the same box)
|
||||
# - "0.0.0.0" = all interfaces (use when Caddy is on a different box)
|
||||
# - Specific IP = bind to specific network interface
|
||||
fulcrum_tcp_bind: "0.0.0.0" # Default: localhost (change to "0.0.0.0" if Caddy is on different box)
|
||||
fulcrum_ssl_bind: "0.0.0.0" # Binding address for SSL port
|
||||
# If Caddy is on a different box, set this to the IP address that Caddy will use to connect
|
||||
|
||||
# SSL/TLS Configuration
|
||||
fulcrum_ssl_enabled: true
|
||||
fulcrum_ssl_cert_path: "{{ fulcrum_config_dir }}/fulcrum.crt"
|
||||
fulcrum_ssl_key_path: "{{ fulcrum_config_dir }}/fulcrum.key"
|
||||
fulcrum_ssl_cert_days: 3650 # 10 years validity for self-signed cert
|
||||
|
||||
# Port forwarding configuration (for public access via VPS)
|
||||
fulcrum_tailscale_hostname: "{{ service_settings.fulcrum.tailscale_hostname }}"
|
||||
|
||||
# Performance
|
||||
# db_mem will be calculated as 75% of available RAM automatically in playbook
|
||||
# db_mem is computed as this share of RAM unless fulcrum_db_mem_mb is set
|
||||
# explicitly. On a 5931 MB host 75% is 4448 MB, which leaves ~1.4 GB for the
|
||||
# OS and for Fulcrum's non-cache memory; the live config had been hand-tuned
|
||||
# down to 2048 and that setting is respected below.
|
||||
fulcrum_db_mem_percent: 0.75 # 75% of RAM for database cache
|
||||
|
||||
# Configuration options
|
||||
fulcrum_anon_logs: true # Anonymize client IPs and TxIDs in logs
|
||||
fulcrum_peering: false # Disable peering with other Fulcrum servers
|
||||
fulcrum_zmq_allow_hashtx: true # Allow ZMQ hashtx notifications
|
||||
|
||||
# Service user
|
||||
fulcrum_user: fulcrum
|
||||
fulcrum_group: fulcrum
|
||||
|
||||
|
||||
# --- Health check -----------------------------------------------------------
|
||||
# Checks the Electrum TCP port and records the answer in its exit code, which
|
||||
# systemd keeps: `systemctl is-failed fulcrum-healthcheck.service`.
|
||||
#
|
||||
# WHERE TO REPORT HEALTH — the one place to plug in monitoring. Empty means
|
||||
# check, exit honestly, report nowhere. Any endpoint accepting an HTTP ping
|
||||
# works; nothing here is specific to a monitoring product.
|
||||
healthcheck_push_url: ""
|
||||
|
||||
# Explicit db_mem in MB. When set it wins over fulcrum_db_mem_percent; empty
|
||||
# means compute from RAM. Set here because the live host had been hand-tuned to
|
||||
# 2048 and a silent jump to 4448 is not something a refactor should do.
|
||||
fulcrum_db_mem_mb_override: 2048
|
||||
Loading…
Add table
Add a link
Reference in a new issue