2025-12-14 18:52:36 +01:00
|
|
|
# Fulcrum Configuration Variables
|
|
|
|
|
|
|
|
|
|
# Version - Pinned to specific release
|
|
|
|
|
fulcrum_version: "2.1.0" # Fulcrum version to install
|
|
|
|
|
|
|
|
|
|
# Directories
|
|
|
|
|
fulcrum_db_dir: /mnt/fulcrum_data/fulcrum_db # Database directory (heavy data on special mount)
|
|
|
|
|
fulcrum_config_dir: /etc/fulcrum # Config file location (standard OS path)
|
|
|
|
|
fulcrum_lib_dir: /var/lib/fulcrum # Other data files (banner, etc.) on OS disk
|
|
|
|
|
fulcrum_binary_path: /usr/local/bin/Fulcrum
|
|
|
|
|
|
|
|
|
|
# Network - Bitcoin RPC connection
|
|
|
|
|
# Bitcoin Knots is on a different host (knots_box_local)
|
ansible: delete the duplicated vars files, move globals to group_vars/all
Three files existed only as second copies of things group_vars/all already
auto-loads, and 34 playbooks named them in vars_files: - which outranks
group_vars, so the copies won. The day someone edited one and not the other,
those plays would silently keep the stale value. infra_vars.yml was already
drifting: group_vars/all/main.yml had grown age_backup_recipient and
backup_pull_public_key that it lacked.
infra_vars.yml - a strict subset of group_vars/all/main.yml
infra_secrets.yml - decrypts byte-identical to group_vars/all/vault.yml
infra_secrets.yml.example - documented Uptime Kuma credentials as the reason
the file exists, which stopped being true
Deleted, along with 62 vars_files entries across 34 playbooks (12 of which
named ../../group_vars/all/main.yml directly - same defect, a vars_files entry
duplicating an auto-loaded file at higher precedence than the file itself).
Checked before touching anything: infra_secrets.yml was listed LAST in 10 plays,
after services_config.yml, so removal would flip precedence if the two shared a
key. They share none, and neither does services_config.yml with
group_vars/all/main.yml, so the removal is provably inert.
services_config.yml was the last one standing. It held four unrelated things:
caddy_sites_dir - an identical copy of roles/caddy_site/defaults/.
Deleted; the role default is now the only one.
*.tailscale_hostname (x3) - a THIRD copy of each box's identity, which
inventory.ini already holds as ansible_host.
Deleted. Edge plays now read
hostvars['<host>'].ansible_host - verified an
edge play resolves that with nothing loaded and
the other host in no play. Three copies of one
name is how bitcoin_rpc_host ended up labelled
"knots_box" while pointing at fulcrum-box.
subdomains, ntfy topic, - genuinely global: their readers span managed,
headscale namespace monitoring, vpn_control and edge, so no single
group covers them. Moved to group_vars/all/main.yml
where they auto-load. The ntfy_topic and
headscale_namespace indirection through
service_settings collapses to the global name.
the four cross-host ports - the only entries with a real justification.
Left in place; they move in the next commit.
Also dead, all Uptime Kuma residue or duplication:
phoenixd_monitor_name, forgejo_runner healthcheck_timeout_seconds/retries,
fulcrum_tailscale_hostname, and bitcoin_knots_version - the last being a
v-prefixed copy of bitcoin_knots_version_short that nothing read, two
hand-maintained copies of one version string.
Corrected a false comment: services_config.yml claimed the uptime_kuma subdomain
"no longer resolves to anything". It resolves to 164.92.239.72 and answers HTTP
302, and 11 playbooks still template it. Same wrong premise as PLAN_3.
Verification: all 37 playbooks' --list-tasks output is byte-identical before and
after. A probe resolving all 22 values services_config.yml used to supply returns
21 identical and one intended deletion (caddy_sites_dir, now role-only - confirmed
the role still resolves it: "Ensure Caddy sites-enabled directory exists" comes
back ok against the real path). memos check-diff identical before and after.
Syntax passes on every playbook.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 20:58:46 +02:00
|
|
|
# Using RPC user/password authentication (credentials from group_vars/all/vault.yml)
|
fulcrum: convert to a role, de-Uptime-Kuma the health check
685-line playbook becomes 33 lines plus a 426-line role
(install/service/healthcheck phases, six templates, one handler).
fulcrum_vars.yml is deleted; its content is the role's defaults.
Verified: fulcrum untouched - active since 2026-07-29 (no restart), 192G datadir,
height 966842, bitcoind and db_mem unchanged on disk. Second run changed=1 (the
arming run, changed_when: false aside). vipy changed=0.
THREE PRE-EXISTING LANDMINES the check-mode diff caught, any of which a faithful
extraction would have detonated:
- bitcoin_rpc_host was "192.168.1.140", commented "IP of knots_box_local". But
.140 is fulcrum-box ITSELF; knots-box is .135. The DHCP leases had reshuffled -
the fifth instance of this same disease in this estate. The live config had
been hand-corrected to knots-box; running the playbook would have reverted it
and pointed Fulcrum at itself. Now addressed by Tailscale name.
- The `Restart fulcrum` handler was guarded by uptime_kuma_enabled, so the three
tasks that notify it (SSL cert, fulcrum.conf, systemd unit) could not restart
anything. A config change applied to disk, reported success, and silently never
took effect. That is worse than the other banner casualties: it makes the
deployment itself lie. Ungated.
- db_mem was about to go 2048 -> 4448 (75% of 5931MB RAM), leaving ~1.4GB for the
OS and Fulcrum's non-cache memory. The live value had been hand-tuned down.
fulcrum_db_mem_mb_override pins it. Note set_fact outranks role defaults, so
the calculation itself has to honour the override.
MY OWN ERROR, third instance: retyping `copy:` as `template:` lost `owner:` on
the banner and on fulcrum.conf. Rather than keep catching these by eye, every
managed path's owner/group/mode is now compared against `git show HEAD:`
mechanically - 12/12 match.
The health check timer had not fired since 2026-02-17 while reporting `active`
and `enabled`. It is OnBootSec + OnUnitActiveSec with no OnCalendar: OnBootSec
elapses once, and OnUnitActiveSec needs the SERVICE to have run this boot to have
anything to schedule from. Restarting the timer does not supply that; running the
service does, so the role now runs the check once after enabling. Also dropped
`Requires=fulcrum.service` from the timer - on a timer that means "stop watching
when the watched thing stops".
Diagnostic note: NextElapseUSecRealtime is always empty for a monotonic timer, so
it reads as broken even when healthy. I misread it once and wrongly called the
timer dead. Use NextElapseUSecMonotonic or systemctl list-timers.
fulcrum_ssl_port and fulcrum_tailscale_hostname moved to services_config.yml -
the socket-proxy play on the edge host needs them and a role default cannot reach
a second play.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 18:02:00 +02:00
|
|
|
# Addressed by Tailscale name, never a LAN IP. This was
|
|
|
|
|
# bitcoin_rpc_host: "192.168.1.140" # IP of knots_box_local
|
|
|
|
|
# but .140 is fulcrum-box ITSELF - knots-box is .135. The DHCP leases had
|
|
|
|
|
# reshuffled (the same drift that transposed the inventory), so running this
|
|
|
|
|
# playbook would have pointed Fulcrum at itself and broken indexing. The live
|
|
|
|
|
# config had already been hand-corrected to knots-box; this makes the repo
|
|
|
|
|
# agree with it.
|
|
|
|
|
bitcoin_rpc_host: "knots-box"
|
2025-12-14 18:52:36 +01:00
|
|
|
bitcoin_rpc_port: 8332 # Bitcoin Knots RPC port
|
ansible: delete the duplicated vars files, move globals to group_vars/all
Three files existed only as second copies of things group_vars/all already
auto-loads, and 34 playbooks named them in vars_files: - which outranks
group_vars, so the copies won. The day someone edited one and not the other,
those plays would silently keep the stale value. infra_vars.yml was already
drifting: group_vars/all/main.yml had grown age_backup_recipient and
backup_pull_public_key that it lacked.
infra_vars.yml - a strict subset of group_vars/all/main.yml
infra_secrets.yml - decrypts byte-identical to group_vars/all/vault.yml
infra_secrets.yml.example - documented Uptime Kuma credentials as the reason
the file exists, which stopped being true
Deleted, along with 62 vars_files entries across 34 playbooks (12 of which
named ../../group_vars/all/main.yml directly - same defect, a vars_files entry
duplicating an auto-loaded file at higher precedence than the file itself).
Checked before touching anything: infra_secrets.yml was listed LAST in 10 plays,
after services_config.yml, so removal would flip precedence if the two shared a
key. They share none, and neither does services_config.yml with
group_vars/all/main.yml, so the removal is provably inert.
services_config.yml was the last one standing. It held four unrelated things:
caddy_sites_dir - an identical copy of roles/caddy_site/defaults/.
Deleted; the role default is now the only one.
*.tailscale_hostname (x3) - a THIRD copy of each box's identity, which
inventory.ini already holds as ansible_host.
Deleted. Edge plays now read
hostvars['<host>'].ansible_host - verified an
edge play resolves that with nothing loaded and
the other host in no play. Three copies of one
name is how bitcoin_rpc_host ended up labelled
"knots_box" while pointing at fulcrum-box.
subdomains, ntfy topic, - genuinely global: their readers span managed,
headscale namespace monitoring, vpn_control and edge, so no single
group covers them. Moved to group_vars/all/main.yml
where they auto-load. The ntfy_topic and
headscale_namespace indirection through
service_settings collapses to the global name.
the four cross-host ports - the only entries with a real justification.
Left in place; they move in the next commit.
Also dead, all Uptime Kuma residue or duplication:
phoenixd_monitor_name, forgejo_runner healthcheck_timeout_seconds/retries,
fulcrum_tailscale_hostname, and bitcoin_knots_version - the last being a
v-prefixed copy of bitcoin_knots_version_short that nothing read, two
hand-maintained copies of one version string.
Corrected a false comment: services_config.yml claimed the uptime_kuma subdomain
"no longer resolves to anything". It resolves to 164.92.239.72 and answers HTTP
302, and 11 playbooks still template it. Same wrong premise as PLAN_3.
Verification: all 37 playbooks' --list-tasks output is byte-identical before and
after. A probe resolving all 22 values services_config.yml used to supply returns
21 identical and one intended deletion (caddy_sites_dir, now role-only - confirmed
the role still resolves it: "Ensure Caddy sites-enabled directory exists" comes
back ok against the real path). memos check-diff identical before and after.
Syntax passes on every playbook.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 20:58:46 +02:00
|
|
|
# Note: bitcoin_rpc_user and bitcoin_rpc_password are loaded from group_vars/all/vault.yml
|
2025-12-14 18:52:36 +01:00
|
|
|
|
|
|
|
|
# Network - Fulcrum server
|
|
|
|
|
fulcrum_tcp_port: 50001
|
ansible: move the cross-host ports to host_vars, delete services_config.yml
The four ports were the only entries in services_config.yml with a real
justification: each is read twice, by the role that deploys the service on its
own box AND by a socket-proxy or Caddy play that runs on the EDGE host and
publishes it. A role default is invisible to that second play.
But the shape was wrong in two ways. The file had to be named in vars_files: by
30 plays - opt-in configuration that someone will eventually forget - and five
role defaults silently interpolated service_settings.*, so bitcoin_knots,
fulcrum, datum_gateway and mempool were not self-contained: using any of them
without that one vars_file entry broke it.
Each port now lives in host_vars/<owning box>/main.yml:
host_vars/knots_box_local/main.yml bitcoin_p2p_port, datum_gateway_api_port,
datum_gateway_stratum_port
host_vars/fulcrum_box_local/main.yml fulcrum_ssl_port
host_vars/mempool_box_local/main.yml mempool_frontend_port
host_vars auto-loads and outranks role defaults, so the owning role picks the
value up with no vars_files at all, and the edge play reads the same single
definition as hostvars['<host>'].<name>. The role defaults keep the protocol
standard (8333, 50002, ...) so each role still works standalone, with the live
deployment's value in host_vars winning.
Also fixed a fourth copy of an inventory identity: the mempool Caddy play had
"mempool-box:{{ ... }}" hardcoded in the upstream. It now derives the host from
hostvars['mempool_box_local'].ansible_host, so inventory is the only place any
box's name is written down.
services_config.yml is deleted, with 25 more vars_files entries across 19
playbooks. Between this and the previous commit, 87 vars_files entries are gone
and every variable in the repo now comes from group_vars/all, host_vars,
inventory, a role default, or that service's own *_vars.yml.
Verification: an edge-host probe resolves all eight ports and hostnames to
byte-identical values to the ones services_config.yml used to supply. Each
owning host resolves its own port through host_vars. All 37 playbooks'
--list-tasks output is unchanged. The four edge plays that consume these values
all check-diff changed=0 - the socket-proxy and Caddy units on vipy are
byte-identical, which is the direct proof the rewiring landed on the same
values. fulcrum and datum-gateway check-diff exactly as before (ok=28/changed=1
and ok=15/changed=1, both the known timer re-arm).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 21:02:57 +02:00
|
|
|
# The edge host's socket-proxy/Caddy play needs this too, and a role default is
|
|
|
|
|
# invisible outside this role. The authoritative value for the live deployment is
|
|
|
|
|
# in host_vars/fulcrum_box_local/main.yml, which outranks this; the value here is the
|
|
|
|
|
# protocol standard, so the role still works standalone.
|
|
|
|
|
fulcrum_ssl_port: 50002
|
2025-12-24 10:27:35 +01:00
|
|
|
# Binding address for Fulcrum TCP/SSL server:
|
2025-12-14 18:52:36 +01:00
|
|
|
# - "127.0.0.1" = localhost only (use when Caddy is on the same box)
|
|
|
|
|
# - "0.0.0.0" = all interfaces (use when Caddy is on a different box)
|
|
|
|
|
# - Specific IP = bind to specific network interface
|
|
|
|
|
fulcrum_tcp_bind: "0.0.0.0" # Default: localhost (change to "0.0.0.0" if Caddy is on different box)
|
2025-12-24 10:27:35 +01:00
|
|
|
fulcrum_ssl_bind: "0.0.0.0" # Binding address for SSL port
|
2025-12-14 18:52:36 +01:00
|
|
|
# If Caddy is on a different box, set this to the IP address that Caddy will use to connect
|
|
|
|
|
|
2025-12-24 10:27:35 +01:00
|
|
|
# SSL/TLS Configuration
|
|
|
|
|
fulcrum_ssl_enabled: true
|
|
|
|
|
fulcrum_ssl_cert_path: "{{ fulcrum_config_dir }}/fulcrum.crt"
|
|
|
|
|
fulcrum_ssl_key_path: "{{ fulcrum_config_dir }}/fulcrum.key"
|
|
|
|
|
fulcrum_ssl_cert_days: 3650 # 10 years validity for self-signed cert
|
|
|
|
|
|
|
|
|
|
|
2025-12-14 18:52:36 +01:00
|
|
|
# Performance
|
|
|
|
|
# db_mem will be calculated as 75% of available RAM automatically in playbook
|
fulcrum: convert to a role, de-Uptime-Kuma the health check
685-line playbook becomes 33 lines plus a 426-line role
(install/service/healthcheck phases, six templates, one handler).
fulcrum_vars.yml is deleted; its content is the role's defaults.
Verified: fulcrum untouched - active since 2026-07-29 (no restart), 192G datadir,
height 966842, bitcoind and db_mem unchanged on disk. Second run changed=1 (the
arming run, changed_when: false aside). vipy changed=0.
THREE PRE-EXISTING LANDMINES the check-mode diff caught, any of which a faithful
extraction would have detonated:
- bitcoin_rpc_host was "192.168.1.140", commented "IP of knots_box_local". But
.140 is fulcrum-box ITSELF; knots-box is .135. The DHCP leases had reshuffled -
the fifth instance of this same disease in this estate. The live config had
been hand-corrected to knots-box; running the playbook would have reverted it
and pointed Fulcrum at itself. Now addressed by Tailscale name.
- The `Restart fulcrum` handler was guarded by uptime_kuma_enabled, so the three
tasks that notify it (SSL cert, fulcrum.conf, systemd unit) could not restart
anything. A config change applied to disk, reported success, and silently never
took effect. That is worse than the other banner casualties: it makes the
deployment itself lie. Ungated.
- db_mem was about to go 2048 -> 4448 (75% of 5931MB RAM), leaving ~1.4GB for the
OS and Fulcrum's non-cache memory. The live value had been hand-tuned down.
fulcrum_db_mem_mb_override pins it. Note set_fact outranks role defaults, so
the calculation itself has to honour the override.
MY OWN ERROR, third instance: retyping `copy:` as `template:` lost `owner:` on
the banner and on fulcrum.conf. Rather than keep catching these by eye, every
managed path's owner/group/mode is now compared against `git show HEAD:`
mechanically - 12/12 match.
The health check timer had not fired since 2026-02-17 while reporting `active`
and `enabled`. It is OnBootSec + OnUnitActiveSec with no OnCalendar: OnBootSec
elapses once, and OnUnitActiveSec needs the SERVICE to have run this boot to have
anything to schedule from. Restarting the timer does not supply that; running the
service does, so the role now runs the check once after enabling. Also dropped
`Requires=fulcrum.service` from the timer - on a timer that means "stop watching
when the watched thing stops".
Diagnostic note: NextElapseUSecRealtime is always empty for a monotonic timer, so
it reads as broken even when healthy. I misread it once and wrongly called the
timer dead. Use NextElapseUSecMonotonic or systemctl list-timers.
fulcrum_ssl_port and fulcrum_tailscale_hostname moved to services_config.yml -
the socket-proxy play on the edge host needs them and a role default cannot reach
a second play.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 18:02:00 +02:00
|
|
|
# db_mem is computed as this share of RAM unless fulcrum_db_mem_mb is set
|
|
|
|
|
# explicitly. On a 5931 MB host 75% is 4448 MB, which leaves ~1.4 GB for the
|
|
|
|
|
# OS and for Fulcrum's non-cache memory; the live config had been hand-tuned
|
|
|
|
|
# down to 2048 and that setting is respected below.
|
2025-12-14 18:52:36 +01:00
|
|
|
fulcrum_db_mem_percent: 0.75 # 75% of RAM for database cache
|
|
|
|
|
|
|
|
|
|
# Configuration options
|
|
|
|
|
fulcrum_anon_logs: true # Anonymize client IPs and TxIDs in logs
|
|
|
|
|
fulcrum_peering: false # Disable peering with other Fulcrum servers
|
|
|
|
|
fulcrum_zmq_allow_hashtx: true # Allow ZMQ hashtx notifications
|
|
|
|
|
|
|
|
|
|
# Service user
|
|
|
|
|
fulcrum_user: fulcrum
|
|
|
|
|
fulcrum_group: fulcrum
|
|
|
|
|
|
fulcrum: convert to a role, de-Uptime-Kuma the health check
685-line playbook becomes 33 lines plus a 426-line role
(install/service/healthcheck phases, six templates, one handler).
fulcrum_vars.yml is deleted; its content is the role's defaults.
Verified: fulcrum untouched - active since 2026-07-29 (no restart), 192G datadir,
height 966842, bitcoind and db_mem unchanged on disk. Second run changed=1 (the
arming run, changed_when: false aside). vipy changed=0.
THREE PRE-EXISTING LANDMINES the check-mode diff caught, any of which a faithful
extraction would have detonated:
- bitcoin_rpc_host was "192.168.1.140", commented "IP of knots_box_local". But
.140 is fulcrum-box ITSELF; knots-box is .135. The DHCP leases had reshuffled -
the fifth instance of this same disease in this estate. The live config had
been hand-corrected to knots-box; running the playbook would have reverted it
and pointed Fulcrum at itself. Now addressed by Tailscale name.
- The `Restart fulcrum` handler was guarded by uptime_kuma_enabled, so the three
tasks that notify it (SSL cert, fulcrum.conf, systemd unit) could not restart
anything. A config change applied to disk, reported success, and silently never
took effect. That is worse than the other banner casualties: it makes the
deployment itself lie. Ungated.
- db_mem was about to go 2048 -> 4448 (75% of 5931MB RAM), leaving ~1.4GB for the
OS and Fulcrum's non-cache memory. The live value had been hand-tuned down.
fulcrum_db_mem_mb_override pins it. Note set_fact outranks role defaults, so
the calculation itself has to honour the override.
MY OWN ERROR, third instance: retyping `copy:` as `template:` lost `owner:` on
the banner and on fulcrum.conf. Rather than keep catching these by eye, every
managed path's owner/group/mode is now compared against `git show HEAD:`
mechanically - 12/12 match.
The health check timer had not fired since 2026-02-17 while reporting `active`
and `enabled`. It is OnBootSec + OnUnitActiveSec with no OnCalendar: OnBootSec
elapses once, and OnUnitActiveSec needs the SERVICE to have run this boot to have
anything to schedule from. Restarting the timer does not supply that; running the
service does, so the role now runs the check once after enabling. Also dropped
`Requires=fulcrum.service` from the timer - on a timer that means "stop watching
when the watched thing stops".
Diagnostic note: NextElapseUSecRealtime is always empty for a monotonic timer, so
it reads as broken even when healthy. I misread it once and wrongly called the
timer dead. Use NextElapseUSecMonotonic or systemctl list-timers.
fulcrum_ssl_port and fulcrum_tailscale_hostname moved to services_config.yml -
the socket-proxy play on the edge host needs them and a role default cannot reach
a second play.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 18:02:00 +02:00
|
|
|
|
|
|
|
|
# --- Health check -----------------------------------------------------------
|
|
|
|
|
# Checks the Electrum TCP port and records the answer in its exit code, which
|
|
|
|
|
# systemd keeps: `systemctl is-failed fulcrum-healthcheck.service`.
|
|
|
|
|
#
|
|
|
|
|
# WHERE TO REPORT HEALTH — the one place to plug in monitoring. Empty means
|
|
|
|
|
# check, exit honestly, report nowhere. Any endpoint accepting an HTTP ping
|
|
|
|
|
# works; nothing here is specific to a monitoring product.
|
|
|
|
|
healthcheck_push_url: ""
|
|
|
|
|
|
|
|
|
|
# Explicit db_mem in MB. When set it wins over fulcrum_db_mem_percent; empty
|
|
|
|
|
# means compute from RAM. Set here because the live host had been hand-tuned to
|
|
|
|
|
# 2048 and a silent jump to 4448 is not something a refactor should do.
|
|
|
|
|
fulcrum_db_mem_mb_override: 2048
|