685-line playbook becomes 33 lines plus a 426-line role (install/service/healthcheck phases, six templates, one handler). fulcrum_vars.yml is deleted; its content is the role's defaults. Verified: fulcrum untouched - active since 2026-07-29 (no restart), 192G datadir, height 966842, bitcoind and db_mem unchanged on disk. Second run changed=1 (the arming run, changed_when: false aside). vipy changed=0. THREE PRE-EXISTING LANDMINES the check-mode diff caught, any of which a faithful extraction would have detonated: - bitcoin_rpc_host was "192.168.1.140", commented "IP of knots_box_local". But .140 is fulcrum-box ITSELF; knots-box is .135. The DHCP leases had reshuffled - the fifth instance of this same disease in this estate. The live config had been hand-corrected to knots-box; running the playbook would have reverted it and pointed Fulcrum at itself. Now addressed by Tailscale name. - The `Restart fulcrum` handler was guarded by uptime_kuma_enabled, so the three tasks that notify it (SSL cert, fulcrum.conf, systemd unit) could not restart anything. A config change applied to disk, reported success, and silently never took effect. That is worse than the other banner casualties: it makes the deployment itself lie. Ungated. - db_mem was about to go 2048 -> 4448 (75% of 5931MB RAM), leaving ~1.4GB for the OS and Fulcrum's non-cache memory. The live value had been hand-tuned down. fulcrum_db_mem_mb_override pins it. Note set_fact outranks role defaults, so the calculation itself has to honour the override. MY OWN ERROR, third instance: retyping `copy:` as `template:` lost `owner:` on the banner and on fulcrum.conf. Rather than keep catching these by eye, every managed path's owner/group/mode is now compared against `git show HEAD:` mechanically - 12/12 match. The health check timer had not fired since 2026-02-17 while reporting `active` and `enabled`. It is OnBootSec + OnUnitActiveSec with no OnCalendar: OnBootSec elapses once, and OnUnitActiveSec needs the SERVICE to have run this boot to have anything to schedule from. Restarting the timer does not supply that; running the service does, so the role now runs the check once after enabling. Also dropped `Requires=fulcrum.service` from the timer - on a timer that means "stop watching when the watched thing stops". Diagnostic note: NextElapseUSecRealtime is always empty for a monotonic timer, so it reads as broken even when healthy. I misread it once and wrongly called the timer dead. Use NextElapseUSecMonotonic or systemctl list-timers. fulcrum_ssl_port and fulcrum_tailscale_hostname moved to services_config.yml - the socket-proxy play on the edge host needs them and a role default cannot reach a second play. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
62 lines
2.4 KiB
YAML
62 lines
2.4 KiB
YAML
---
|
|
# Everything here answers "is Fulcrum healthy" and records the answer. The
|
|
# Uptime Kuma specifics that used to follow — an embedded Python script creating
|
|
# monitors over the API, a /tmp credentials file, push-URL extraction and a
|
|
# systemd Environment= rewrite — are gone. Where it reports is now one variable,
|
|
# healthcheck_push_url. See the role README.
|
|
- name: Create Fulcrum health check script
|
|
ansible.builtin.template:
|
|
src: healthcheck.sh.j2
|
|
dest: /usr/local/bin/fulcrum-healthcheck-push.sh
|
|
owner: root
|
|
group: root
|
|
mode: '0755'
|
|
validate: "bash -n %s"
|
|
|
|
- name: Create systemd service for Fulcrum health check
|
|
ansible.builtin.template:
|
|
src: healthcheck.service.j2
|
|
dest: /etc/systemd/system/fulcrum-healthcheck.service
|
|
owner: root
|
|
group: root
|
|
mode: '0644'
|
|
|
|
- name: Create systemd timer for Fulcrum health check
|
|
ansible.builtin.template:
|
|
src: healthcheck.timer.j2
|
|
dest: /etc/systemd/system/fulcrum-healthcheck.timer
|
|
owner: root
|
|
group: root
|
|
mode: '0644'
|
|
|
|
- name: Reload systemd daemon for health check
|
|
systemd:
|
|
daemon_reload: yes
|
|
|
|
# state: restarted, not started. The hand-written timer had got itself stuck
|
|
# `active` with no next elapse and had not fired since 2026-02-17; `started` on
|
|
# an already-active timer is a no-op and would have left it stuck. Restarting
|
|
# re-arms it. See the note in healthcheck.timer.j2.
|
|
- name: Enable and restart the Fulcrum health check timer
|
|
systemd:
|
|
name: fulcrum-healthcheck.timer
|
|
enabled: yes
|
|
state: restarted
|
|
daemon_reload: yes
|
|
|
|
# Run the check once, which is both a smoke test and the thing that actually
|
|
# arms the timer.
|
|
#
|
|
# This timer is OnBootSec + OnUnitActiveSec with no OnCalendar. OnBootSec is
|
|
# monotonic and had long since elapsed; OnUnitActiveSec schedules relative to the
|
|
# SERVICE last being active, and the service had not run since 2026-02-17 — so
|
|
# there was no reference to schedule from and the timer sat `active` and
|
|
# `enabled` with NextElapseUSecMonotonic=infinity. Restarting the timer alone
|
|
# does not supply that reference; running the service does.
|
|
#
|
|
# (Diagnosing this is easy to get wrong: NextElapseUSecRealtime is always empty
|
|
# for a monotonic timer, so it looks broken even when it is fine. Read
|
|
# NextElapseUSecMonotonic, or just use `systemctl list-timers`.)
|
|
- name: Run the Fulcrum health check once to arm the timer
|
|
command: systemctl start fulcrum-healthcheck.service
|
|
changed_when: false
|