personal_infra/ansible/roles/mempool/tasks/healthcheck.yml

59 lines
1.8 KiB
YAML
Raw Normal View History

mempool: convert to a role, de-Uptime-Kuma the health checks 745-line playbook becomes 37 lines (the role, plus the Caddy play for the edge host) and a 408-line role with docker/deploy/healthcheck phases and six templates. mempool_vars.yml is deleted; its content is the role's defaults. Three health checks are kept, not collapsed: Mempool is three moving parts and knowing which one is down is the point. Each has its own script, unit, timer and push_url, driven by a mempool_healthchecks list. The Uptime Kuma specifics are gone - the embedded Python creating monitors over the API, the /tmp credentials file, the push-URL file read back and parsed, three Environment= rewrites - and the three live push URLs are preserved from the vault, so reporting is unchanged. `Enable and start health check timers` and `Display deployment status` were both guarded by uptime_kuma_enabled despite being deployment tasks. Third service in a row with that pattern: the deprecation banner was applied to contiguous blocks, so anything sitting near the push plumbing was disabled with it. Ungated. TWO OWNERSHIP PROBLEMS, different in kind: - MINE: I wrote `owner: root` on docker-compose.yml where the original says `owner: "{{ ansible_user }}"`. A straight violation of extract-mechanically- change-nothing, caught only by reading the check-mode diff line by line. Reverted to match the original. - PRE-EXISTING, and dangerous: the playbook declared `owner: "{{ ansible_user }}"` (1000) on the MariaDB data directory, which the container owns as uid 999. Confirmed against `git show HEAD:` before concluding it was not mine. It had drifted since the containers were created and went unnoticed because the playbook had not been run since. This was not academic. The first real run pulled a newer mariadb:10.11 and recreated mempool-db; with the chown still in place MariaDB would have come back to a data directory it could not write. The role now ensures the directory exists and leaves ownership to the container. Verified after the run: /opt/mempool/mysql is still 999:999 and all three containers are healthy. This is a deliberate behaviour change, not part of the extraction. It is in this commit rather than a follow-up because the faithful version was never safe to run, so there was no intermediate state worth recording as verified. mempool_frontend_port moved to services_config.yml: two hosts need it (this role deploys the frontend, the Caddy play proxies to it from the edge host) and a role default is invisible to the second play. caddy_site's parameter assert caught this loudly - "'mempool_frontend_port' is undefined" - rather than silently. Verified: check-mode diff clean apart from unavoidable check-mode artifacts; first run ok=24 changed=5, zero failures; second run changed=2 - the two bare `command:` tasks (pull, compose up) that have no changed_when and always report changed. That is the idempotent floor. All three health checks report ExecMainStatus 0 with their push URLs intact. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-12 18:49:11 +02:00
---
# Three checks, one per moving part. The Uptime Kuma specifics that used to
# follow — an embedded Python script creating monitors over the API, a /tmp
# credentials file, a push-URL file read back and parsed, and three systemd
# Environment= rewrites — are gone. Where each reports is now hc.push_url.
- name: Create Mempool health check scripts
ansible.builtin.template:
src: "healthcheck-{{ hc.name }}.sh.j2"
dest: "/usr/local/bin/mempool-{{ hc.name }}-healthcheck-push.sh"
owner: root
group: root
mode: '0755'
validate: "bash -n %s"
loop: "{{ mempool_healthchecks }}"
loop_control:
loop_var: hc
label: "{{ hc.name }}"
- name: Create systemd services for health checks
ansible.builtin.template:
src: healthcheck.service.j2
dest: "/etc/systemd/system/mempool-{{ hc.name }}-healthcheck.service"
owner: root
group: root
uptime kuma: remove every live reference, repoint the probes to Gatus Nothing in the repo pushes to, authenticates against, or is gated by Uptime Kuma any more. ── The sixth instance of the banner bug ──────────────────────────────────── memos had `Restart memos` guarded by `uptime_kuma_enabled`, because the deprecation banner was placed immediately above it and swept it in. It is a HANDLER, so every memos config change since 2026-09-11 applied to disk and silently never restarted the service. Ungated. That is the same failure found in forgejo-runner's self-assert, phoenixd's timer enable, mempool's three timer enables, fulcrum's restart handler and bitcoind's restart handler. Every guard was read and asked "monitoring or deployment?" before being deleted, which is the only reason this was caught. ── What was removed ──────────────────────────────────────────────────────── 30 uptime_kuma_enabled guards across 7 unconverted service playbooks, and the 29 Kuma monitor-creation tasks they gated (embedded Python that drove the Kuma API, temp credential files, cleanup) 7 dead uptime_kuma_api_url definitions 7 stale DEPRECATED banners uptime_kuma_enabled and subdomains.uptime_kuma from group_vars/all healthcheck_push_urls from the vault - 30 push tokens services/ntfy/setup_ntfy_uptime_kuma_notification.yml -> archive/ The explanatory comments in the six converted roles are KEPT on purpose. They record why a handler is ungated, and deleting the explanation invites someone to helpfully re-add the guard. ── The probes moved rather than died ─────────────────────────────────────── Eight per-service health checks were still pushing to Kuma. They are not superseded by infra/401: that answers "is the unit running", these answer "does the service actually respond" - an RPC call to bitcoind, a TCP connect to Fulcrum's Electrum port, an HTTP fetch from Mempool's backend. A process can be perfectly `active` and useless. So they were repointed, not deleted. Gatus external endpoints take a POST with a bearer token and success=true|false where Kuma took a GET with ?status=up, so report() now maps up/down to true/false internally and no call site changed. Registered by infra/403 as the `probe` group, one token per host. Two bugs fixed while in there: * forgejo-runner's check only ever reported SUCCESS - it exited before pushing when the runner was down, so a failure was invisible until the heartbeat window expired. Reporting the failure is the entire point of a check. * All six healthcheck .service units were mode 0644 and now carry a bearer token. They are 0600. Verified: 91 endpoints, 91 UP, 0 DOWN. Every probe triggered by hand and confirmed arriving. Zero Kuma URLs left in the vault, zero live references in any playbook or role. Still standing, deliberately: the Kuma container on watchtower, its Caddy vhost, and the uptime.contrapeso.xyz DNS record. Turning the service off is a separate decision from removing the code that talked to it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 10:30:28 +02:00
mode: "0600"
mempool: convert to a role, de-Uptime-Kuma the health checks 745-line playbook becomes 37 lines (the role, plus the Caddy play for the edge host) and a 408-line role with docker/deploy/healthcheck phases and six templates. mempool_vars.yml is deleted; its content is the role's defaults. Three health checks are kept, not collapsed: Mempool is three moving parts and knowing which one is down is the point. Each has its own script, unit, timer and push_url, driven by a mempool_healthchecks list. The Uptime Kuma specifics are gone - the embedded Python creating monitors over the API, the /tmp credentials file, the push-URL file read back and parsed, three Environment= rewrites - and the three live push URLs are preserved from the vault, so reporting is unchanged. `Enable and start health check timers` and `Display deployment status` were both guarded by uptime_kuma_enabled despite being deployment tasks. Third service in a row with that pattern: the deprecation banner was applied to contiguous blocks, so anything sitting near the push plumbing was disabled with it. Ungated. TWO OWNERSHIP PROBLEMS, different in kind: - MINE: I wrote `owner: root` on docker-compose.yml where the original says `owner: "{{ ansible_user }}"`. A straight violation of extract-mechanically- change-nothing, caught only by reading the check-mode diff line by line. Reverted to match the original. - PRE-EXISTING, and dangerous: the playbook declared `owner: "{{ ansible_user }}"` (1000) on the MariaDB data directory, which the container owns as uid 999. Confirmed against `git show HEAD:` before concluding it was not mine. It had drifted since the containers were created and went unnoticed because the playbook had not been run since. This was not academic. The first real run pulled a newer mariadb:10.11 and recreated mempool-db; with the chown still in place MariaDB would have come back to a data directory it could not write. The role now ensures the directory exists and leaves ownership to the container. Verified after the run: /opt/mempool/mysql is still 999:999 and all three containers are healthy. This is a deliberate behaviour change, not part of the extraction. It is in this commit rather than a follow-up because the faithful version was never safe to run, so there was no intermediate state worth recording as verified. mempool_frontend_port moved to services_config.yml: two hosts need it (this role deploys the frontend, the Caddy play proxies to it from the edge host) and a role default is invisible to the second play. caddy_site's parameter assert caught this loudly - "'mempool_frontend_port' is undefined" - rather than silently. Verified: check-mode diff clean apart from unavoidable check-mode artifacts; first run ok=24 changed=5, zero failures; second run changed=2 - the two bare `command:` tasks (pull, compose up) that have no changed_when and always report changed. That is the idempotent floor. All three health checks report ExecMainStatus 0 with their push URLs intact. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-12 18:49:11 +02:00
loop: "{{ mempool_healthchecks }}"
loop_control:
loop_var: hc
label: "{{ hc.name }}"
- name: Create systemd timers for health checks
ansible.builtin.template:
src: healthcheck.timer.j2
dest: "/etc/systemd/system/mempool-{{ hc.name }}-healthcheck.timer"
owner: root
group: root
mode: '0644'
loop: "{{ mempool_healthchecks }}"
loop_control:
loop_var: hc
label: "{{ hc.name }}"
- name: Reload systemd daemon
systemd:
daemon_reload: yes
# Ungated on purpose: enabling a timer is deployment, not monitoring. The
# deprecation banner swept this up with the push plumbing, so Ansible stopped
# managing three timers that are in fact running on the host.
- name: Enable and start health check timers
systemd:
name: "mempool-{{ hc.name }}-healthcheck.timer"
enabled: yes
state: started
loop: "{{ mempool_healthchecks }}"
loop_control:
loop_var: hc
label: "{{ hc.name }}"