personal_infra/ansible/roles/healthcheck/tasks/main.yml

91 lines
3.4 KiB
YAML
Raw Normal View History

monitoring: recover host checks for the whole estate, reported to Gatus Five checks, 27 endpoints, replacing what Uptime Kuma used to watch: is it up every 5min, all hosts is disk full daily, all hosts is CPU hot every 5min, nodito is ZFS broken daily, nodito is UPS online every 5min, nodito Two roles, kept separate so neither knows about the other - they meet at a URL and a token, the same way caddy_site and each service meet at a vhost: roles/gatus_endpoint runs on the observability host, writes ONE file into /opt/gatus/config/endpoints/. Gatus merges every *.yaml there and appends lists, so callers compose without coordinating. roles/healthcheck runs on the monitored host: a check script, a systemd service, a timer, and an optional push. Ships a library of check bodies under templates/checks/. Everything PUSHES. Gatus never reaches out, which matters because nodito and its VMs are behind NAT, and because four of the five checks are internal state with no pollable surface at all. Liveness pushes too, deliberately: a heartbeat proves the host is running AND can reach the internet, where an ICMP probe from one vantage point only proves it answers pings from there. And since Gatus alerts when a heartbeat window expires, a check that stops running raises the alarm by itself - a dead timer looks exactly like a dead host, which is the correct reading. One bearer token per host, generated straight into the vault and never printed. A token only writes results for its own host's endpoints, so a compromised host can lie about itself, which it could do anyway. Three things learned from the source that shaped this: * Gatus polls its own config every 30s and reloads (main.listenToConfigurationFileChanges), so gatus_endpoint needs no restart handler - writing the file IS the deploy. * ...but on a reload it panics if the new config fails to parse, unless skip-invalid-config-update is set. Endpoint files are contributed by other playbooks, so one malformed file would take the monitor down at the worst possible moment. Now set. * The push URL uses a key Gatus computes, not the name you write: sanitize(group) + "_" + sanitize(name), lowercased with / _ . , space # + & replaced by "-" (config/key/key.go). So knots_box_local is knots-box-local in the URL. The playbook derives it rather than hand-writing. storage: maximum-number-of-results 900, up from upstream's 100. Gatus bounds the database by COUNT and trims inline on insert, so there is no retention job and no way to fill a disk - but history depth is then a function of check frequency, and 100 results at a 5-minute interval is 8 hours. 900 is ~3 days of liveness and ~2.5 years of the daily disk check. The uptime table is separate and its 30-day retention is hard-coded upstream. A bug worth recording: the first deploy shipped five scripts that all died with "syntax error: unexpected end of file". Jinja strips an included template's trailing newline and trim_blocks then eats the newline after {% endif %}, so the closing brace of check() landed on the same line as the body's last statement - `return 0}`. Every check was broken and the deploy still reported failed=0, because the role's "run once" task has failed_when: false and reports the result as a debug message nobody read. The blank line that fixes it is now load-bearing and commented as such. Verified by triggering every unit by hand rather than waiting on timers: all checks exit 0 on all hosts, and Gatus shows 26 UP / 1 DOWN. The one DOWN is liveness_watchtower, which is honest - that host currently refuses SSH (TCP connects, no banner exchange) and is excluded from this deploy. It is also the box still running Uptime Kuma. Known waste, not yet fixed: healthcheck installs its dependencies per CHECK rather than per HOST, so apt runs 29 times estate-wide for a curl that is already present, and daemon_reload runs 4x per host. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 08:54:10 +02:00
---
- name: "Assert healthcheck '{{ healthcheck_name }}' is fully specified"
ansible.builtin.assert:
that:
- healthcheck_name | length > 0
- healthcheck_description | length > 0
- (healthcheck_check | length > 0) != (healthcheck_command | length > 0)
- not (healthcheck_push_url | length > 0) or (healthcheck_push_token | length > 0)
fail_msg: >-
healthcheck needs a name, a description, exactly one of healthcheck_check
or healthcheck_command, and a token whenever a push URL is set. A push URL
with no token would report to Gatus and be rejected 401 on every run.
monitoring: retire the Uptime-Kuma-era checks, add ZFS pool capacity Five things deprecated, each verified against the DEPLOYED script before being deleted rather than assumed superseded: infra/410_disk_usage_alerts.yml -> disk-usage check (infra/400) infra/420_system_healthcheck.yml -> liveness check (infra/400) infra/430_cpu_temp_alerts.yml -> cpu-temp check (infra/400) 32_zfs play 2 (monitoring half) -> zfs-health check (infra/400) 34_nut play 2 (entirely) -> ups-status check (infra/400) Nothing is lost by the swap. The old system_healthcheck.sh only computed uptime and pushed, which is exactly a liveness heartbeat. The old disk monitor was WEAKER than its replacement: it checked "/" alone at 80%, where the new one walks every real filesystem at 85%. Deleting the playbooks was not the hard part. The units they installed live on the hosts, enabled, and keep firing regardless of what the repo says - two of them were still pushing to uptime.contrapeso.xyz every 15 minutes across nine machines. A playbook deleted without a cleanup leaves its output running forever with nothing left to explain it. So infra/409_remove_legacy_monitoring stops, disables and removes the units, deletes /opt/{disk-monitoring, system-healthcheck,nodito-monitoring,zfs-monitoring}, and removes the orphaned hand-written ups-heartbeat.sh. It ends by grepping for any surviving Kuma reference and reporting it. Kept permanently and idempotent, so a rebuilt or restored host cannot quietly bring them back. A trap avoided: the monthly ZFS scrub lived INSIDE 32_zfs play 2. Deleting the play wholesale would have silently stopped scrubbing the pool - and an unscrubbed pool makes the health check meaningless, because it would have nothing true to report. That play is now scrub-only and check-runs ok=5 changed=0. ZFS pool capacity added as a sixth condition to the zfs-health check. `zpool status` reports a 95% full pool as perfectly ONLINE, so capacity has to be read separately with `zpool list` - and it is the failure you get warning of rather than the one you discover. Threshold 80%, because ZFS allocation degrades badly past roughly that and fragmentation is hard to undo. The pool is at 47%. Verified both directions: passes on the real pool, and a simulated 91% exits 1. Also fixed the waste recorded in c2de6db: the healthcheck role installed its dependencies once per CHECK rather than per HOST - 29 apt transactions estate-wide for a curl already present, and the slowest part of every deploy. It now deduplicates within a play run, and the redundant standalone daemon_reload is gone (the systemd task already does one). site.yml updated, which exposed that services/gatus was never in it. It now runs before the three registration playbooks, since registering endpoints against a Gatus that is not yet serving would simply fail. Verified: no legacy timer remains on any host; 83 endpoints, 83 UP, 0 DOWN. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 10:18:57 +02:00
# Deduplicated across the whole play run. This role is included once PER CHECK,
# and a host with several checks was otherwise running apt several times to
# install a curl that was already there - 29 apt transactions estate-wide, and
# the slowest thing in the deploy by a wide margin. The fact below remembers
# what has already been ensured on this host.
monitoring: recover host checks for the whole estate, reported to Gatus Five checks, 27 endpoints, replacing what Uptime Kuma used to watch: is it up every 5min, all hosts is disk full daily, all hosts is CPU hot every 5min, nodito is ZFS broken daily, nodito is UPS online every 5min, nodito Two roles, kept separate so neither knows about the other - they meet at a URL and a token, the same way caddy_site and each service meet at a vhost: roles/gatus_endpoint runs on the observability host, writes ONE file into /opt/gatus/config/endpoints/. Gatus merges every *.yaml there and appends lists, so callers compose without coordinating. roles/healthcheck runs on the monitored host: a check script, a systemd service, a timer, and an optional push. Ships a library of check bodies under templates/checks/. Everything PUSHES. Gatus never reaches out, which matters because nodito and its VMs are behind NAT, and because four of the five checks are internal state with no pollable surface at all. Liveness pushes too, deliberately: a heartbeat proves the host is running AND can reach the internet, where an ICMP probe from one vantage point only proves it answers pings from there. And since Gatus alerts when a heartbeat window expires, a check that stops running raises the alarm by itself - a dead timer looks exactly like a dead host, which is the correct reading. One bearer token per host, generated straight into the vault and never printed. A token only writes results for its own host's endpoints, so a compromised host can lie about itself, which it could do anyway. Three things learned from the source that shaped this: * Gatus polls its own config every 30s and reloads (main.listenToConfigurationFileChanges), so gatus_endpoint needs no restart handler - writing the file IS the deploy. * ...but on a reload it panics if the new config fails to parse, unless skip-invalid-config-update is set. Endpoint files are contributed by other playbooks, so one malformed file would take the monitor down at the worst possible moment. Now set. * The push URL uses a key Gatus computes, not the name you write: sanitize(group) + "_" + sanitize(name), lowercased with / _ . , space # + & replaced by "-" (config/key/key.go). So knots_box_local is knots-box-local in the URL. The playbook derives it rather than hand-writing. storage: maximum-number-of-results 900, up from upstream's 100. Gatus bounds the database by COUNT and trims inline on insert, so there is no retention job and no way to fill a disk - but history depth is then a function of check frequency, and 100 results at a 5-minute interval is 8 hours. 900 is ~3 days of liveness and ~2.5 years of the daily disk check. The uptime table is separate and its 30-day retention is hard-coded upstream. A bug worth recording: the first deploy shipped five scripts that all died with "syntax error: unexpected end of file". Jinja strips an included template's trailing newline and trim_blocks then eats the newline after {% endif %}, so the closing brace of check() landed on the same line as the body's last statement - `return 0}`. Every check was broken and the deploy still reported failed=0, because the role's "run once" task has failed_when: false and reports the result as a debug message nobody read. The blank line that fixes it is now load-bearing and commented as such. Verified by triggering every unit by hand rather than waiting on timers: all checks exit 0 on all hosts, and Gatus shows 26 UP / 1 DOWN. The one DOWN is liveness_watchtower, which is honest - that host currently refuses SSH (TCP connects, no banner exchange) and is excluded from this deploy. It is also the box still running Uptime Kuma. Known waste, not yet fixed: healthcheck installs its dependencies per CHECK rather than per HOST, so apt runs 29 times estate-wide for a curl that is already present, and daemon_reload runs 4x per host. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 08:54:10 +02:00
- name: Install healthcheck dependencies
ansible.builtin.package:
monitoring: retire the Uptime-Kuma-era checks, add ZFS pool capacity Five things deprecated, each verified against the DEPLOYED script before being deleted rather than assumed superseded: infra/410_disk_usage_alerts.yml -> disk-usage check (infra/400) infra/420_system_healthcheck.yml -> liveness check (infra/400) infra/430_cpu_temp_alerts.yml -> cpu-temp check (infra/400) 32_zfs play 2 (monitoring half) -> zfs-health check (infra/400) 34_nut play 2 (entirely) -> ups-status check (infra/400) Nothing is lost by the swap. The old system_healthcheck.sh only computed uptime and pushed, which is exactly a liveness heartbeat. The old disk monitor was WEAKER than its replacement: it checked "/" alone at 80%, where the new one walks every real filesystem at 85%. Deleting the playbooks was not the hard part. The units they installed live on the hosts, enabled, and keep firing regardless of what the repo says - two of them were still pushing to uptime.contrapeso.xyz every 15 minutes across nine machines. A playbook deleted without a cleanup leaves its output running forever with nothing left to explain it. So infra/409_remove_legacy_monitoring stops, disables and removes the units, deletes /opt/{disk-monitoring, system-healthcheck,nodito-monitoring,zfs-monitoring}, and removes the orphaned hand-written ups-heartbeat.sh. It ends by grepping for any surviving Kuma reference and reporting it. Kept permanently and idempotent, so a rebuilt or restored host cannot quietly bring them back. A trap avoided: the monthly ZFS scrub lived INSIDE 32_zfs play 2. Deleting the play wholesale would have silently stopped scrubbing the pool - and an unscrubbed pool makes the health check meaningless, because it would have nothing true to report. That play is now scrub-only and check-runs ok=5 changed=0. ZFS pool capacity added as a sixth condition to the zfs-health check. `zpool status` reports a 95% full pool as perfectly ONLINE, so capacity has to be read separately with `zpool list` - and it is the failure you get warning of rather than the one you discover. Threshold 80%, because ZFS allocation degrades badly past roughly that and fragmentation is hard to undo. The pool is at 47%. Verified both directions: passes on the real pool, and a simulated 91% exits 1. Also fixed the waste recorded in c2de6db: the healthcheck role installed its dependencies once per CHECK rather than per HOST - 29 apt transactions estate-wide for a curl already present, and the slowest part of every deploy. It now deduplicates within a play run, and the redundant standalone daemon_reload is gone (the systemd task already does one). site.yml updated, which exposed that services/gatus was never in it. It now runs before the three registration playbooks, since registering endpoints against a Gatus that is not yet serving would simply fail. Verified: no legacy timer remains on any host; 83 endpoints, 83 UP, 0 DOWN. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 10:18:57 +02:00
name: "{{ healthcheck_wanted_packages }}"
monitoring: recover host checks for the whole estate, reported to Gatus Five checks, 27 endpoints, replacing what Uptime Kuma used to watch: is it up every 5min, all hosts is disk full daily, all hosts is CPU hot every 5min, nodito is ZFS broken daily, nodito is UPS online every 5min, nodito Two roles, kept separate so neither knows about the other - they meet at a URL and a token, the same way caddy_site and each service meet at a vhost: roles/gatus_endpoint runs on the observability host, writes ONE file into /opt/gatus/config/endpoints/. Gatus merges every *.yaml there and appends lists, so callers compose without coordinating. roles/healthcheck runs on the monitored host: a check script, a systemd service, a timer, and an optional push. Ships a library of check bodies under templates/checks/. Everything PUSHES. Gatus never reaches out, which matters because nodito and its VMs are behind NAT, and because four of the five checks are internal state with no pollable surface at all. Liveness pushes too, deliberately: a heartbeat proves the host is running AND can reach the internet, where an ICMP probe from one vantage point only proves it answers pings from there. And since Gatus alerts when a heartbeat window expires, a check that stops running raises the alarm by itself - a dead timer looks exactly like a dead host, which is the correct reading. One bearer token per host, generated straight into the vault and never printed. A token only writes results for its own host's endpoints, so a compromised host can lie about itself, which it could do anyway. Three things learned from the source that shaped this: * Gatus polls its own config every 30s and reloads (main.listenToConfigurationFileChanges), so gatus_endpoint needs no restart handler - writing the file IS the deploy. * ...but on a reload it panics if the new config fails to parse, unless skip-invalid-config-update is set. Endpoint files are contributed by other playbooks, so one malformed file would take the monitor down at the worst possible moment. Now set. * The push URL uses a key Gatus computes, not the name you write: sanitize(group) + "_" + sanitize(name), lowercased with / _ . , space # + & replaced by "-" (config/key/key.go). So knots_box_local is knots-box-local in the URL. The playbook derives it rather than hand-writing. storage: maximum-number-of-results 900, up from upstream's 100. Gatus bounds the database by COUNT and trims inline on insert, so there is no retention job and no way to fill a disk - but history depth is then a function of check frequency, and 100 results at a 5-minute interval is 8 hours. 900 is ~3 days of liveness and ~2.5 years of the daily disk check. The uptime table is separate and its 30-day retention is hard-coded upstream. A bug worth recording: the first deploy shipped five scripts that all died with "syntax error: unexpected end of file". Jinja strips an included template's trailing newline and trim_blocks then eats the newline after {% endif %}, so the closing brace of check() landed on the same line as the body's last statement - `return 0}`. Every check was broken and the deploy still reported failed=0, because the role's "run once" task has failed_when: false and reports the result as a debug message nobody read. The blank line that fixes it is now load-bearing and commented as such. Verified by triggering every unit by hand rather than waiting on timers: all checks exit 0 on all hosts, and Gatus shows 26 UP / 1 DOWN. The one DOWN is liveness_watchtower, which is honest - that host currently refuses SSH (TCP connects, no banner exchange) and is excluded from this deploy. It is also the box still running Uptime Kuma. Known waste, not yet fixed: healthcheck installs its dependencies per CHECK rather than per HOST, so apt runs 29 times estate-wide for a curl that is already present, and daemon_reload runs 4x per host. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 08:54:10 +02:00
state: present
monitoring: retire the Uptime-Kuma-era checks, add ZFS pool capacity Five things deprecated, each verified against the DEPLOYED script before being deleted rather than assumed superseded: infra/410_disk_usage_alerts.yml -> disk-usage check (infra/400) infra/420_system_healthcheck.yml -> liveness check (infra/400) infra/430_cpu_temp_alerts.yml -> cpu-temp check (infra/400) 32_zfs play 2 (monitoring half) -> zfs-health check (infra/400) 34_nut play 2 (entirely) -> ups-status check (infra/400) Nothing is lost by the swap. The old system_healthcheck.sh only computed uptime and pushed, which is exactly a liveness heartbeat. The old disk monitor was WEAKER than its replacement: it checked "/" alone at 80%, where the new one walks every real filesystem at 85%. Deleting the playbooks was not the hard part. The units they installed live on the hosts, enabled, and keep firing regardless of what the repo says - two of them were still pushing to uptime.contrapeso.xyz every 15 minutes across nine machines. A playbook deleted without a cleanup leaves its output running forever with nothing left to explain it. So infra/409_remove_legacy_monitoring stops, disables and removes the units, deletes /opt/{disk-monitoring, system-healthcheck,nodito-monitoring,zfs-monitoring}, and removes the orphaned hand-written ups-heartbeat.sh. It ends by grepping for any surviving Kuma reference and reporting it. Kept permanently and idempotent, so a rebuilt or restored host cannot quietly bring them back. A trap avoided: the monthly ZFS scrub lived INSIDE 32_zfs play 2. Deleting the play wholesale would have silently stopped scrubbing the pool - and an unscrubbed pool makes the health check meaningless, because it would have nothing true to report. That play is now scrub-only and check-runs ok=5 changed=0. ZFS pool capacity added as a sixth condition to the zfs-health check. `zpool status` reports a 95% full pool as perfectly ONLINE, so capacity has to be read separately with `zpool list` - and it is the failure you get warning of rather than the one you discover. Threshold 80%, because ZFS allocation degrades badly past roughly that and fragmentation is hard to undo. The pool is at 47%. Verified both directions: passes on the real pool, and a simulated 91% exits 1. Also fixed the waste recorded in c2de6db: the healthcheck role installed its dependencies once per CHECK rather than per HOST - 29 apt transactions estate-wide for a curl already present, and the slowest part of every deploy. It now deduplicates within a play run, and the redundant standalone daemon_reload is gone (the systemd task already does one). site.yml updated, which exposed that services/gatus was never in it. It now runs before the three registration playbooks, since registering endpoints against a Gatus that is not yet serving would simply fail. Verified: no legacy timer remains on any host; 83 endpoints, 83 UP, 0 DOWN. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 10:18:57 +02:00
vars:
healthcheck_wanted_packages: >-
{{ (healthcheck_packages | default(['curl']))
| difference(healthcheck_installed_packages | default([])) }}
when: healthcheck_wanted_packages | length > 0
- name: Remember which dependencies this host already has
ansible.builtin.set_fact:
healthcheck_installed_packages: >-
{{ (healthcheck_installed_packages | default([]))
| union(healthcheck_packages | default(['curl'])) }}
monitoring: recover host checks for the whole estate, reported to Gatus Five checks, 27 endpoints, replacing what Uptime Kuma used to watch: is it up every 5min, all hosts is disk full daily, all hosts is CPU hot every 5min, nodito is ZFS broken daily, nodito is UPS online every 5min, nodito Two roles, kept separate so neither knows about the other - they meet at a URL and a token, the same way caddy_site and each service meet at a vhost: roles/gatus_endpoint runs on the observability host, writes ONE file into /opt/gatus/config/endpoints/. Gatus merges every *.yaml there and appends lists, so callers compose without coordinating. roles/healthcheck runs on the monitored host: a check script, a systemd service, a timer, and an optional push. Ships a library of check bodies under templates/checks/. Everything PUSHES. Gatus never reaches out, which matters because nodito and its VMs are behind NAT, and because four of the five checks are internal state with no pollable surface at all. Liveness pushes too, deliberately: a heartbeat proves the host is running AND can reach the internet, where an ICMP probe from one vantage point only proves it answers pings from there. And since Gatus alerts when a heartbeat window expires, a check that stops running raises the alarm by itself - a dead timer looks exactly like a dead host, which is the correct reading. One bearer token per host, generated straight into the vault and never printed. A token only writes results for its own host's endpoints, so a compromised host can lie about itself, which it could do anyway. Three things learned from the source that shaped this: * Gatus polls its own config every 30s and reloads (main.listenToConfigurationFileChanges), so gatus_endpoint needs no restart handler - writing the file IS the deploy. * ...but on a reload it panics if the new config fails to parse, unless skip-invalid-config-update is set. Endpoint files are contributed by other playbooks, so one malformed file would take the monitor down at the worst possible moment. Now set. * The push URL uses a key Gatus computes, not the name you write: sanitize(group) + "_" + sanitize(name), lowercased with / _ . , space # + & replaced by "-" (config/key/key.go). So knots_box_local is knots-box-local in the URL. The playbook derives it rather than hand-writing. storage: maximum-number-of-results 900, up from upstream's 100. Gatus bounds the database by COUNT and trims inline on insert, so there is no retention job and no way to fill a disk - but history depth is then a function of check frequency, and 100 results at a 5-minute interval is 8 hours. 900 is ~3 days of liveness and ~2.5 years of the daily disk check. The uptime table is separate and its 30-day retention is hard-coded upstream. A bug worth recording: the first deploy shipped five scripts that all died with "syntax error: unexpected end of file". Jinja strips an included template's trailing newline and trim_blocks then eats the newline after {% endif %}, so the closing brace of check() landed on the same line as the body's last statement - `return 0}`. Every check was broken and the deploy still reported failed=0, because the role's "run once" task has failed_when: false and reports the result as a debug message nobody read. The blank line that fixes it is now load-bearing and commented as such. Verified by triggering every unit by hand rather than waiting on timers: all checks exit 0 on all hosts, and Gatus shows 26 UP / 1 DOWN. The one DOWN is liveness_watchtower, which is honest - that host currently refuses SSH (TCP connects, no banner exchange) and is excluded from this deploy. It is also the box still running Uptime Kuma. Known waste, not yet fixed: healthcheck installs its dependencies per CHECK rather than per HOST, so apt runs 29 times estate-wide for a curl that is already present, and daemon_reload runs 4x per host. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 08:54:10 +02:00
- name: Create the healthcheck log directory
ansible.builtin.file:
path: "{{ healthcheck_log_dir }}"
state: directory
owner: root
group: root
mode: "0750"
- name: "Install the {{ healthcheck_name }} check script"
ansible.builtin.template:
src: healthcheck.sh.j2
dest: "{{ healthcheck_script_dir }}/{{ healthcheck_name }}-healthcheck.sh"
owner: root
group: root
mode: "0755"
# The token is in this unit file, so it must not be world-readable.
- name: "Install the {{ healthcheck_name }} systemd service"
ansible.builtin.template:
src: healthcheck.service.j2
dest: "/etc/systemd/system/{{ healthcheck_name }}-healthcheck.service"
owner: root
group: root
mode: "0600"
- name: "Install the {{ healthcheck_name }} systemd timer"
ansible.builtin.template:
src: healthcheck.timer.j2
dest: "/etc/systemd/system/{{ healthcheck_name }}-healthcheck.timer"
owner: root
group: root
mode: "0644"
# `restarted`, not `started`: started is a no-op on an already-active timer, so
# a changed interval or a stuck timer would never be picked up.
- name: "Enable and start the {{ healthcheck_name }} timer"
ansible.builtin.systemd:
name: "{{ healthcheck_name }}-healthcheck.timer"
enabled: yes
state: restarted
daemon_reload: yes
- name: "Run the {{ healthcheck_name }} check once now"
ansible.builtin.command: "{{ healthcheck_script_dir }}/{{ healthcheck_name }}-healthcheck.sh"
environment:
HEALTHCHECK_PUSH_URL: "{{ healthcheck_push_url }}"
HEALTHCHECK_PUSH_TOKEN: "{{ healthcheck_push_token }}"
register: healthcheck_first_run
changed_when: false
failed_when: false
- name: "Report the first {{ healthcheck_name }} result"
ansible.builtin.debug:
msg: >-
{{ healthcheck_name }}: {{ 'HEALTHY' if healthcheck_first_run.rc == 0
else 'UNHEALTHY (rc=' ~ healthcheck_first_run.rc ~ ')' }}