monitoring: recover host checks for the whole estate, reported to Gatus

Five checks, 27 endpoints, replacing what Uptime Kuma used to watch:

  is it up      every 5min, all hosts
  is disk full  daily,      all hosts
  is CPU hot    every 5min, nodito
  is ZFS broken daily,      nodito
  is UPS online every 5min, nodito

Two roles, kept separate so neither knows about the other - they meet at a URL
and a token, the same way caddy_site and each service meet at a vhost:

  roles/gatus_endpoint  runs on the observability host, writes ONE file into
                        /opt/gatus/config/endpoints/. Gatus merges every *.yaml
                        there and appends lists, so callers compose without
                        coordinating.
  roles/healthcheck     runs on the monitored host: a check script, a systemd
                        service, a timer, and an optional push. Ships a library
                        of check bodies under templates/checks/.

Everything PUSHES. Gatus never reaches out, which matters because nodito and its
VMs are behind NAT, and because four of the five checks are internal state with
no pollable surface at all. Liveness pushes too, deliberately: a heartbeat proves
the host is running AND can reach the internet, where an ICMP probe from one
vantage point only proves it answers pings from there. And since Gatus alerts
when a heartbeat window expires, a check that stops running raises the alarm by
itself - a dead timer looks exactly like a dead host, which is the correct
reading.

One bearer token per host, generated straight into the vault and never printed.
A token only writes results for its own host's endpoints, so a compromised host
can lie about itself, which it could do anyway.

Three things learned from the source that shaped this:

  * Gatus polls its own config every 30s and reloads
    (main.listenToConfigurationFileChanges), so gatus_endpoint needs no restart
    handler - writing the file IS the deploy.
  * ...but on a reload it panics if the new config fails to parse, unless
    skip-invalid-config-update is set. Endpoint files are contributed by other
    playbooks, so one malformed file would take the monitor down at the worst
    possible moment. Now set.
  * The push URL uses a key Gatus computes, not the name you write:
    sanitize(group) + "_" + sanitize(name), lowercased with / _ . , space # + &
    replaced by "-" (config/key/key.go). So knots_box_local is
    knots-box-local in the URL. The playbook derives it rather than hand-writing.

storage: maximum-number-of-results 900, up from upstream's 100. Gatus bounds the
database by COUNT and trims inline on insert, so there is no retention job and
no way to fill a disk - but history depth is then a function of check frequency,
and 100 results at a 5-minute interval is 8 hours. 900 is ~3 days of liveness
and ~2.5 years of the daily disk check. The uptime table is separate and its
30-day retention is hard-coded upstream.

A bug worth recording: the first deploy shipped five scripts that all died with
"syntax error: unexpected end of file". Jinja strips an included template's
trailing newline and trim_blocks then eats the newline after {% endif %}, so the
closing brace of check() landed on the same line as the body's last statement -
`return 0}`. Every check was broken and the deploy still reported failed=0,
because the role's "run once" task has failed_when: false and reports the result
as a debug message nobody read. The blank line that fixes it is now load-bearing
and commented as such.

Verified by triggering every unit by hand rather than waiting on timers: all
checks exit 0 on all hosts, and Gatus shows 26 UP / 1 DOWN. The one DOWN is
liveness_watchtower, which is honest - that host currently refuses SSH (TCP
connects, no banner exchange) and is excluded from this deploy. It is also the
box still running Uptime Kuma.

Known waste, not yet fixed: healthcheck installs its dependencies per CHECK
rather than per HOST, so apt runs 29 times estate-wide for a curl that is
already present, and daemon_reload runs 4x per host.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
counterweight 2026-09-14 08:54:10 +02:00
parent fa9f7d10cd
commit ede407ebe4
Signed by: counterweight
GPG key ID: 883EDBAA726BD96C
17 changed files with 887 additions and 200 deletions

View file

@ -0,0 +1,38 @@
---
# One health check: a script, a systemd service, a timer, and an optional push.
#
# The exit code is the answer and systemd keeps it:
# systemctl is-failed <name>-healthcheck.service
# Reporting anywhere else is optional and generic. Point healthcheck_push_url at
# Gatus, or at whatever replaces it, or at nothing.
healthcheck_name: "" # e.g. disk-usage -> disk-usage-healthcheck
healthcheck_description: ""
# The check itself. Pick ONE:
# healthcheck_check: a template under templates/checks/ (without .sh.j2)
# healthcheck_command: a shell one-liner that exits 0 for healthy
healthcheck_check: ""
healthcheck_command: ""
# systemd timer. OnUnitActiveSec unless healthcheck_on_calendar is set.
healthcheck_interval: "5min"
healthcheck_on_calendar: ""
healthcheck_boot_delay: "2min"
# ── Reporting ────────────────────────────────────────────────────────────────
# Gatus external endpoints:
# POST {url}?success=true|false&error=...
# Authorization: Bearer {token}
# Empty url = check and log only, which is a valid state and not an error.
healthcheck_push_url: ""
healthcheck_push_token: ""
healthcheck_script_dir: /usr/local/bin
healthcheck_log_dir: /var/log/healthchecks
# Per-check knobs, consumed by the templates under checks/
healthcheck_disk_threshold: 85 # percent
healthcheck_cpu_temp_threshold: 80 # celsius
healthcheck_zfs_pool: ""
healthcheck_ups_name: ""

View file

@ -0,0 +1,78 @@
---
- name: "Assert healthcheck '{{ healthcheck_name }}' is fully specified"
ansible.builtin.assert:
that:
- healthcheck_name | length > 0
- healthcheck_description | length > 0
- (healthcheck_check | length > 0) != (healthcheck_command | length > 0)
- not (healthcheck_push_url | length > 0) or (healthcheck_push_token | length > 0)
fail_msg: >-
healthcheck needs a name, a description, exactly one of healthcheck_check
or healthcheck_command, and a token whenever a push URL is set. A push URL
with no token would report to Gatus and be rejected 401 on every run.
- name: Install healthcheck dependencies
ansible.builtin.package:
name: "{{ healthcheck_packages | default(['curl']) }}"
state: present
- name: Create the healthcheck log directory
ansible.builtin.file:
path: "{{ healthcheck_log_dir }}"
state: directory
owner: root
group: root
mode: "0750"
- name: "Install the {{ healthcheck_name }} check script"
ansible.builtin.template:
src: healthcheck.sh.j2
dest: "{{ healthcheck_script_dir }}/{{ healthcheck_name }}-healthcheck.sh"
owner: root
group: root
mode: "0755"
# The token is in this unit file, so it must not be world-readable.
- name: "Install the {{ healthcheck_name }} systemd service"
ansible.builtin.template:
src: healthcheck.service.j2
dest: "/etc/systemd/system/{{ healthcheck_name }}-healthcheck.service"
owner: root
group: root
mode: "0600"
- name: "Install the {{ healthcheck_name }} systemd timer"
ansible.builtin.template:
src: healthcheck.timer.j2
dest: "/etc/systemd/system/{{ healthcheck_name }}-healthcheck.timer"
owner: root
group: root
mode: "0644"
- name: Reload systemd
ansible.builtin.systemd:
daemon_reload: yes
# `restarted`, not `started`: started is a no-op on an already-active timer, so
# a changed interval or a stuck timer would never be picked up.
- name: "Enable and start the {{ healthcheck_name }} timer"
ansible.builtin.systemd:
name: "{{ healthcheck_name }}-healthcheck.timer"
enabled: yes
state: restarted
daemon_reload: yes
- name: "Run the {{ healthcheck_name }} check once now"
ansible.builtin.command: "{{ healthcheck_script_dir }}/{{ healthcheck_name }}-healthcheck.sh"
environment:
HEALTHCHECK_PUSH_URL: "{{ healthcheck_push_url }}"
HEALTHCHECK_PUSH_TOKEN: "{{ healthcheck_push_token }}"
register: healthcheck_first_run
changed_when: false
failed_when: false
- name: "Report the first {{ healthcheck_name }} result"
ansible.builtin.debug:
msg: >-
{{ healthcheck_name }}: {{ 'HEALTHY' if healthcheck_first_run.rc == 0
else 'UNHEALTHY (rc=' ~ healthcheck_first_run.rc ~ ')' }}

View file

@ -0,0 +1,28 @@
# Hottest core across every thermal zone and hwmon sensor lm-sensors knows
# about. Reading the hottest rather than an average is deliberate: one core
# throttling is a real problem that an average hides.
local threshold={{ healthcheck_cpu_temp_threshold }}
local hottest=0 label=""
while read -r t; do
[ -z "$t" ] && continue
t=${t%.*}
if [ "$t" -gt "$hottest" ]; then hottest=$t; fi
done < <(sensors -u 2>/dev/null | awk '/_input:/ && /temp/ {print $2}')
# Fall back to the kernel thermal zones if lm-sensors reports nothing.
if [ "$hottest" -eq 0 ]; then
for z in /sys/class/thermal/thermal_zone*/temp; do
[ -r "$z" ] || continue
local milli; milli=$(cat "$z" 2>/dev/null) || continue
local c=$((milli / 1000))
if [ "$c" -gt "$hottest" ]; then hottest=$c; label=$(cat "${z%/temp}/type" 2>/dev/null); fi
done
fi
if [ "$hottest" -eq 0 ]; then
MESSAGE="no temperature sensors readable"
return 1
fi
MESSAGE="${hottest}C${label:+ (${label})}"
[ "$hottest" -lt "$threshold" ]

View file

@ -0,0 +1,19 @@
# Every real filesystem must be under the threshold. tmpfs, devtmpfs,
# squashfs and overlay are excluded: they are either RAM, read-only, or
# container layers, and none of them fills up in a way an operator can act on.
local threshold={{ healthcheck_disk_threshold }}
local worst=0 worst_mount="" over=""
while read -r pct mount; do
pct=${pct%\%}
[ -z "$pct" ] && continue
if [ "$pct" -gt "$worst" ]; then worst=$pct; worst_mount=$mount; fi
if [ "$pct" -ge "$threshold" ]; then over="${over}${over:+, }${mount} ${pct}%"; fi
done < <(df -P -x tmpfs -x devtmpfs -x squashfs -x overlay --output=pcent,target 2>/dev/null | tail -n +2)
if [ -n "$over" ]; then
MESSAGE="over ${threshold}%: ${over}"
return 1
fi
MESSAGE="max ${worst}% on ${worst_mount:-/}"
return 0

View file

@ -0,0 +1,6 @@
# Liveness has no test to run: the fact that this script executed at all is
# the signal. What proves the host is alive is the PUSH arriving at Gatus,
# and what detects the host being dead is the heartbeat window expiring with
# no push. So this always succeeds - the reporting is the check.
MESSAGE="up since $(uptime -p 2>/dev/null || echo unknown)"
return 0

View file

@ -0,0 +1,18 @@
# OL means on line power. Anything else - OB (on battery), LB (low battery),
# or no answer at all - is a failure worth waking up for, because the
# hypervisor has a finite number of minutes left.
local ups="{{ healthcheck_ups_name }}"
local status charge runtime load
status=$(upsc "${ups}@localhost" ups.status 2>/dev/null)
if [ -z "$status" ]; then
MESSAGE="cannot reach UPS ${ups} via upsd"
return 1
fi
charge=$(upsc "${ups}@localhost" battery.charge 2>/dev/null)
runtime=$(upsc "${ups}@localhost" battery.runtime 2>/dev/null)
load=$(upsc "${ups}@localhost" ups.load 2>/dev/null)
MESSAGE="status=${status} charge=${charge}% runtime=${runtime}s load=${load}%"
[[ "$status" == *"OL"* ]]

View file

@ -0,0 +1,52 @@
# Five conditions, all of which have to hold. Ported from the check that
# infra/nodito/32_zfs_pool_setup_playbook.yml deployed, which was correct -
# only its reporting was tied to Uptime Kuma.
local pool="{{ healthcheck_zfs_pool }}"
local json issues=""
json=$(zpool status -j "$pool" 2>&1) || { MESSAGE="zpool status failed: ${json}"; return 1; }
# 1. pool state
local state
state=$(echo "$json" | jq -r --arg p "$pool" '.pools[$p].state')
[ "$state" = "ONLINE" ] || issues="${issues}${issues:+; }pool ${state}"
# 2. every vdev and device ONLINE
local bad
bad=$(echo "$json" | jq -r --arg p "$pool" '
.pools[$p].vdevs[] | .. | objects
| select(.state? and .state != "ONLINE")
| "\(.name // "unknown"):\(.state)"' 2>/dev/null | paste -sd, -)
[ -z "$bad" ] || issues="${issues}${issues:+; }devices ${bad}"
# 3. resilver in progress
local fn st
fn=$(echo "$json" | jq -r --arg p "$pool" '.pools[$p].scan_stats.function // "NONE"')
st=$(echo "$json" | jq -r --arg p "$pool" '.pools[$p].scan_stats.state // "NONE"')
if [ "$fn" = "RESILVER" ] && [ "$st" = "SCANNING" ]; then
issues="${issues}${issues:+; }resilvering"
fi
# 4. read/write/checksum errors. ZFS reports these as strings.
local errs
errs=$(echo "$json" | jq -r --arg p "$pool" '
.pools[$p].vdevs[] | .. | objects
| select(.name? and ((.read_errors // "0" | tonumber) > 0
or (.write_errors // "0" | tonumber) > 0
or (.checksum_errors // "0" | tonumber) > 0))
| "\(.name) r=\(.read_errors) w=\(.write_errors) c=\(.checksum_errors)"' 2>/dev/null | paste -sd, -)
[ -z "$errs" ] || issues="${issues}${issues:+; }errors ${errs}"
# 5. errors from the last scrub
local scan_err
scan_err=$(echo "$json" | jq -r --arg p "$pool" '.pools[$p].scan_stats.errors // "0"')
if [ -n "$scan_err" ] && [ "$scan_err" != "0" ] && [ "$scan_err" != "null" ]; then
issues="${issues}${issues:+; }scan errors ${scan_err}"
fi
if [ -n "$issues" ]; then MESSAGE="$issues"; return 1; fi
local scrub
scrub=$(echo "$json" | jq -r --arg p "$pool" '.pools[$p].scan_stats.start_time // "never"')
MESSAGE="${pool} ONLINE, last scrub ${scrub}"
return 0

View file

@ -0,0 +1,16 @@
[Unit]
Description={{ healthcheck_description }}
After=network-online.target
Wants=network-online.target
[Service]
Type=oneshot
User=root
ExecStart={{ healthcheck_script_dir }}/{{ healthcheck_name }}-healthcheck.sh
Environment=HEALTHCHECK_PUSH_URL={{ healthcheck_push_url }}
Environment=HEALTHCHECK_PUSH_TOKEN={{ healthcheck_push_token }}
StandardOutput=journal
StandardError=journal
[Install]
WantedBy=multi-user.target

View file

@ -0,0 +1,68 @@
#!/bin/bash
# {{ healthcheck_name }} — {{ healthcheck_description }}
# Managed by Ansible (roles/healthcheck). Do not edit on the host.
#
# The exit code is the real answer; systemd stores it:
# systemctl is-failed {{ healthcheck_name }}-healthcheck.service
# The push below is an optional extra, and having no URL is normal.
set -uo pipefail
LOG_FILE="{{ healthcheck_log_dir }}/{{ healthcheck_name }}.log"
PUSH_URL="${HEALTHCHECK_PUSH_URL:-}"
PUSH_TOKEN="${HEALTHCHECK_PUSH_TOKEN:-}"
log() { echo "$(date '+%Y-%m-%d %H:%M:%S') - $*" >> "$LOG_FILE"; }
# Report to Gatus as an external endpoint. Note this is a POST with a bearer
# token, not a GET with a query string - it is not the shape Uptime Kuma used.
report() {
local success="$1" message="$2"
# No push URL is normal, not an error: the exit code below is a complete
# answer for anything reading unit state.
[ -n "$PUSH_URL" ] || return 0
local encoded
encoded=$(printf '%s' "$message" | sed 's/%/%25/g; s/ /%20/g; s/&/%26/g; s/+/%2B/g; s/#/%23/g')
local code
code=$(curl -s -o /dev/null -w '%{http_code}' -X POST \
--max-time 15 --retry 2 --retry-delay 3 \
-H "Authorization: Bearer ${PUSH_TOKEN}" \
"${PUSH_URL}?success=${success}&error=${encoded}" 2>/dev/null)
if [ "$code" = "200" ]; then
log "reported success=${success}"
else
log "ERROR: report failed (HTTP ${code})"
return 1
fi
}
# ── the check ────────────────────────────────────────────────────────────────
#
# The blank line before the closing brace below is load-bearing. Jinja strips an
# included template's trailing newline, and trim_blocks (on by default in
# Ansible) then eats the newline after {% raw %}{% endif %}{% endraw %} - so without it the brace lands
# on the same line as the check body's last statement, producing `return 0}` and
# a script that dies with "syntax error: unexpected end of file".
check() {
{% if healthcheck_check %}
{% include 'checks/' ~ healthcheck_check ~ '.sh.j2' %}
{% else %}
{{ healthcheck_command }}
{% endif %}
}
MESSAGE=""
if check; then
log "OK${MESSAGE:+ - $MESSAGE}"
report "true" "${MESSAGE:-ok}"
exit 0
else
log "FAILED${MESSAGE:+ - $MESSAGE}"
report "false" "${MESSAGE:-check failed}"
exit 1
fi

View file

@ -0,0 +1,18 @@
[Unit]
Description=Run {{ healthcheck_description }}
Requires={{ healthcheck_name }}-healthcheck.service
[Timer]
OnBootSec={{ healthcheck_boot_delay }}
{% if healthcheck_on_calendar %}
OnCalendar={{ healthcheck_on_calendar }}
{% else %}
OnUnitActiveSec={{ healthcheck_interval }}
{% endif %}
# Run a missed occurrence on the next boot rather than silently skipping it.
Persistent=true
# Spread the pushes out so twelve hosts do not all report in the same second.
RandomizedDelaySec={{ healthcheck_randomized_delay | default('30') }}
[Install]
WantedBy=timers.target