personal_infra/ansible/roles/healthcheck/templates/healthcheck.sh.j2
counterweight efa9eb55ca
monitoring: systemd services, domain expiry, DNS correctness, public endpoints
Four more check types, 42 endpoints, taking the estate from 39 to 81.

── systemd services (infra/401) ────────────────────────────────────────────
Every unit we deploy, checked every 5 minutes with a 16-minute heartbeat.

This closes the gap that let a real bug run unnoticed earlier today: a backup
script left forgejo, lnbits, headscale and memos stopped, and NOTHING caught
it. The dumps exited 0, the artefacts were correct, the deploy said failed=0,
and liveness only proves the HOST is up - not that anything on it serves.

One endpoint PER UNIT but only ONE timer per host: the check iterates that
host's units and pushes a result for each, the way check-backups.sh reports per
source. Four units on vipy would otherwise mean four scripts, services and
timers. A single host-level red light would also say "something on vipy is
down" without saying which, which is not the question anyone has.

Keys are host-qualified because unit names collide - caddy runs on four
machines. Which units a host runs lives in host_vars/<host>/monitored_services,
because "what runs here" is a property of the machine, the same reasoning as
the cross-host ports.

nut-driver-enumerator is deliberately excluded: it is a oneshot generator that
is `enabled` but always `inactive`, so it would report down forever. Checked
live before excluding it.

── domain, DNS and public endpoints (infra/402) ────────────────────────────
The first checks that PULL rather than push, and that is the right way round:
all three are about how the outside world sees us, so they must be measured
from outside. Nothing is installed anywhere - no script, no timer, no token.
They also have no heartbeat, because a heartbeat answers "did the reporter
report"; when Gatus does the checking itself, failure is immediate.

  domain   1 endpoint,  24h, [DOMAIN_EXPIRATION] > 336h (two weeks)
  dns      11 endpoints, 24h, [DNS_RCODE] == NOERROR and [BODY] == the IP
  public   14 endpoints, 5m,  11 HTTPS + 3 TCP

Expected A records are derived from inventory (hostvars[host].ansible_host),
not written down again. The estate's recurring bug is an address recorded in a
second place and left behind when the machine moved; asserting against
inventory means a renumbered box is one edit, not two.

Expected HTTP status was checked live per site rather than assumed. Two return
401 - the Gatus dashboard and the DATUM dashboard, both behind basic auth - and
that is what is asserted: expecting 200 there would go green precisely when the
auth broke. headscale asserts /health rather than /, which is a 404 by design.
Every HTTPS check also carries [CERTIFICATE_EXPIRATION] > 168h, which is free
on an endpoint already being polled and catches a renewal that silently stops.

Two things learned the hard way, both now in comments:

  * A domain-expiry endpoint needs a URL SCHEME. Gatus derives the endpoint type
    from the prefix (endpoint.Type()), so a bare "contrapeso.xyz" is UNKNOWN and
    the whole config is rejected. It is https:// plus a DOMAIN_EXPIRATION
    condition and no status assertion, so the registrar's parking page at the
    apex is irrelevant. Upstream also enforces a 5m minimum interval for that
    placeholder, because it uses a free whois service.

  * That rejection proved skip-invalid-config-update was worth adding. Gatus
    logged "the configuration file was updated, but it is not valid, the old
    configuration will continue being used" and kept running. Without it the
    reload path calls panic() and one malformed contributed file takes the
    monitor down.

Verified: 15/15 service endpoints UP; 26/27 public-facing UP. The single DOWN is
public_memos, correctly - memos-box is powered off, and the condition result
reads [STATUS] (502) == 200. memos-box and arbret-staging-box were shut down
deliberately from the Proxmox UI to test the liveness endpoints; their endpoints
stay registered and will go green when the VMs return.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 09:33:57 +02:00

82 lines
3.1 KiB
Django/Jinja

#!/bin/bash
# {{ healthcheck_name }}{{ healthcheck_description }}
# Managed by Ansible (roles/healthcheck). Do not edit on the host.
#
# The exit code is the real answer; systemd stores it:
# systemctl is-failed {{ healthcheck_name }}-healthcheck.service
# The push below is an optional extra, and having no URL is normal.
set -uo pipefail
LOG_FILE="{{ healthcheck_log_dir }}/{{ healthcheck_name }}.log"
PUSH_URL="${HEALTHCHECK_PUSH_URL:-}"
PUSH_TOKEN="${HEALTHCHECK_PUSH_TOKEN:-}"
# Set only by checks that report several results; see healthcheck_push_base.
PUSH_BASE="${HEALTHCHECK_PUSH_BASE:-}"
log() { echo "$(date '+%Y-%m-%d %H:%M:%S') - $*" >> "$LOG_FILE"; }
# Report to Gatus as an external endpoint. Note this is a POST with a bearer
# token, not a GET with a query string - it is not the shape Uptime Kuma used.
report() {
local success="$1" message="$2"
# No push URL is normal, not an error: the exit code below is a complete
# answer for anything reading unit state.
[ -n "$PUSH_URL" ] || return 0
local encoded
encoded=$(printf '%s' "$message" | sed 's/%/%25/g; s/ /%20/g; s/&/%26/g; s/+/%2B/g; s/#/%23/g')
local code
code=$(curl -s -o /dev/null -w '%{http_code}' -X POST \
--max-time 15 --retry 2 --retry-delay 3 \
-H "Authorization: Bearer ${PUSH_TOKEN}" \
"${PUSH_URL}?success=${success}&error=${encoded}" 2>/dev/null)
if [ "$code" = "200" ]; then
log "reported success=${success}"
else
log "ERROR: report failed (HTTP ${code})"
return 1
fi
}
# Report to an arbitrary endpoint key under PUSH_BASE. Used by checks that
# produce one result per item rather than a single verdict.
report_key() {
local key="$1" success="$2" message="$3"
[ -n "$PUSH_BASE" ] || return 0
local encoded
encoded=$(printf '%s' "$message" | sed 's/%/%25/g; s/ /%20/g; s/&/%26/g; s/+/%2B/g; s/#/%23/g')
curl -s -o /dev/null --max-time 15 --retry 2 --retry-delay 3 -X POST \
-H "Authorization: Bearer ${PUSH_TOKEN}" \
"${PUSH_BASE}/${key}/external?success=${success}&error=${encoded}" 2>/dev/null || true
}
# ── the check ────────────────────────────────────────────────────────────────
#
# The blank line before the closing brace below is load-bearing. Jinja strips an
# included template's trailing newline, and trim_blocks (on by default in
# Ansible) then eats the newline after {% raw %}{% endif %}{% endraw %} - so without it the brace lands
# on the same line as the check body's last statement, producing `return 0}` and
# a script that dies with "syntax error: unexpected end of file".
check() {
{% if healthcheck_check %}
{% include 'checks/' ~ healthcheck_check ~ '.sh.j2' %}
{% else %}
{{ healthcheck_command }}
{% endif %}
}
MESSAGE=""
if check; then
log "OK${MESSAGE:+ - $MESSAGE}"
report "true" "${MESSAGE:-ok}"
exit 0
else
log "FAILED${MESSAGE:+ - $MESSAGE}"
report "false" "${MESSAGE:-check failed}"
exit 1
fi