personal_infra/ansible/infra/400_host_monitoring.yml

176 lines
7 KiB
YAML
Raw Normal View History

monitoring: recover host checks for the whole estate, reported to Gatus Five checks, 27 endpoints, replacing what Uptime Kuma used to watch: is it up every 5min, all hosts is disk full daily, all hosts is CPU hot every 5min, nodito is ZFS broken daily, nodito is UPS online every 5min, nodito Two roles, kept separate so neither knows about the other - they meet at a URL and a token, the same way caddy_site and each service meet at a vhost: roles/gatus_endpoint runs on the observability host, writes ONE file into /opt/gatus/config/endpoints/. Gatus merges every *.yaml there and appends lists, so callers compose without coordinating. roles/healthcheck runs on the monitored host: a check script, a systemd service, a timer, and an optional push. Ships a library of check bodies under templates/checks/. Everything PUSHES. Gatus never reaches out, which matters because nodito and its VMs are behind NAT, and because four of the five checks are internal state with no pollable surface at all. Liveness pushes too, deliberately: a heartbeat proves the host is running AND can reach the internet, where an ICMP probe from one vantage point only proves it answers pings from there. And since Gatus alerts when a heartbeat window expires, a check that stops running raises the alarm by itself - a dead timer looks exactly like a dead host, which is the correct reading. One bearer token per host, generated straight into the vault and never printed. A token only writes results for its own host's endpoints, so a compromised host can lie about itself, which it could do anyway. Three things learned from the source that shaped this: * Gatus polls its own config every 30s and reloads (main.listenToConfigurationFileChanges), so gatus_endpoint needs no restart handler - writing the file IS the deploy. * ...but on a reload it panics if the new config fails to parse, unless skip-invalid-config-update is set. Endpoint files are contributed by other playbooks, so one malformed file would take the monitor down at the worst possible moment. Now set. * The push URL uses a key Gatus computes, not the name you write: sanitize(group) + "_" + sanitize(name), lowercased with / _ . , space # + & replaced by "-" (config/key/key.go). So knots_box_local is knots-box-local in the URL. The playbook derives it rather than hand-writing. storage: maximum-number-of-results 900, up from upstream's 100. Gatus bounds the database by COUNT and trims inline on insert, so there is no retention job and no way to fill a disk - but history depth is then a function of check frequency, and 100 results at a 5-minute interval is 8 hours. 900 is ~3 days of liveness and ~2.5 years of the daily disk check. The uptime table is separate and its 30-day retention is hard-coded upstream. A bug worth recording: the first deploy shipped five scripts that all died with "syntax error: unexpected end of file". Jinja strips an included template's trailing newline and trim_blocks then eats the newline after {% endif %}, so the closing brace of check() landed on the same line as the body's last statement - `return 0}`. Every check was broken and the deploy still reported failed=0, because the role's "run once" task has failed_when: false and reports the result as a debug message nobody read. The blank line that fixes it is now load-bearing and commented as such. Verified by triggering every unit by hand rather than waiting on timers: all checks exit 0 on all hosts, and Gatus shows 26 UP / 1 DOWN. The one DOWN is liveness_watchtower, which is honest - that host currently refuses SSH (TCP connects, no banner exchange) and is excluded from this deploy. It is also the box still running Uptime Kuma. Known waste, not yet fixed: healthcheck installs its dependencies per CHECK rather than per HOST, so apt runs 29 times estate-wide for a curl that is already present, and daemon_reload runs 4x per host. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 08:54:10 +02:00
---
# Host-level monitoring for the whole estate, reported to Gatus.
#
# Every check here PUSHES. Gatus never reaches out, which matters because nodito
# and its VMs sit behind NAT, and because four of the five checks are internal
# state with no pollable surface at all - disk usage, CPU temperature, ZFS pool
# health and UPS mains status cannot be observed from outside the machine.
#
# Liveness is a push too, and that is a choice rather than a limitation. A
# heartbeat proves the host is running AND can reach the internet; an ICMP probe
# from one vantage point only proves it answers pings from there. And because
# Gatus alerts when a heartbeat window expires, a check that stops running
# raises the alarm by itself - a dead timer looks exactly like a dead host,
# which is the correct reading.
#
# Each host has ONE bearer token, shared across its own checks: a token can only
# write results for that host's endpoints, so a compromised host can lie about
# itself, which it could do anyway.
#
# The push URL must use Gatus's own key format (config/key/key.go):
# key = sanitize(group) + "_" + sanitize(name)
# where sanitize lowercases and replaces / _ . , space # + & with "-". So
# knots_box_local becomes knots-box-local in the URL but stays readable in the
# name. host_key below is the Jinja equivalent; do not hand-write these.
# ─────────────────────────────────────────────────────────────────────────────
# Register everything with Gatus.
#
# This play runs FIRST on purpose. Gatus reloads its config within 30s, and the
# host plays below take minutes, so every endpoint exists before its first push
# arrives. Registering afterwards would 404 every first report.
#
# Heartbeat windows are several times the check interval, so one missed run - a
# slow apt run, a reboot - does not raise an alarm, but a check that has
# genuinely stopped does.
# ─────────────────────────────────────────────────────────────────────────────
- name: Register the host checks with Gatus
hosts: observability
become: yes
vars:
monitored: "{{ groups['managed'] | sort }}"
tasks:
- name: Build the liveness endpoint list
ansible.builtin.set_fact:
liveness_endpoints: "{{ liveness_endpoints | default([]) + [{
'name': item,
'group': 'liveness',
'token': gatus_push_tokens[item],
'heartbeat': '16m'}] }}"
loop: "{{ monitored }}"
- name: Build the disk endpoint list
ansible.builtin.set_fact:
disk_endpoints: "{{ disk_endpoints | default([]) + [{
'name': item,
'group': 'disk',
'token': gatus_push_tokens[item],
'heartbeat': '30h'}] }}"
loop: "{{ monitored }}"
- name: Register liveness endpoints
ansible.builtin.include_role:
name: gatus_endpoint
vars:
gatus_endpoint_name: liveness
gatus_endpoint_external: "{{ liveness_endpoints }}"
- name: Register disk endpoints
ansible.builtin.include_role:
name: gatus_endpoint
vars:
gatus_endpoint_name: disk
gatus_endpoint_external: "{{ disk_endpoints }}"
- name: Register the hypervisor endpoints
ansible.builtin.include_role:
name: gatus_endpoint
vars:
gatus_endpoint_name: hypervisor
gatus_endpoint_external:
- name: cpu
group: hypervisor
token: "{{ gatus_push_tokens['nodito'] }}"
heartbeat: "16m"
- name: zfs
group: hypervisor
token: "{{ gatus_push_tokens['nodito'] }}"
heartbeat: "30h"
- name: ups
group: hypervisor
token: "{{ gatus_push_tokens['nodito'] }}"
heartbeat: "16m"
- name: Deploy host liveness and disk checks
hosts: managed
become: yes
vars:
gatus_api: "https://{{ subdomains.gatus }}.{{ root_domain }}/api/v1/endpoints"
host_key: "{{ inventory_hostname | lower | regex_replace('[/_.,# +&]', '-') }}"
host_token: "{{ gatus_push_tokens[inventory_hostname] }}"
tasks:
- name: Is the host up?
ansible.builtin.include_role:
name: healthcheck
vars:
healthcheck_name: liveness
healthcheck_description: "Liveness heartbeat for {{ inventory_hostname }}"
healthcheck_check: liveness
healthcheck_interval: "5min"
healthcheck_boot_delay: "1min"
healthcheck_push_url: "{{ gatus_api }}/liveness_{{ host_key }}/external"
healthcheck_push_token: "{{ host_token }}"
- name: Is the disk packed?
ansible.builtin.include_role:
name: healthcheck
vars:
healthcheck_name: disk-usage
healthcheck_description: "Disk usage for {{ inventory_hostname }}"
healthcheck_check: disk-usage
# Daily. RandomizedDelaySec spreads twelve hosts across the hour rather
# than having them all report in the same second.
healthcheck_on_calendar: "*-*-* 07:00:00"
healthcheck_randomized_delay: "3600"
healthcheck_boot_delay: "5min"
healthcheck_push_url: "{{ gatus_api }}/disk_{{ host_key }}/external"
healthcheck_push_token: "{{ host_token }}"
- name: Deploy the hypervisor-only checks
hosts: hypervisor
become: yes
vars:
gatus_api: "https://{{ subdomains.gatus }}.{{ root_domain }}/api/v1/endpoints"
host_token: "{{ gatus_push_tokens[inventory_hostname] }}"
tasks:
- name: Is the CPU hot?
ansible.builtin.include_role:
name: healthcheck
vars:
healthcheck_name: cpu-temp
healthcheck_description: "CPU temperature for {{ inventory_hostname }}"
healthcheck_check: cpu-temp
healthcheck_packages: [curl, lm-sensors]
healthcheck_interval: "5min"
healthcheck_push_url: "{{ gatus_api }}/hypervisor_cpu/external"
healthcheck_push_token: "{{ host_token }}"
- name: Is ZFS broken?
ansible.builtin.include_role:
name: healthcheck
vars:
healthcheck_name: zfs-health
healthcheck_description: "ZFS pool health for {{ zfs_pool_name }}"
healthcheck_check: zfs-health
healthcheck_packages: [curl, jq]
healthcheck_zfs_pool: "{{ zfs_pool_name }}"
healthcheck_on_calendar: "*-*-* 07:20:00"
healthcheck_boot_delay: "10min"
healthcheck_push_url: "{{ gatus_api }}/hypervisor_zfs/external"
healthcheck_push_token: "{{ host_token }}"
- name: Is the UPS online?
ansible.builtin.include_role:
name: healthcheck
vars:
healthcheck_name: ups-status
healthcheck_description: "UPS mains status for {{ ups_name }}"
healthcheck_check: ups-status
healthcheck_ups_name: "{{ ups_name }}"
healthcheck_interval: "5min"
healthcheck_push_url: "{{ gatus_api }}/hypervisor_ups/external"
healthcheck_push_token: "{{ host_token }}"