monitoring: recover host checks for the whole estate, reported to Gatus
Five checks, 27 endpoints, replacing what Uptime Kuma used to watch:
is it up every 5min, all hosts
is disk full daily, all hosts
is CPU hot every 5min, nodito
is ZFS broken daily, nodito
is UPS online every 5min, nodito
Two roles, kept separate so neither knows about the other - they meet at a URL
and a token, the same way caddy_site and each service meet at a vhost:
roles/gatus_endpoint runs on the observability host, writes ONE file into
/opt/gatus/config/endpoints/. Gatus merges every *.yaml
there and appends lists, so callers compose without
coordinating.
roles/healthcheck runs on the monitored host: a check script, a systemd
service, a timer, and an optional push. Ships a library
of check bodies under templates/checks/.
Everything PUSHES. Gatus never reaches out, which matters because nodito and its
VMs are behind NAT, and because four of the five checks are internal state with
no pollable surface at all. Liveness pushes too, deliberately: a heartbeat proves
the host is running AND can reach the internet, where an ICMP probe from one
vantage point only proves it answers pings from there. And since Gatus alerts
when a heartbeat window expires, a check that stops running raises the alarm by
itself - a dead timer looks exactly like a dead host, which is the correct
reading.
One bearer token per host, generated straight into the vault and never printed.
A token only writes results for its own host's endpoints, so a compromised host
can lie about itself, which it could do anyway.
Three things learned from the source that shaped this:
* Gatus polls its own config every 30s and reloads
(main.listenToConfigurationFileChanges), so gatus_endpoint needs no restart
handler - writing the file IS the deploy.
* ...but on a reload it panics if the new config fails to parse, unless
skip-invalid-config-update is set. Endpoint files are contributed by other
playbooks, so one malformed file would take the monitor down at the worst
possible moment. Now set.
* The push URL uses a key Gatus computes, not the name you write:
sanitize(group) + "_" + sanitize(name), lowercased with / _ . , space # + &
replaced by "-" (config/key/key.go). So knots_box_local is
knots-box-local in the URL. The playbook derives it rather than hand-writing.
storage: maximum-number-of-results 900, up from upstream's 100. Gatus bounds the
database by COUNT and trims inline on insert, so there is no retention job and
no way to fill a disk - but history depth is then a function of check frequency,
and 100 results at a 5-minute interval is 8 hours. 900 is ~3 days of liveness
and ~2.5 years of the daily disk check. The uptime table is separate and its
30-day retention is hard-coded upstream.
A bug worth recording: the first deploy shipped five scripts that all died with
"syntax error: unexpected end of file". Jinja strips an included template's
trailing newline and trim_blocks then eats the newline after {% endif %}, so the
closing brace of check() landed on the same line as the body's last statement -
`return 0}`. Every check was broken and the deploy still reported failed=0,
because the role's "run once" task has failed_when: false and reports the result
as a debug message nobody read. The blank line that fixes it is now load-bearing
and commented as such.
Verified by triggering every unit by hand rather than waiting on timers: all
checks exit 0 on all hosts, and Gatus shows 26 UP / 1 DOWN. The one DOWN is
liveness_watchtower, which is honest - that host currently refuses SSH (TCP
connects, no banner exchange) and is excluded from this deploy. It is also the
box still running Uptime Kuma.
Known waste, not yet fixed: healthcheck installs its dependencies per CHECK
rather than per HOST, so apt runs 29 times estate-wide for a curl that is
already present, and daemon_reload runs 4x per host.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
parent
fa9f7d10cd
commit
ede407ebe4
17 changed files with 887 additions and 200 deletions
|
|
@ -57,6 +57,21 @@ gatus_storage_type: sqlite
|
|||
gatus_storage_path: "/data/gatus.db"
|
||||
gatus_storage_caching: true
|
||||
|
||||
# Per-endpoint row caps. Gatus bounds the database by COUNT, not by time, and
|
||||
# trims inline on insert (storage/store/sql/sql.go, InsertEndpointResult) - so
|
||||
# there is no retention job to write and no way for this to fill a disk.
|
||||
#
|
||||
# History depth is therefore a function of check frequency, not of days:
|
||||
# 900 results is ~3 days of a 5-minute liveness check, and ~2.5 years of a daily
|
||||
# disk check. Upstream's default is 100, which would have been 8 hours of
|
||||
# liveness - not enough to still see a weekend incident on Monday.
|
||||
#
|
||||
# Note the uptime table is separate and its 30-day retention is hard-coded
|
||||
# upstream (uptimeRetention), so uptime percentages top out at 30 days whatever
|
||||
# this is set to.
|
||||
gatus_storage_max_results: 900
|
||||
gatus_storage_max_events: 50
|
||||
|
||||
# ── Alerting ─────────────────────────────────────────────────────────────────
|
||||
# Pass-through: rendered verbatim under `alerting:`, so any provider Gatus
|
||||
# supports works without touching this role. Empty means "check and record,
|
||||
|
|
@ -69,6 +84,10 @@ gatus_default_alerts: []
|
|||
gatus_basic_auth: {}
|
||||
|
||||
gatus_maintenance: {}
|
||||
|
||||
# Keep serving the previous config if a contributed endpoint file is malformed,
|
||||
# instead of panicking. See the note in config.yaml.j2.
|
||||
gatus_skip_invalid_config_update: true
|
||||
gatus_log_level: INFO
|
||||
|
||||
# ── Self-check ───────────────────────────────────────────────────────────────
|
||||
|
|
|
|||
|
|
@ -23,12 +23,21 @@ web:
|
|||
address: 0.0.0.0
|
||||
port: {{ gatus_port }}
|
||||
|
||||
# Gatus reloads when its config changes. If the NEW config fails to parse it
|
||||
# calls panic() - unless this is set, in which case it logs the error and keeps
|
||||
# running on the old config. Endpoint files are contributed by other playbooks,
|
||||
# so one malformed file would otherwise take the monitor down, which is the
|
||||
# worst possible time to lose it.
|
||||
skip-invalid-config-update: {{ gatus_skip_invalid_config_update | bool | lower }}
|
||||
|
||||
storage:
|
||||
type: {{ gatus_storage_type }}
|
||||
{% if gatus_storage_type != 'memory' %}
|
||||
path: {{ gatus_storage_path }}
|
||||
{% endif %}
|
||||
caching: {{ gatus_storage_caching | bool | lower }}
|
||||
maximum-number-of-results: {{ gatus_storage_max_results }}
|
||||
maximum-number-of-events: {{ gatus_storage_max_events }}
|
||||
|
||||
ui:
|
||||
title: {{ gatus_ui_title }}
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue