monitoring: systemd services, domain expiry, DNS correctness, public endpoints
Four more check types, 42 endpoints, taking the estate from 39 to 81.
── systemd services (infra/401) ────────────────────────────────────────────
Every unit we deploy, checked every 5 minutes with a 16-minute heartbeat.
This closes the gap that let a real bug run unnoticed earlier today: a backup
script left forgejo, lnbits, headscale and memos stopped, and NOTHING caught
it. The dumps exited 0, the artefacts were correct, the deploy said failed=0,
and liveness only proves the HOST is up - not that anything on it serves.
One endpoint PER UNIT but only ONE timer per host: the check iterates that
host's units and pushes a result for each, the way check-backups.sh reports per
source. Four units on vipy would otherwise mean four scripts, services and
timers. A single host-level red light would also say "something on vipy is
down" without saying which, which is not the question anyone has.
Keys are host-qualified because unit names collide - caddy runs on four
machines. Which units a host runs lives in host_vars/<host>/monitored_services,
because "what runs here" is a property of the machine, the same reasoning as
the cross-host ports.
nut-driver-enumerator is deliberately excluded: it is a oneshot generator that
is `enabled` but always `inactive`, so it would report down forever. Checked
live before excluding it.
── domain, DNS and public endpoints (infra/402) ────────────────────────────
The first checks that PULL rather than push, and that is the right way round:
all three are about how the outside world sees us, so they must be measured
from outside. Nothing is installed anywhere - no script, no timer, no token.
They also have no heartbeat, because a heartbeat answers "did the reporter
report"; when Gatus does the checking itself, failure is immediate.
domain 1 endpoint, 24h, [DOMAIN_EXPIRATION] > 336h (two weeks)
dns 11 endpoints, 24h, [DNS_RCODE] == NOERROR and [BODY] == the IP
public 14 endpoints, 5m, 11 HTTPS + 3 TCP
Expected A records are derived from inventory (hostvars[host].ansible_host),
not written down again. The estate's recurring bug is an address recorded in a
second place and left behind when the machine moved; asserting against
inventory means a renumbered box is one edit, not two.
Expected HTTP status was checked live per site rather than assumed. Two return
401 - the Gatus dashboard and the DATUM dashboard, both behind basic auth - and
that is what is asserted: expecting 200 there would go green precisely when the
auth broke. headscale asserts /health rather than /, which is a 404 by design.
Every HTTPS check also carries [CERTIFICATE_EXPIRATION] > 168h, which is free
on an endpoint already being polled and catches a renewal that silently stops.
Two things learned the hard way, both now in comments:
* A domain-expiry endpoint needs a URL SCHEME. Gatus derives the endpoint type
from the prefix (endpoint.Type()), so a bare "contrapeso.xyz" is UNKNOWN and
the whole config is rejected. It is https:// plus a DOMAIN_EXPIRATION
condition and no status assertion, so the registrar's parking page at the
apex is irrelevant. Upstream also enforces a 5m minimum interval for that
placeholder, because it uses a free whois service.
* That rejection proved skip-invalid-config-update was worth adding. Gatus
logged "the configuration file was updated, but it is not valid, the old
configuration will continue being used" and kept running. Without it the
reload path calls panic() and one malformed contributed file takes the
monitor down.
Verified: 15/15 service endpoints UP; 26/27 public-facing UP. The single DOWN is
public_memos, correctly - memos-box is powered off, and the condition result
reads [STATUS] (502) == 200. memos-box and arbret-staging-box were shut down
deliberately from the Proxmox UI to test the liveness endpoints; their endpoints
stay registered and will go green when the VMs return.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 09:33:57 +02:00
|
|
|
---
|
|
|
|
|
# Domain expiry, DNS correctness, and public endpoint reachability.
|
|
|
|
|
#
|
|
|
|
|
# These are the first checks in the estate that PULL rather than push, and that
|
|
|
|
|
# is the right way round for them: all three are about how the outside world
|
|
|
|
|
# sees us, so they must be measured from outside. Gatus polls from the
|
|
|
|
|
# observability host and needs nothing installed anywhere else - there is no
|
|
|
|
|
# script, no timer and no token, because nothing is reporting in.
|
|
|
|
|
#
|
|
|
|
|
# That also means these have no heartbeat. A heartbeat answers "did the thing
|
|
|
|
|
# that was supposed to report in do so"; when Gatus does the checking itself,
|
|
|
|
|
# failure is immediate and self-evident.
|
|
|
|
|
|
|
|
|
|
- name: Register the public-facing checks with Gatus
|
|
|
|
|
hosts: observability
|
|
|
|
|
become: yes
|
|
|
|
|
|
|
|
|
|
vars:
|
|
|
|
|
# Expected A records, derived from inventory rather than written down again.
|
|
|
|
|
# The estate's recurring bug is an address recorded in a second place and
|
|
|
|
|
# then left behind when the machine moved, so the check asserts against
|
|
|
|
|
# ansible_host - if a box is renumbered, inventory is the one edit.
|
|
|
|
|
dns_records:
|
|
|
|
|
- {sub: "{{ subdomains.gatus }}", host: monitoring}
|
|
|
|
|
- {sub: "{{ subdomains.headscale }}", host: spacey}
|
|
|
|
|
- {sub: "{{ subdomains.vaultwarden }}", host: vipy}
|
|
|
|
|
- {sub: "{{ subdomains.forgejo }}", host: vipy}
|
|
|
|
|
- {sub: "{{ subdomains.lnbits }}", host: vipy}
|
|
|
|
|
- {sub: "{{ subdomains.ntfy_emergency_app }}", host: vipy}
|
|
|
|
|
- {sub: "{{ subdomains.personal_blog }}", host: vipy}
|
|
|
|
|
- {sub: "{{ subdomains.memos }}", host: vipy}
|
|
|
|
|
- {sub: "{{ subdomains.mempool }}", host: vipy}
|
|
|
|
|
- {sub: "{{ subdomains.datum_gateway }}", host: vipy}
|
|
|
|
|
|
|
|
|
|
# A public resolver on purpose: this must test what the internet sees, not
|
|
|
|
|
# what a local cache or the tailnet's MagicDNS happens to answer.
|
|
|
|
|
dns_resolver: "1.1.1.1"
|
|
|
|
|
|
|
|
|
|
# Expected status per site, checked live before being written down.
|
|
|
|
|
# 401 is the CORRECT answer for the two behind basic auth - asserting 200
|
|
|
|
|
# there would go green precisely when the auth broke.
|
|
|
|
|
public_sites:
|
|
|
|
|
- {name: gatus, sub: "{{ subdomains.gatus }}", path: "/", status: 401}
|
|
|
|
|
- {name: headscale, sub: "{{ subdomains.headscale }}", path: "/health", status: 200}
|
|
|
|
|
- {name: vaultwarden, sub: "{{ subdomains.vaultwarden }}", path: "/", status: 200}
|
|
|
|
|
- {name: forgejo, sub: "{{ subdomains.forgejo }}", path: "/", status: 200}
|
|
|
|
|
- {name: lnbits, sub: "{{ subdomains.lnbits }}", path: "/", status: 200}
|
|
|
|
|
- {name: avisame, sub: "{{ subdomains.ntfy_emergency_app }}", path: "/", status: 200}
|
|
|
|
|
- {name: blog, sub: "{{ subdomains.personal_blog }}", path: "/", status: 200}
|
|
|
|
|
- {name: memos, sub: "{{ subdomains.memos }}", path: "/", status: 200}
|
|
|
|
|
- {name: mempool, sub: "{{ subdomains.mempool }}", path: "/", status: 200}
|
|
|
|
|
- {name: datum, sub: "{{ subdomains.datum_gateway }}", path: "/", status: 401}
|
|
|
|
|
|
|
|
|
|
# Ports published from the edge host by socket_proxy.
|
|
|
|
|
public_tcp:
|
|
|
|
|
- {name: bitcoin-p2p, host: vipy, port: "{{ hostvars['knots_box_local'].bitcoin_p2p_port }}"}
|
|
|
|
|
- {name: fulcrum-ssl, host: vipy, port: "{{ hostvars['fulcrum_box_local'].fulcrum_ssl_port }}"}
|
|
|
|
|
- {name: datum-stratum, host: vipy, port: "{{ hostvars['knots_box_local'].datum_gateway_stratum_port }}"}
|
|
|
|
|
|
|
|
|
|
tasks:
|
|
|
|
|
# ── Domain expiry ────────────────────────────────────────────────────────
|
2026-09-14 09:53:15 +02:00
|
|
|
# Each domain needs a URL SCHEME: Gatus derives the endpoint type from the
|
|
|
|
|
# prefix (endpoint.Type()), so a bare "example.com" is UNKNOWN and the whole
|
|
|
|
|
# config is rejected. No status is asserted, only the WHOIS/RDAP expiry, so
|
|
|
|
|
# whatever the apex serves - a real site, or the registrar's parking page -
|
|
|
|
|
# is irrelevant.
|
|
|
|
|
#
|
|
|
|
|
# 24h, and upstream enforces a 5m minimum for DOMAIN_EXPIRATION anyway
|
|
|
|
|
# because it uses a free whois service that must not be hammered.
|
|
|
|
|
# 336h = 14 days of runway, because renewal is a manual act at the registrar.
|
|
|
|
|
- name: Build the domain endpoints
|
monitoring: systemd services, domain expiry, DNS correctness, public endpoints
Four more check types, 42 endpoints, taking the estate from 39 to 81.
── systemd services (infra/401) ────────────────────────────────────────────
Every unit we deploy, checked every 5 minutes with a 16-minute heartbeat.
This closes the gap that let a real bug run unnoticed earlier today: a backup
script left forgejo, lnbits, headscale and memos stopped, and NOTHING caught
it. The dumps exited 0, the artefacts were correct, the deploy said failed=0,
and liveness only proves the HOST is up - not that anything on it serves.
One endpoint PER UNIT but only ONE timer per host: the check iterates that
host's units and pushes a result for each, the way check-backups.sh reports per
source. Four units on vipy would otherwise mean four scripts, services and
timers. A single host-level red light would also say "something on vipy is
down" without saying which, which is not the question anyone has.
Keys are host-qualified because unit names collide - caddy runs on four
machines. Which units a host runs lives in host_vars/<host>/monitored_services,
because "what runs here" is a property of the machine, the same reasoning as
the cross-host ports.
nut-driver-enumerator is deliberately excluded: it is a oneshot generator that
is `enabled` but always `inactive`, so it would report down forever. Checked
live before excluding it.
── domain, DNS and public endpoints (infra/402) ────────────────────────────
The first checks that PULL rather than push, and that is the right way round:
all three are about how the outside world sees us, so they must be measured
from outside. Nothing is installed anywhere - no script, no timer, no token.
They also have no heartbeat, because a heartbeat answers "did the reporter
report"; when Gatus does the checking itself, failure is immediate.
domain 1 endpoint, 24h, [DOMAIN_EXPIRATION] > 336h (two weeks)
dns 11 endpoints, 24h, [DNS_RCODE] == NOERROR and [BODY] == the IP
public 14 endpoints, 5m, 11 HTTPS + 3 TCP
Expected A records are derived from inventory (hostvars[host].ansible_host),
not written down again. The estate's recurring bug is an address recorded in a
second place and left behind when the machine moved; asserting against
inventory means a renumbered box is one edit, not two.
Expected HTTP status was checked live per site rather than assumed. Two return
401 - the Gatus dashboard and the DATUM dashboard, both behind basic auth - and
that is what is asserted: expecting 200 there would go green precisely when the
auth broke. headscale asserts /health rather than /, which is a 404 by design.
Every HTTPS check also carries [CERTIFICATE_EXPIRATION] > 168h, which is free
on an endpoint already being polled and catches a renewal that silently stops.
Two things learned the hard way, both now in comments:
* A domain-expiry endpoint needs a URL SCHEME. Gatus derives the endpoint type
from the prefix (endpoint.Type()), so a bare "contrapeso.xyz" is UNKNOWN and
the whole config is rejected. It is https:// plus a DOMAIN_EXPIRATION
condition and no status assertion, so the registrar's parking page at the
apex is irrelevant. Upstream also enforces a 5m minimum interval for that
placeholder, because it uses a free whois service.
* That rejection proved skip-invalid-config-update was worth adding. Gatus
logged "the configuration file was updated, but it is not valid, the old
configuration will continue being used" and kept running. Without it the
reload path calls panic() and one malformed contributed file takes the
monitor down.
Verified: 15/15 service endpoints UP; 26/27 public-facing UP. The single DOWN is
public_memos, correctly - memos-box is powered off, and the condition result
reads [STATUS] (502) == 200. memos-box and arbret-staging-box were shut down
deliberately from the Proxmox UI to test the liveness endpoints; their endpoints
stay registered and will go green when the VMs return.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 09:33:57 +02:00
|
|
|
ansible.builtin.set_fact:
|
2026-09-14 09:53:15 +02:00
|
|
|
domain_endpoints: "{{ domain_endpoints | default([]) + [{
|
|
|
|
|
'name': item,
|
|
|
|
|
'group': 'domain',
|
|
|
|
|
'url': 'https://' ~ item,
|
|
|
|
|
'interval': '24h',
|
alerting: Signal via signal-cli-rest-api, and faster failure detection
Three changes: detection windows tightened, a Signal transport deployed, and
alerts attached to all 84 non-transport endpoints.
── Failing faster ──────────────────────────────────────────────────────────
The heartbeat is a TICKER, not a deadline: Gatus wakes every interval and asks
"did anything arrive in the last interval", so real detection is 1-2x the
window. And a window can only ever be as tight as the push frequency - which is
why two checks moved rather than just having their numbers changed.
liveness, cpu, ups, service-health, probes 16m -> 11m
disk, zfs daily/30h -> 6-hourly/7h
backup store + pull job daily/30h -> 6-hourly/7h
DNS records 24h -> 6h
backup dump 30h -> 26h
backup-dump stays slow because the dump genuinely is daily. The store-side check
catches the same fault within 6h by reading the source's dump timestamp out of
the artefact filename, so 26h is a backstop rather than the primary signal.
── Thresholds differ by check type, deliberately ───────────────────────────
`failure-threshold` counts consecutive failures, but "consecutive" is a
different amount of wall-clock time per check: a push endpoint produces one
failure per heartbeat window, a pulled one per interval. The default of 3 would
mean 33 minutes on an 11m heartbeat and over a day on a 7h one.
More importantly the heartbeat window ALREADY encodes the tolerance - an 11m
window on a 5-minute push is exactly "one missed push forgiven" - so stacking a
threshold of 3 triples a tolerance that was already chosen. Hence:
push/heartbeat endpoints failure-threshold 1
pulled, 5m (public) failure-threshold 3 (= 15 minutes)
pulled, 6h/24h (dns, domain) failure-threshold 1
── The Signal transport ────────────────────────────────────────────────────
roles/signal_api runs signal-cli-rest-api on the observability host, pinned by
digest, MODE=native.
It publishes NO PORTS. The API has no authentication of any kind - anything that
reaches it can send messages as you and read your Signal. Gatus talks to it over
a shared docker network by service name, which is also WHY the network exists:
Gatus runs in a container, so the host's loopback is unreachable from it and a
port published on 127.0.0.1 would not have worked.
MODE=native and not json-rpc because this VPS has 464MB of RAM and already runs
Gatus and Caddy. The json-rpc modes hold a resident JVM; native runs a binary
per request, and alerts are rare enough that startup cost per alert is the right
trade.
Monitored - Gatus polls /v1/health over the same network path the alerts take,
so it proves the delivery route rather than mere container liveness. Deliberately
NOT backed up: the data directory holds Signal private keys and the recovery
path is to link the device again from the phone.
Neither the signal-api endpoint nor Gatus's self-check carries a Signal alert.
If either is down, Signal is precisely what cannot deliver the alert.
── Four traps hit while linking, all now in the role README ────────────────
* /v1/qrcodelink is BROKEN in native mode - returns "no data to encode" while
the binary itself emits a perfectly good URI. Not worked around by switching
MODE, which would put a JVM in the path of every alert permanently.
* The data dir must be owned by uid 1000, not root. A root-owned 0700 dir
cannot be traversed by the container user, so linking silently never
completes and /v1/accounts returns "Failed to read local accounts list".
* `docker exec` runs as ROOT while the service runs as uid 1000, so without
--config the account is written to /root/... on the container's ephemeral
layer. It reports success and is destroyed on the next recreate.
* The phone reporting "network error" was IPv6: chat.signal.org resolves to
dualstack AAAA records first, the container has no IPv6 address, and this
host's IPv6 path is broken - the same edge that 404'd the Go tarball. Fixed
with a mounted gai.conf that prefers IPv4.
Verified: provider loads (configuredProviders=[signal]), a test message was
delivered and confirmed received, 84 endpoints carry alerts, 86 UP / 0 DOWN.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 21:40:37 +02:00
|
|
|
'conditions': ['[DOMAIN_EXPIRATION] > 336h'],
|
|
|
|
|
'alerts': [{'type': 'signal', 'failure-threshold': 1,
|
|
|
|
|
'success-threshold': 1, 'send-on-resolved': true,
|
|
|
|
|
'minimum-reminder-interval': '168h'}]}] }}"
|
2026-09-14 09:53:15 +02:00
|
|
|
loop: "{{ monitored_domains }}"
|
monitoring: systemd services, domain expiry, DNS correctness, public endpoints
Four more check types, 42 endpoints, taking the estate from 39 to 81.
── systemd services (infra/401) ────────────────────────────────────────────
Every unit we deploy, checked every 5 minutes with a 16-minute heartbeat.
This closes the gap that let a real bug run unnoticed earlier today: a backup
script left forgejo, lnbits, headscale and memos stopped, and NOTHING caught
it. The dumps exited 0, the artefacts were correct, the deploy said failed=0,
and liveness only proves the HOST is up - not that anything on it serves.
One endpoint PER UNIT but only ONE timer per host: the check iterates that
host's units and pushes a result for each, the way check-backups.sh reports per
source. Four units on vipy would otherwise mean four scripts, services and
timers. A single host-level red light would also say "something on vipy is
down" without saying which, which is not the question anyone has.
Keys are host-qualified because unit names collide - caddy runs on four
machines. Which units a host runs lives in host_vars/<host>/monitored_services,
because "what runs here" is a property of the machine, the same reasoning as
the cross-host ports.
nut-driver-enumerator is deliberately excluded: it is a oneshot generator that
is `enabled` but always `inactive`, so it would report down forever. Checked
live before excluding it.
── domain, DNS and public endpoints (infra/402) ────────────────────────────
The first checks that PULL rather than push, and that is the right way round:
all three are about how the outside world sees us, so they must be measured
from outside. Nothing is installed anywhere - no script, no timer, no token.
They also have no heartbeat, because a heartbeat answers "did the reporter
report"; when Gatus does the checking itself, failure is immediate.
domain 1 endpoint, 24h, [DOMAIN_EXPIRATION] > 336h (two weeks)
dns 11 endpoints, 24h, [DNS_RCODE] == NOERROR and [BODY] == the IP
public 14 endpoints, 5m, 11 HTTPS + 3 TCP
Expected A records are derived from inventory (hostvars[host].ansible_host),
not written down again. The estate's recurring bug is an address recorded in a
second place and left behind when the machine moved; asserting against
inventory means a renumbered box is one edit, not two.
Expected HTTP status was checked live per site rather than assumed. Two return
401 - the Gatus dashboard and the DATUM dashboard, both behind basic auth - and
that is what is asserted: expecting 200 there would go green precisely when the
auth broke. headscale asserts /health rather than /, which is a 404 by design.
Every HTTPS check also carries [CERTIFICATE_EXPIRATION] > 168h, which is free
on an endpoint already being polled and catches a renewal that silently stops.
Two things learned the hard way, both now in comments:
* A domain-expiry endpoint needs a URL SCHEME. Gatus derives the endpoint type
from the prefix (endpoint.Type()), so a bare "contrapeso.xyz" is UNKNOWN and
the whole config is rejected. It is https:// plus a DOMAIN_EXPIRATION
condition and no status assertion, so the registrar's parking page at the
apex is irrelevant. Upstream also enforces a 5m minimum interval for that
placeholder, because it uses a free whois service.
* That rejection proved skip-invalid-config-update was worth adding. Gatus
logged "the configuration file was updated, but it is not valid, the old
configuration will continue being used" and kept running. Without it the
reload path calls panic() and one malformed contributed file takes the
monitor down.
Verified: 15/15 service endpoints UP; 26/27 public-facing UP. The single DOWN is
public_memos, correctly - memos-box is powered off, and the condition result
reads [STATUS] (502) == 200. memos-box and arbret-staging-box were shut down
deliberately from the Proxmox UI to test the liveness endpoints; their endpoints
stay registered and will go green when the VMs return.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 09:33:57 +02:00
|
|
|
|
|
|
|
|
# ── DNS ──────────────────────────────────────────────────────────────────
|
alerting: Signal via signal-cli-rest-api, and faster failure detection
Three changes: detection windows tightened, a Signal transport deployed, and
alerts attached to all 84 non-transport endpoints.
── Failing faster ──────────────────────────────────────────────────────────
The heartbeat is a TICKER, not a deadline: Gatus wakes every interval and asks
"did anything arrive in the last interval", so real detection is 1-2x the
window. And a window can only ever be as tight as the push frequency - which is
why two checks moved rather than just having their numbers changed.
liveness, cpu, ups, service-health, probes 16m -> 11m
disk, zfs daily/30h -> 6-hourly/7h
backup store + pull job daily/30h -> 6-hourly/7h
DNS records 24h -> 6h
backup dump 30h -> 26h
backup-dump stays slow because the dump genuinely is daily. The store-side check
catches the same fault within 6h by reading the source's dump timestamp out of
the artefact filename, so 26h is a backstop rather than the primary signal.
── Thresholds differ by check type, deliberately ───────────────────────────
`failure-threshold` counts consecutive failures, but "consecutive" is a
different amount of wall-clock time per check: a push endpoint produces one
failure per heartbeat window, a pulled one per interval. The default of 3 would
mean 33 minutes on an 11m heartbeat and over a day on a 7h one.
More importantly the heartbeat window ALREADY encodes the tolerance - an 11m
window on a 5-minute push is exactly "one missed push forgiven" - so stacking a
threshold of 3 triples a tolerance that was already chosen. Hence:
push/heartbeat endpoints failure-threshold 1
pulled, 5m (public) failure-threshold 3 (= 15 minutes)
pulled, 6h/24h (dns, domain) failure-threshold 1
── The Signal transport ────────────────────────────────────────────────────
roles/signal_api runs signal-cli-rest-api on the observability host, pinned by
digest, MODE=native.
It publishes NO PORTS. The API has no authentication of any kind - anything that
reaches it can send messages as you and read your Signal. Gatus talks to it over
a shared docker network by service name, which is also WHY the network exists:
Gatus runs in a container, so the host's loopback is unreachable from it and a
port published on 127.0.0.1 would not have worked.
MODE=native and not json-rpc because this VPS has 464MB of RAM and already runs
Gatus and Caddy. The json-rpc modes hold a resident JVM; native runs a binary
per request, and alerts are rare enough that startup cost per alert is the right
trade.
Monitored - Gatus polls /v1/health over the same network path the alerts take,
so it proves the delivery route rather than mere container liveness. Deliberately
NOT backed up: the data directory holds Signal private keys and the recovery
path is to link the device again from the phone.
Neither the signal-api endpoint nor Gatus's self-check carries a Signal alert.
If either is down, Signal is precisely what cannot deliver the alert.
── Four traps hit while linking, all now in the role README ────────────────
* /v1/qrcodelink is BROKEN in native mode - returns "no data to encode" while
the binary itself emits a perfectly good URI. Not worked around by switching
MODE, which would put a JVM in the path of every alert permanently.
* The data dir must be owned by uid 1000, not root. A root-owned 0700 dir
cannot be traversed by the container user, so linking silently never
completes and /v1/accounts returns "Failed to read local accounts list".
* `docker exec` runs as ROOT while the service runs as uid 1000, so without
--config the account is written to /root/... on the container's ephemeral
layer. It reports success and is destroyed on the next recreate.
* The phone reporting "network error" was IPv6: chat.signal.org resolves to
dualstack AAAA records first, the container has no IPv6 address, and this
host's IPv6 path is broken - the same edge that 404'd the Go tarball. Fixed
with a mounted gai.conf that prefers IPv4.
Verified: provider loads (configuredProviders=[signal]), a test message was
delivered and confirmed received, 84 endpoints carry alerts, 86 UP / 0 DOWN.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 21:40:37 +02:00
|
|
|
# 6h, not daily: a DNS query is cheap and a wrong record is an outage. The
|
|
|
|
|
# domain check stays at 24h because it does a WHOIS/RDAP lookup against a
|
|
|
|
|
# free service. Alert on the first failure - at a 6h interval, waiting for
|
|
|
|
|
# three would be nearly a day.
|
monitoring: systemd services, domain expiry, DNS correctness, public endpoints
Four more check types, 42 endpoints, taking the estate from 39 to 81.
── systemd services (infra/401) ────────────────────────────────────────────
Every unit we deploy, checked every 5 minutes with a 16-minute heartbeat.
This closes the gap that let a real bug run unnoticed earlier today: a backup
script left forgejo, lnbits, headscale and memos stopped, and NOTHING caught
it. The dumps exited 0, the artefacts were correct, the deploy said failed=0,
and liveness only proves the HOST is up - not that anything on it serves.
One endpoint PER UNIT but only ONE timer per host: the check iterates that
host's units and pushes a result for each, the way check-backups.sh reports per
source. Four units on vipy would otherwise mean four scripts, services and
timers. A single host-level red light would also say "something on vipy is
down" without saying which, which is not the question anyone has.
Keys are host-qualified because unit names collide - caddy runs on four
machines. Which units a host runs lives in host_vars/<host>/monitored_services,
because "what runs here" is a property of the machine, the same reasoning as
the cross-host ports.
nut-driver-enumerator is deliberately excluded: it is a oneshot generator that
is `enabled` but always `inactive`, so it would report down forever. Checked
live before excluding it.
── domain, DNS and public endpoints (infra/402) ────────────────────────────
The first checks that PULL rather than push, and that is the right way round:
all three are about how the outside world sees us, so they must be measured
from outside. Nothing is installed anywhere - no script, no timer, no token.
They also have no heartbeat, because a heartbeat answers "did the reporter
report"; when Gatus does the checking itself, failure is immediate.
domain 1 endpoint, 24h, [DOMAIN_EXPIRATION] > 336h (two weeks)
dns 11 endpoints, 24h, [DNS_RCODE] == NOERROR and [BODY] == the IP
public 14 endpoints, 5m, 11 HTTPS + 3 TCP
Expected A records are derived from inventory (hostvars[host].ansible_host),
not written down again. The estate's recurring bug is an address recorded in a
second place and left behind when the machine moved; asserting against
inventory means a renumbered box is one edit, not two.
Expected HTTP status was checked live per site rather than assumed. Two return
401 - the Gatus dashboard and the DATUM dashboard, both behind basic auth - and
that is what is asserted: expecting 200 there would go green precisely when the
auth broke. headscale asserts /health rather than /, which is a 404 by design.
Every HTTPS check also carries [CERTIFICATE_EXPIRATION] > 168h, which is free
on an endpoint already being polled and catches a renewal that silently stops.
Two things learned the hard way, both now in comments:
* A domain-expiry endpoint needs a URL SCHEME. Gatus derives the endpoint type
from the prefix (endpoint.Type()), so a bare "contrapeso.xyz" is UNKNOWN and
the whole config is rejected. It is https:// plus a DOMAIN_EXPIRATION
condition and no status assertion, so the registrar's parking page at the
apex is irrelevant. Upstream also enforces a 5m minimum interval for that
placeholder, because it uses a free whois service.
* That rejection proved skip-invalid-config-update was worth adding. Gatus
logged "the configuration file was updated, but it is not valid, the old
configuration will continue being used" and kept running. Without it the
reload path calls panic() and one malformed contributed file takes the
monitor down.
Verified: 15/15 service endpoints UP; 26/27 public-facing UP. The single DOWN is
public_memos, correctly - memos-box is powered off, and the condition result
reads [STATUS] (502) == 200. memos-box and arbret-staging-box were shut down
deliberately from the Proxmox UI to test the liveness endpoints; their endpoints
stay registered and will go green when the VMs return.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 09:33:57 +02:00
|
|
|
- name: Build the DNS endpoints
|
|
|
|
|
ansible.builtin.set_fact:
|
|
|
|
|
dns_endpoints: "{{ dns_endpoints | default([]) + [{
|
|
|
|
|
'name': item.sub ~ '.' ~ root_domain,
|
|
|
|
|
'group': 'dns',
|
|
|
|
|
'url': dns_resolver,
|
alerting: Signal via signal-cli-rest-api, and faster failure detection
Three changes: detection windows tightened, a Signal transport deployed, and
alerts attached to all 84 non-transport endpoints.
── Failing faster ──────────────────────────────────────────────────────────
The heartbeat is a TICKER, not a deadline: Gatus wakes every interval and asks
"did anything arrive in the last interval", so real detection is 1-2x the
window. And a window can only ever be as tight as the push frequency - which is
why two checks moved rather than just having their numbers changed.
liveness, cpu, ups, service-health, probes 16m -> 11m
disk, zfs daily/30h -> 6-hourly/7h
backup store + pull job daily/30h -> 6-hourly/7h
DNS records 24h -> 6h
backup dump 30h -> 26h
backup-dump stays slow because the dump genuinely is daily. The store-side check
catches the same fault within 6h by reading the source's dump timestamp out of
the artefact filename, so 26h is a backstop rather than the primary signal.
── Thresholds differ by check type, deliberately ───────────────────────────
`failure-threshold` counts consecutive failures, but "consecutive" is a
different amount of wall-clock time per check: a push endpoint produces one
failure per heartbeat window, a pulled one per interval. The default of 3 would
mean 33 minutes on an 11m heartbeat and over a day on a 7h one.
More importantly the heartbeat window ALREADY encodes the tolerance - an 11m
window on a 5-minute push is exactly "one missed push forgiven" - so stacking a
threshold of 3 triples a tolerance that was already chosen. Hence:
push/heartbeat endpoints failure-threshold 1
pulled, 5m (public) failure-threshold 3 (= 15 minutes)
pulled, 6h/24h (dns, domain) failure-threshold 1
── The Signal transport ────────────────────────────────────────────────────
roles/signal_api runs signal-cli-rest-api on the observability host, pinned by
digest, MODE=native.
It publishes NO PORTS. The API has no authentication of any kind - anything that
reaches it can send messages as you and read your Signal. Gatus talks to it over
a shared docker network by service name, which is also WHY the network exists:
Gatus runs in a container, so the host's loopback is unreachable from it and a
port published on 127.0.0.1 would not have worked.
MODE=native and not json-rpc because this VPS has 464MB of RAM and already runs
Gatus and Caddy. The json-rpc modes hold a resident JVM; native runs a binary
per request, and alerts are rare enough that startup cost per alert is the right
trade.
Monitored - Gatus polls /v1/health over the same network path the alerts take,
so it proves the delivery route rather than mere container liveness. Deliberately
NOT backed up: the data directory holds Signal private keys and the recovery
path is to link the device again from the phone.
Neither the signal-api endpoint nor Gatus's self-check carries a Signal alert.
If either is down, Signal is precisely what cannot deliver the alert.
── Four traps hit while linking, all now in the role README ────────────────
* /v1/qrcodelink is BROKEN in native mode - returns "no data to encode" while
the binary itself emits a perfectly good URI. Not worked around by switching
MODE, which would put a JVM in the path of every alert permanently.
* The data dir must be owned by uid 1000, not root. A root-owned 0700 dir
cannot be traversed by the container user, so linking silently never
completes and /v1/accounts returns "Failed to read local accounts list".
* `docker exec` runs as ROOT while the service runs as uid 1000, so without
--config the account is written to /root/... on the container's ephemeral
layer. It reports success and is destroyed on the next recreate.
* The phone reporting "network error" was IPv6: chat.signal.org resolves to
dualstack AAAA records first, the container has no IPv6 address, and this
host's IPv6 path is broken - the same edge that 404'd the Go tarball. Fixed
with a mounted gai.conf that prefers IPv4.
Verified: provider loads (configuredProviders=[signal]), a test message was
delivered and confirmed received, 84 endpoints carry alerts, 86 UP / 0 DOWN.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 21:40:37 +02:00
|
|
|
'interval': '6h',
|
monitoring: systemd services, domain expiry, DNS correctness, public endpoints
Four more check types, 42 endpoints, taking the estate from 39 to 81.
── systemd services (infra/401) ────────────────────────────────────────────
Every unit we deploy, checked every 5 minutes with a 16-minute heartbeat.
This closes the gap that let a real bug run unnoticed earlier today: a backup
script left forgejo, lnbits, headscale and memos stopped, and NOTHING caught
it. The dumps exited 0, the artefacts were correct, the deploy said failed=0,
and liveness only proves the HOST is up - not that anything on it serves.
One endpoint PER UNIT but only ONE timer per host: the check iterates that
host's units and pushes a result for each, the way check-backups.sh reports per
source. Four units on vipy would otherwise mean four scripts, services and
timers. A single host-level red light would also say "something on vipy is
down" without saying which, which is not the question anyone has.
Keys are host-qualified because unit names collide - caddy runs on four
machines. Which units a host runs lives in host_vars/<host>/monitored_services,
because "what runs here" is a property of the machine, the same reasoning as
the cross-host ports.
nut-driver-enumerator is deliberately excluded: it is a oneshot generator that
is `enabled` but always `inactive`, so it would report down forever. Checked
live before excluding it.
── domain, DNS and public endpoints (infra/402) ────────────────────────────
The first checks that PULL rather than push, and that is the right way round:
all three are about how the outside world sees us, so they must be measured
from outside. Nothing is installed anywhere - no script, no timer, no token.
They also have no heartbeat, because a heartbeat answers "did the reporter
report"; when Gatus does the checking itself, failure is immediate.
domain 1 endpoint, 24h, [DOMAIN_EXPIRATION] > 336h (two weeks)
dns 11 endpoints, 24h, [DNS_RCODE] == NOERROR and [BODY] == the IP
public 14 endpoints, 5m, 11 HTTPS + 3 TCP
Expected A records are derived from inventory (hostvars[host].ansible_host),
not written down again. The estate's recurring bug is an address recorded in a
second place and left behind when the machine moved; asserting against
inventory means a renumbered box is one edit, not two.
Expected HTTP status was checked live per site rather than assumed. Two return
401 - the Gatus dashboard and the DATUM dashboard, both behind basic auth - and
that is what is asserted: expecting 200 there would go green precisely when the
auth broke. headscale asserts /health rather than /, which is a 404 by design.
Every HTTPS check also carries [CERTIFICATE_EXPIRATION] > 168h, which is free
on an endpoint already being polled and catches a renewal that silently stops.
Two things learned the hard way, both now in comments:
* A domain-expiry endpoint needs a URL SCHEME. Gatus derives the endpoint type
from the prefix (endpoint.Type()), so a bare "contrapeso.xyz" is UNKNOWN and
the whole config is rejected. It is https:// plus a DOMAIN_EXPIRATION
condition and no status assertion, so the registrar's parking page at the
apex is irrelevant. Upstream also enforces a 5m minimum interval for that
placeholder, because it uses a free whois service.
* That rejection proved skip-invalid-config-update was worth adding. Gatus
logged "the configuration file was updated, but it is not valid, the old
configuration will continue being used" and kept running. Without it the
reload path calls panic() and one malformed contributed file takes the
monitor down.
Verified: 15/15 service endpoints UP; 26/27 public-facing UP. The single DOWN is
public_memos, correctly - memos-box is powered off, and the condition result
reads [STATUS] (502) == 200. memos-box and arbret-staging-box were shut down
deliberately from the Proxmox UI to test the liveness endpoints; their endpoints
stay registered and will go green when the VMs return.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 09:33:57 +02:00
|
|
|
'dns': {'query-type': 'A', 'query-name': item.sub ~ '.' ~ root_domain},
|
|
|
|
|
'conditions': ['[DNS_RCODE] == NOERROR',
|
alerting: Signal via signal-cli-rest-api, and faster failure detection
Three changes: detection windows tightened, a Signal transport deployed, and
alerts attached to all 84 non-transport endpoints.
── Failing faster ──────────────────────────────────────────────────────────
The heartbeat is a TICKER, not a deadline: Gatus wakes every interval and asks
"did anything arrive in the last interval", so real detection is 1-2x the
window. And a window can only ever be as tight as the push frequency - which is
why two checks moved rather than just having their numbers changed.
liveness, cpu, ups, service-health, probes 16m -> 11m
disk, zfs daily/30h -> 6-hourly/7h
backup store + pull job daily/30h -> 6-hourly/7h
DNS records 24h -> 6h
backup dump 30h -> 26h
backup-dump stays slow because the dump genuinely is daily. The store-side check
catches the same fault within 6h by reading the source's dump timestamp out of
the artefact filename, so 26h is a backstop rather than the primary signal.
── Thresholds differ by check type, deliberately ───────────────────────────
`failure-threshold` counts consecutive failures, but "consecutive" is a
different amount of wall-clock time per check: a push endpoint produces one
failure per heartbeat window, a pulled one per interval. The default of 3 would
mean 33 minutes on an 11m heartbeat and over a day on a 7h one.
More importantly the heartbeat window ALREADY encodes the tolerance - an 11m
window on a 5-minute push is exactly "one missed push forgiven" - so stacking a
threshold of 3 triples a tolerance that was already chosen. Hence:
push/heartbeat endpoints failure-threshold 1
pulled, 5m (public) failure-threshold 3 (= 15 minutes)
pulled, 6h/24h (dns, domain) failure-threshold 1
── The Signal transport ────────────────────────────────────────────────────
roles/signal_api runs signal-cli-rest-api on the observability host, pinned by
digest, MODE=native.
It publishes NO PORTS. The API has no authentication of any kind - anything that
reaches it can send messages as you and read your Signal. Gatus talks to it over
a shared docker network by service name, which is also WHY the network exists:
Gatus runs in a container, so the host's loopback is unreachable from it and a
port published on 127.0.0.1 would not have worked.
MODE=native and not json-rpc because this VPS has 464MB of RAM and already runs
Gatus and Caddy. The json-rpc modes hold a resident JVM; native runs a binary
per request, and alerts are rare enough that startup cost per alert is the right
trade.
Monitored - Gatus polls /v1/health over the same network path the alerts take,
so it proves the delivery route rather than mere container liveness. Deliberately
NOT backed up: the data directory holds Signal private keys and the recovery
path is to link the device again from the phone.
Neither the signal-api endpoint nor Gatus's self-check carries a Signal alert.
If either is down, Signal is precisely what cannot deliver the alert.
── Four traps hit while linking, all now in the role README ────────────────
* /v1/qrcodelink is BROKEN in native mode - returns "no data to encode" while
the binary itself emits a perfectly good URI. Not worked around by switching
MODE, which would put a JVM in the path of every alert permanently.
* The data dir must be owned by uid 1000, not root. A root-owned 0700 dir
cannot be traversed by the container user, so linking silently never
completes and /v1/accounts returns "Failed to read local accounts list".
* `docker exec` runs as ROOT while the service runs as uid 1000, so without
--config the account is written to /root/... on the container's ephemeral
layer. It reports success and is destroyed on the next recreate.
* The phone reporting "network error" was IPv6: chat.signal.org resolves to
dualstack AAAA records first, the container has no IPv6 address, and this
host's IPv6 path is broken - the same edge that 404'd the Go tarball. Fixed
with a mounted gai.conf that prefers IPv4.
Verified: provider loads (configuredProviders=[signal]), a test message was
delivered and confirmed received, 84 endpoints carry alerts, 86 UP / 0 DOWN.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 21:40:37 +02:00
|
|
|
'[BODY] == ' ~ hostvars[item.host].ansible_host],
|
|
|
|
|
'alerts': [{'type': 'signal', 'failure-threshold': 1,
|
|
|
|
|
'success-threshold': 1, 'send-on-resolved': true,
|
|
|
|
|
'minimum-reminder-interval': '24h'}]}] }}"
|
monitoring: systemd services, domain expiry, DNS correctness, public endpoints
Four more check types, 42 endpoints, taking the estate from 39 to 81.
── systemd services (infra/401) ────────────────────────────────────────────
Every unit we deploy, checked every 5 minutes with a 16-minute heartbeat.
This closes the gap that let a real bug run unnoticed earlier today: a backup
script left forgejo, lnbits, headscale and memos stopped, and NOTHING caught
it. The dumps exited 0, the artefacts were correct, the deploy said failed=0,
and liveness only proves the HOST is up - not that anything on it serves.
One endpoint PER UNIT but only ONE timer per host: the check iterates that
host's units and pushes a result for each, the way check-backups.sh reports per
source. Four units on vipy would otherwise mean four scripts, services and
timers. A single host-level red light would also say "something on vipy is
down" without saying which, which is not the question anyone has.
Keys are host-qualified because unit names collide - caddy runs on four
machines. Which units a host runs lives in host_vars/<host>/monitored_services,
because "what runs here" is a property of the machine, the same reasoning as
the cross-host ports.
nut-driver-enumerator is deliberately excluded: it is a oneshot generator that
is `enabled` but always `inactive`, so it would report down forever. Checked
live before excluding it.
── domain, DNS and public endpoints (infra/402) ────────────────────────────
The first checks that PULL rather than push, and that is the right way round:
all three are about how the outside world sees us, so they must be measured
from outside. Nothing is installed anywhere - no script, no timer, no token.
They also have no heartbeat, because a heartbeat answers "did the reporter
report"; when Gatus does the checking itself, failure is immediate.
domain 1 endpoint, 24h, [DOMAIN_EXPIRATION] > 336h (two weeks)
dns 11 endpoints, 24h, [DNS_RCODE] == NOERROR and [BODY] == the IP
public 14 endpoints, 5m, 11 HTTPS + 3 TCP
Expected A records are derived from inventory (hostvars[host].ansible_host),
not written down again. The estate's recurring bug is an address recorded in a
second place and left behind when the machine moved; asserting against
inventory means a renumbered box is one edit, not two.
Expected HTTP status was checked live per site rather than assumed. Two return
401 - the Gatus dashboard and the DATUM dashboard, both behind basic auth - and
that is what is asserted: expecting 200 there would go green precisely when the
auth broke. headscale asserts /health rather than /, which is a 404 by design.
Every HTTPS check also carries [CERTIFICATE_EXPIRATION] > 168h, which is free
on an endpoint already being polled and catches a renewal that silently stops.
Two things learned the hard way, both now in comments:
* A domain-expiry endpoint needs a URL SCHEME. Gatus derives the endpoint type
from the prefix (endpoint.Type()), so a bare "contrapeso.xyz" is UNKNOWN and
the whole config is rejected. It is https:// plus a DOMAIN_EXPIRATION
condition and no status assertion, so the registrar's parking page at the
apex is irrelevant. Upstream also enforces a 5m minimum interval for that
placeholder, because it uses a free whois service.
* That rejection proved skip-invalid-config-update was worth adding. Gatus
logged "the configuration file was updated, but it is not valid, the old
configuration will continue being used" and kept running. Without it the
reload path calls panic() and one malformed contributed file takes the
monitor down.
Verified: 15/15 service endpoints UP; 26/27 public-facing UP. The single DOWN is
public_memos, correctly - memos-box is powered off, and the condition result
reads [STATUS] (502) == 200. memos-box and arbret-staging-box were shut down
deliberately from the Proxmox UI to test the liveness endpoints; their endpoints
stay registered and will go green when the VMs return.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 09:33:57 +02:00
|
|
|
loop: "{{ dns_records }}"
|
|
|
|
|
|
|
|
|
|
# ── Public HTTP ──────────────────────────────────────────────────────────
|
alerting: Signal via signal-cli-rest-api, and faster failure detection
Three changes: detection windows tightened, a Signal transport deployed, and
alerts attached to all 84 non-transport endpoints.
── Failing faster ──────────────────────────────────────────────────────────
The heartbeat is a TICKER, not a deadline: Gatus wakes every interval and asks
"did anything arrive in the last interval", so real detection is 1-2x the
window. And a window can only ever be as tight as the push frequency - which is
why two checks moved rather than just having their numbers changed.
liveness, cpu, ups, service-health, probes 16m -> 11m
disk, zfs daily/30h -> 6-hourly/7h
backup store + pull job daily/30h -> 6-hourly/7h
DNS records 24h -> 6h
backup dump 30h -> 26h
backup-dump stays slow because the dump genuinely is daily. The store-side check
catches the same fault within 6h by reading the source's dump timestamp out of
the artefact filename, so 26h is a backstop rather than the primary signal.
── Thresholds differ by check type, deliberately ───────────────────────────
`failure-threshold` counts consecutive failures, but "consecutive" is a
different amount of wall-clock time per check: a push endpoint produces one
failure per heartbeat window, a pulled one per interval. The default of 3 would
mean 33 minutes on an 11m heartbeat and over a day on a 7h one.
More importantly the heartbeat window ALREADY encodes the tolerance - an 11m
window on a 5-minute push is exactly "one missed push forgiven" - so stacking a
threshold of 3 triples a tolerance that was already chosen. Hence:
push/heartbeat endpoints failure-threshold 1
pulled, 5m (public) failure-threshold 3 (= 15 minutes)
pulled, 6h/24h (dns, domain) failure-threshold 1
── The Signal transport ────────────────────────────────────────────────────
roles/signal_api runs signal-cli-rest-api on the observability host, pinned by
digest, MODE=native.
It publishes NO PORTS. The API has no authentication of any kind - anything that
reaches it can send messages as you and read your Signal. Gatus talks to it over
a shared docker network by service name, which is also WHY the network exists:
Gatus runs in a container, so the host's loopback is unreachable from it and a
port published on 127.0.0.1 would not have worked.
MODE=native and not json-rpc because this VPS has 464MB of RAM and already runs
Gatus and Caddy. The json-rpc modes hold a resident JVM; native runs a binary
per request, and alerts are rare enough that startup cost per alert is the right
trade.
Monitored - Gatus polls /v1/health over the same network path the alerts take,
so it proves the delivery route rather than mere container liveness. Deliberately
NOT backed up: the data directory holds Signal private keys and the recovery
path is to link the device again from the phone.
Neither the signal-api endpoint nor Gatus's self-check carries a Signal alert.
If either is down, Signal is precisely what cannot deliver the alert.
── Four traps hit while linking, all now in the role README ────────────────
* /v1/qrcodelink is BROKEN in native mode - returns "no data to encode" while
the binary itself emits a perfectly good URI. Not worked around by switching
MODE, which would put a JVM in the path of every alert permanently.
* The data dir must be owned by uid 1000, not root. A root-owned 0700 dir
cannot be traversed by the container user, so linking silently never
completes and /v1/accounts returns "Failed to read local accounts list".
* `docker exec` runs as ROOT while the service runs as uid 1000, so without
--config the account is written to /root/... on the container's ephemeral
layer. It reports success and is destroyed on the next recreate.
* The phone reporting "network error" was IPv6: chat.signal.org resolves to
dualstack AAAA records first, the container has no IPv6 address, and this
host's IPv6 path is broken - the same edge that 404'd the Go tarball. Fixed
with a mounted gai.conf that prefers IPv4.
Verified: provider loads (configuredProviders=[signal]), a test message was
delivered and confirmed received, 84 endpoints carry alerts, 86 UP / 0 DOWN.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 21:40:37 +02:00
|
|
|
# failure-threshold 3 at a 5m interval = 15 minutes. A pulled endpoint has
|
|
|
|
|
# no heartbeat window, so unlike the push checks the tolerance has to live
|
|
|
|
|
# in the threshold - and one failed poll of a public site is usually a blip.
|
monitoring: systemd services, domain expiry, DNS correctness, public endpoints
Four more check types, 42 endpoints, taking the estate from 39 to 81.
── systemd services (infra/401) ────────────────────────────────────────────
Every unit we deploy, checked every 5 minutes with a 16-minute heartbeat.
This closes the gap that let a real bug run unnoticed earlier today: a backup
script left forgejo, lnbits, headscale and memos stopped, and NOTHING caught
it. The dumps exited 0, the artefacts were correct, the deploy said failed=0,
and liveness only proves the HOST is up - not that anything on it serves.
One endpoint PER UNIT but only ONE timer per host: the check iterates that
host's units and pushes a result for each, the way check-backups.sh reports per
source. Four units on vipy would otherwise mean four scripts, services and
timers. A single host-level red light would also say "something on vipy is
down" without saying which, which is not the question anyone has.
Keys are host-qualified because unit names collide - caddy runs on four
machines. Which units a host runs lives in host_vars/<host>/monitored_services,
because "what runs here" is a property of the machine, the same reasoning as
the cross-host ports.
nut-driver-enumerator is deliberately excluded: it is a oneshot generator that
is `enabled` but always `inactive`, so it would report down forever. Checked
live before excluding it.
── domain, DNS and public endpoints (infra/402) ────────────────────────────
The first checks that PULL rather than push, and that is the right way round:
all three are about how the outside world sees us, so they must be measured
from outside. Nothing is installed anywhere - no script, no timer, no token.
They also have no heartbeat, because a heartbeat answers "did the reporter
report"; when Gatus does the checking itself, failure is immediate.
domain 1 endpoint, 24h, [DOMAIN_EXPIRATION] > 336h (two weeks)
dns 11 endpoints, 24h, [DNS_RCODE] == NOERROR and [BODY] == the IP
public 14 endpoints, 5m, 11 HTTPS + 3 TCP
Expected A records are derived from inventory (hostvars[host].ansible_host),
not written down again. The estate's recurring bug is an address recorded in a
second place and left behind when the machine moved; asserting against
inventory means a renumbered box is one edit, not two.
Expected HTTP status was checked live per site rather than assumed. Two return
401 - the Gatus dashboard and the DATUM dashboard, both behind basic auth - and
that is what is asserted: expecting 200 there would go green precisely when the
auth broke. headscale asserts /health rather than /, which is a 404 by design.
Every HTTPS check also carries [CERTIFICATE_EXPIRATION] > 168h, which is free
on an endpoint already being polled and catches a renewal that silently stops.
Two things learned the hard way, both now in comments:
* A domain-expiry endpoint needs a URL SCHEME. Gatus derives the endpoint type
from the prefix (endpoint.Type()), so a bare "contrapeso.xyz" is UNKNOWN and
the whole config is rejected. It is https:// plus a DOMAIN_EXPIRATION
condition and no status assertion, so the registrar's parking page at the
apex is irrelevant. Upstream also enforces a 5m minimum interval for that
placeholder, because it uses a free whois service.
* That rejection proved skip-invalid-config-update was worth adding. Gatus
logged "the configuration file was updated, but it is not valid, the old
configuration will continue being used" and kept running. Without it the
reload path calls panic() and one malformed contributed file takes the
monitor down.
Verified: 15/15 service endpoints UP; 26/27 public-facing UP. The single DOWN is
public_memos, correctly - memos-box is powered off, and the condition result
reads [STATUS] (502) == 200. memos-box and arbret-staging-box were shut down
deliberately from the Proxmox UI to test the liveness endpoints; their endpoints
stay registered and will go green when the VMs return.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 09:33:57 +02:00
|
|
|
- name: Build the public HTTP endpoints
|
|
|
|
|
ansible.builtin.set_fact:
|
|
|
|
|
http_endpoints: "{{ http_endpoints | default([]) + [{
|
|
|
|
|
'name': item.name,
|
|
|
|
|
'group': 'public',
|
|
|
|
|
'url': 'https://' ~ item.sub ~ '.' ~ root_domain ~ item.path,
|
|
|
|
|
'interval': '5m',
|
|
|
|
|
'conditions': ['[STATUS] == ' ~ item.status,
|
alerting: Signal via signal-cli-rest-api, and faster failure detection
Three changes: detection windows tightened, a Signal transport deployed, and
alerts attached to all 84 non-transport endpoints.
── Failing faster ──────────────────────────────────────────────────────────
The heartbeat is a TICKER, not a deadline: Gatus wakes every interval and asks
"did anything arrive in the last interval", so real detection is 1-2x the
window. And a window can only ever be as tight as the push frequency - which is
why two checks moved rather than just having their numbers changed.
liveness, cpu, ups, service-health, probes 16m -> 11m
disk, zfs daily/30h -> 6-hourly/7h
backup store + pull job daily/30h -> 6-hourly/7h
DNS records 24h -> 6h
backup dump 30h -> 26h
backup-dump stays slow because the dump genuinely is daily. The store-side check
catches the same fault within 6h by reading the source's dump timestamp out of
the artefact filename, so 26h is a backstop rather than the primary signal.
── Thresholds differ by check type, deliberately ───────────────────────────
`failure-threshold` counts consecutive failures, but "consecutive" is a
different amount of wall-clock time per check: a push endpoint produces one
failure per heartbeat window, a pulled one per interval. The default of 3 would
mean 33 minutes on an 11m heartbeat and over a day on a 7h one.
More importantly the heartbeat window ALREADY encodes the tolerance - an 11m
window on a 5-minute push is exactly "one missed push forgiven" - so stacking a
threshold of 3 triples a tolerance that was already chosen. Hence:
push/heartbeat endpoints failure-threshold 1
pulled, 5m (public) failure-threshold 3 (= 15 minutes)
pulled, 6h/24h (dns, domain) failure-threshold 1
── The Signal transport ────────────────────────────────────────────────────
roles/signal_api runs signal-cli-rest-api on the observability host, pinned by
digest, MODE=native.
It publishes NO PORTS. The API has no authentication of any kind - anything that
reaches it can send messages as you and read your Signal. Gatus talks to it over
a shared docker network by service name, which is also WHY the network exists:
Gatus runs in a container, so the host's loopback is unreachable from it and a
port published on 127.0.0.1 would not have worked.
MODE=native and not json-rpc because this VPS has 464MB of RAM and already runs
Gatus and Caddy. The json-rpc modes hold a resident JVM; native runs a binary
per request, and alerts are rare enough that startup cost per alert is the right
trade.
Monitored - Gatus polls /v1/health over the same network path the alerts take,
so it proves the delivery route rather than mere container liveness. Deliberately
NOT backed up: the data directory holds Signal private keys and the recovery
path is to link the device again from the phone.
Neither the signal-api endpoint nor Gatus's self-check carries a Signal alert.
If either is down, Signal is precisely what cannot deliver the alert.
── Four traps hit while linking, all now in the role README ────────────────
* /v1/qrcodelink is BROKEN in native mode - returns "no data to encode" while
the binary itself emits a perfectly good URI. Not worked around by switching
MODE, which would put a JVM in the path of every alert permanently.
* The data dir must be owned by uid 1000, not root. A root-owned 0700 dir
cannot be traversed by the container user, so linking silently never
completes and /v1/accounts returns "Failed to read local accounts list".
* `docker exec` runs as ROOT while the service runs as uid 1000, so without
--config the account is written to /root/... on the container's ephemeral
layer. It reports success and is destroyed on the next recreate.
* The phone reporting "network error" was IPv6: chat.signal.org resolves to
dualstack AAAA records first, the container has no IPv6 address, and this
host's IPv6 path is broken - the same edge that 404'd the Go tarball. Fixed
with a mounted gai.conf that prefers IPv4.
Verified: provider loads (configuredProviders=[signal]), a test message was
delivered and confirmed received, 84 endpoints carry alerts, 86 UP / 0 DOWN.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 21:40:37 +02:00
|
|
|
'[CERTIFICATE_EXPIRATION] > 168h'],
|
|
|
|
|
'alerts': [{'type': 'signal', 'failure-threshold': 3,
|
|
|
|
|
'success-threshold': 2, 'send-on-resolved': true,
|
|
|
|
|
'minimum-reminder-interval': '6h'}]}] }}"
|
monitoring: systemd services, domain expiry, DNS correctness, public endpoints
Four more check types, 42 endpoints, taking the estate from 39 to 81.
── systemd services (infra/401) ────────────────────────────────────────────
Every unit we deploy, checked every 5 minutes with a 16-minute heartbeat.
This closes the gap that let a real bug run unnoticed earlier today: a backup
script left forgejo, lnbits, headscale and memos stopped, and NOTHING caught
it. The dumps exited 0, the artefacts were correct, the deploy said failed=0,
and liveness only proves the HOST is up - not that anything on it serves.
One endpoint PER UNIT but only ONE timer per host: the check iterates that
host's units and pushes a result for each, the way check-backups.sh reports per
source. Four units on vipy would otherwise mean four scripts, services and
timers. A single host-level red light would also say "something on vipy is
down" without saying which, which is not the question anyone has.
Keys are host-qualified because unit names collide - caddy runs on four
machines. Which units a host runs lives in host_vars/<host>/monitored_services,
because "what runs here" is a property of the machine, the same reasoning as
the cross-host ports.
nut-driver-enumerator is deliberately excluded: it is a oneshot generator that
is `enabled` but always `inactive`, so it would report down forever. Checked
live before excluding it.
── domain, DNS and public endpoints (infra/402) ────────────────────────────
The first checks that PULL rather than push, and that is the right way round:
all three are about how the outside world sees us, so they must be measured
from outside. Nothing is installed anywhere - no script, no timer, no token.
They also have no heartbeat, because a heartbeat answers "did the reporter
report"; when Gatus does the checking itself, failure is immediate.
domain 1 endpoint, 24h, [DOMAIN_EXPIRATION] > 336h (two weeks)
dns 11 endpoints, 24h, [DNS_RCODE] == NOERROR and [BODY] == the IP
public 14 endpoints, 5m, 11 HTTPS + 3 TCP
Expected A records are derived from inventory (hostvars[host].ansible_host),
not written down again. The estate's recurring bug is an address recorded in a
second place and left behind when the machine moved; asserting against
inventory means a renumbered box is one edit, not two.
Expected HTTP status was checked live per site rather than assumed. Two return
401 - the Gatus dashboard and the DATUM dashboard, both behind basic auth - and
that is what is asserted: expecting 200 there would go green precisely when the
auth broke. headscale asserts /health rather than /, which is a 404 by design.
Every HTTPS check also carries [CERTIFICATE_EXPIRATION] > 168h, which is free
on an endpoint already being polled and catches a renewal that silently stops.
Two things learned the hard way, both now in comments:
* A domain-expiry endpoint needs a URL SCHEME. Gatus derives the endpoint type
from the prefix (endpoint.Type()), so a bare "contrapeso.xyz" is UNKNOWN and
the whole config is rejected. It is https:// plus a DOMAIN_EXPIRATION
condition and no status assertion, so the registrar's parking page at the
apex is irrelevant. Upstream also enforces a 5m minimum interval for that
placeholder, because it uses a free whois service.
* That rejection proved skip-invalid-config-update was worth adding. Gatus
logged "the configuration file was updated, but it is not valid, the old
configuration will continue being used" and kept running. Without it the
reload path calls panic() and one malformed contributed file takes the
monitor down.
Verified: 15/15 service endpoints UP; 26/27 public-facing UP. The single DOWN is
public_memos, correctly - memos-box is powered off, and the condition result
reads [STATUS] (502) == 200. memos-box and arbret-staging-box were shut down
deliberately from the Proxmox UI to test the liveness endpoints; their endpoints
stay registered and will go green when the VMs return.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 09:33:57 +02:00
|
|
|
loop: "{{ public_sites }}"
|
|
|
|
|
|
|
|
|
|
# ── Public TCP ───────────────────────────────────────────────────────────
|
|
|
|
|
- name: Build the public TCP endpoints
|
|
|
|
|
ansible.builtin.set_fact:
|
|
|
|
|
tcp_endpoints: "{{ tcp_endpoints | default([]) + [{
|
|
|
|
|
'name': item.name,
|
|
|
|
|
'group': 'public',
|
|
|
|
|
'url': 'tcp://' ~ hostvars[item.host].ansible_host ~ ':' ~ item.port,
|
|
|
|
|
'interval': '5m',
|
alerting: Signal via signal-cli-rest-api, and faster failure detection
Three changes: detection windows tightened, a Signal transport deployed, and
alerts attached to all 84 non-transport endpoints.
── Failing faster ──────────────────────────────────────────────────────────
The heartbeat is a TICKER, not a deadline: Gatus wakes every interval and asks
"did anything arrive in the last interval", so real detection is 1-2x the
window. And a window can only ever be as tight as the push frequency - which is
why two checks moved rather than just having their numbers changed.
liveness, cpu, ups, service-health, probes 16m -> 11m
disk, zfs daily/30h -> 6-hourly/7h
backup store + pull job daily/30h -> 6-hourly/7h
DNS records 24h -> 6h
backup dump 30h -> 26h
backup-dump stays slow because the dump genuinely is daily. The store-side check
catches the same fault within 6h by reading the source's dump timestamp out of
the artefact filename, so 26h is a backstop rather than the primary signal.
── Thresholds differ by check type, deliberately ───────────────────────────
`failure-threshold` counts consecutive failures, but "consecutive" is a
different amount of wall-clock time per check: a push endpoint produces one
failure per heartbeat window, a pulled one per interval. The default of 3 would
mean 33 minutes on an 11m heartbeat and over a day on a 7h one.
More importantly the heartbeat window ALREADY encodes the tolerance - an 11m
window on a 5-minute push is exactly "one missed push forgiven" - so stacking a
threshold of 3 triples a tolerance that was already chosen. Hence:
push/heartbeat endpoints failure-threshold 1
pulled, 5m (public) failure-threshold 3 (= 15 minutes)
pulled, 6h/24h (dns, domain) failure-threshold 1
── The Signal transport ────────────────────────────────────────────────────
roles/signal_api runs signal-cli-rest-api on the observability host, pinned by
digest, MODE=native.
It publishes NO PORTS. The API has no authentication of any kind - anything that
reaches it can send messages as you and read your Signal. Gatus talks to it over
a shared docker network by service name, which is also WHY the network exists:
Gatus runs in a container, so the host's loopback is unreachable from it and a
port published on 127.0.0.1 would not have worked.
MODE=native and not json-rpc because this VPS has 464MB of RAM and already runs
Gatus and Caddy. The json-rpc modes hold a resident JVM; native runs a binary
per request, and alerts are rare enough that startup cost per alert is the right
trade.
Monitored - Gatus polls /v1/health over the same network path the alerts take,
so it proves the delivery route rather than mere container liveness. Deliberately
NOT backed up: the data directory holds Signal private keys and the recovery
path is to link the device again from the phone.
Neither the signal-api endpoint nor Gatus's self-check carries a Signal alert.
If either is down, Signal is precisely what cannot deliver the alert.
── Four traps hit while linking, all now in the role README ────────────────
* /v1/qrcodelink is BROKEN in native mode - returns "no data to encode" while
the binary itself emits a perfectly good URI. Not worked around by switching
MODE, which would put a JVM in the path of every alert permanently.
* The data dir must be owned by uid 1000, not root. A root-owned 0700 dir
cannot be traversed by the container user, so linking silently never
completes and /v1/accounts returns "Failed to read local accounts list".
* `docker exec` runs as ROOT while the service runs as uid 1000, so without
--config the account is written to /root/... on the container's ephemeral
layer. It reports success and is destroyed on the next recreate.
* The phone reporting "network error" was IPv6: chat.signal.org resolves to
dualstack AAAA records first, the container has no IPv6 address, and this
host's IPv6 path is broken - the same edge that 404'd the Go tarball. Fixed
with a mounted gai.conf that prefers IPv4.
Verified: provider loads (configuredProviders=[signal]), a test message was
delivered and confirmed received, 84 endpoints carry alerts, 86 UP / 0 DOWN.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 21:40:37 +02:00
|
|
|
'conditions': ['[CONNECTED] == true'],
|
|
|
|
|
'alerts': [{'type': 'signal', 'failure-threshold': 3,
|
|
|
|
|
'success-threshold': 2, 'send-on-resolved': true,
|
|
|
|
|
'minimum-reminder-interval': '6h'}]}] }}"
|
monitoring: systemd services, domain expiry, DNS correctness, public endpoints
Four more check types, 42 endpoints, taking the estate from 39 to 81.
── systemd services (infra/401) ────────────────────────────────────────────
Every unit we deploy, checked every 5 minutes with a 16-minute heartbeat.
This closes the gap that let a real bug run unnoticed earlier today: a backup
script left forgejo, lnbits, headscale and memos stopped, and NOTHING caught
it. The dumps exited 0, the artefacts were correct, the deploy said failed=0,
and liveness only proves the HOST is up - not that anything on it serves.
One endpoint PER UNIT but only ONE timer per host: the check iterates that
host's units and pushes a result for each, the way check-backups.sh reports per
source. Four units on vipy would otherwise mean four scripts, services and
timers. A single host-level red light would also say "something on vipy is
down" without saying which, which is not the question anyone has.
Keys are host-qualified because unit names collide - caddy runs on four
machines. Which units a host runs lives in host_vars/<host>/monitored_services,
because "what runs here" is a property of the machine, the same reasoning as
the cross-host ports.
nut-driver-enumerator is deliberately excluded: it is a oneshot generator that
is `enabled` but always `inactive`, so it would report down forever. Checked
live before excluding it.
── domain, DNS and public endpoints (infra/402) ────────────────────────────
The first checks that PULL rather than push, and that is the right way round:
all three are about how the outside world sees us, so they must be measured
from outside. Nothing is installed anywhere - no script, no timer, no token.
They also have no heartbeat, because a heartbeat answers "did the reporter
report"; when Gatus does the checking itself, failure is immediate.
domain 1 endpoint, 24h, [DOMAIN_EXPIRATION] > 336h (two weeks)
dns 11 endpoints, 24h, [DNS_RCODE] == NOERROR and [BODY] == the IP
public 14 endpoints, 5m, 11 HTTPS + 3 TCP
Expected A records are derived from inventory (hostvars[host].ansible_host),
not written down again. The estate's recurring bug is an address recorded in a
second place and left behind when the machine moved; asserting against
inventory means a renumbered box is one edit, not two.
Expected HTTP status was checked live per site rather than assumed. Two return
401 - the Gatus dashboard and the DATUM dashboard, both behind basic auth - and
that is what is asserted: expecting 200 there would go green precisely when the
auth broke. headscale asserts /health rather than /, which is a 404 by design.
Every HTTPS check also carries [CERTIFICATE_EXPIRATION] > 168h, which is free
on an endpoint already being polled and catches a renewal that silently stops.
Two things learned the hard way, both now in comments:
* A domain-expiry endpoint needs a URL SCHEME. Gatus derives the endpoint type
from the prefix (endpoint.Type()), so a bare "contrapeso.xyz" is UNKNOWN and
the whole config is rejected. It is https:// plus a DOMAIN_EXPIRATION
condition and no status assertion, so the registrar's parking page at the
apex is irrelevant. Upstream also enforces a 5m minimum interval for that
placeholder, because it uses a free whois service.
* That rejection proved skip-invalid-config-update was worth adding. Gatus
logged "the configuration file was updated, but it is not valid, the old
configuration will continue being used" and kept running. Without it the
reload path calls panic() and one malformed contributed file takes the
monitor down.
Verified: 15/15 service endpoints UP; 26/27 public-facing UP. The single DOWN is
public_memos, correctly - memos-box is powered off, and the condition result
reads [STATUS] (502) == 200. memos-box and arbret-staging-box were shut down
deliberately from the Proxmox UI to test the liveness endpoints; their endpoints
stay registered and will go green when the VMs return.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 09:33:57 +02:00
|
|
|
loop: "{{ public_tcp }}"
|
|
|
|
|
|
|
|
|
|
- name: Register the public-facing endpoints
|
|
|
|
|
ansible.builtin.include_role:
|
|
|
|
|
name: gatus_endpoint
|
|
|
|
|
vars:
|
|
|
|
|
gatus_endpoint_name: public
|
|
|
|
|
gatus_endpoint_pulled: "{{ domain_endpoints + dns_endpoints + http_endpoints + tcp_endpoints }}"
|