Three changes: detection windows tightened, a Signal transport deployed, and
alerts attached to all 84 non-transport endpoints.
── Failing faster ──────────────────────────────────────────────────────────
The heartbeat is a TICKER, not a deadline: Gatus wakes every interval and asks
"did anything arrive in the last interval", so real detection is 1-2x the
window. And a window can only ever be as tight as the push frequency - which is
why two checks moved rather than just having their numbers changed.
liveness, cpu, ups, service-health, probes 16m -> 11m
disk, zfs daily/30h -> 6-hourly/7h
backup store + pull job daily/30h -> 6-hourly/7h
DNS records 24h -> 6h
backup dump 30h -> 26h
backup-dump stays slow because the dump genuinely is daily. The store-side check
catches the same fault within 6h by reading the source's dump timestamp out of
the artefact filename, so 26h is a backstop rather than the primary signal.
── Thresholds differ by check type, deliberately ───────────────────────────
`failure-threshold` counts consecutive failures, but "consecutive" is a
different amount of wall-clock time per check: a push endpoint produces one
failure per heartbeat window, a pulled one per interval. The default of 3 would
mean 33 minutes on an 11m heartbeat and over a day on a 7h one.
More importantly the heartbeat window ALREADY encodes the tolerance - an 11m
window on a 5-minute push is exactly "one missed push forgiven" - so stacking a
threshold of 3 triples a tolerance that was already chosen. Hence:
push/heartbeat endpoints failure-threshold 1
pulled, 5m (public) failure-threshold 3 (= 15 minutes)
pulled, 6h/24h (dns, domain) failure-threshold 1
── The Signal transport ────────────────────────────────────────────────────
roles/signal_api runs signal-cli-rest-api on the observability host, pinned by
digest, MODE=native.
It publishes NO PORTS. The API has no authentication of any kind - anything that
reaches it can send messages as you and read your Signal. Gatus talks to it over
a shared docker network by service name, which is also WHY the network exists:
Gatus runs in a container, so the host's loopback is unreachable from it and a
port published on 127.0.0.1 would not have worked.
MODE=native and not json-rpc because this VPS has 464MB of RAM and already runs
Gatus and Caddy. The json-rpc modes hold a resident JVM; native runs a binary
per request, and alerts are rare enough that startup cost per alert is the right
trade.
Monitored - Gatus polls /v1/health over the same network path the alerts take,
so it proves the delivery route rather than mere container liveness. Deliberately
NOT backed up: the data directory holds Signal private keys and the recovery
path is to link the device again from the phone.
Neither the signal-api endpoint nor Gatus's self-check carries a Signal alert.
If either is down, Signal is precisely what cannot deliver the alert.
── Four traps hit while linking, all now in the role README ────────────────
* /v1/qrcodelink is BROKEN in native mode - returns "no data to encode" while
the binary itself emits a perfectly good URI. Not worked around by switching
MODE, which would put a JVM in the path of every alert permanently.
* The data dir must be owned by uid 1000, not root. A root-owned 0700 dir
cannot be traversed by the container user, so linking silently never
completes and /v1/accounts returns "Failed to read local accounts list".
* `docker exec` runs as ROOT while the service runs as uid 1000, so without
--config the account is written to /root/... on the container's ephemeral
layer. It reports success and is destroyed on the next recreate.
* The phone reporting "network error" was IPv6: chat.signal.org resolves to
dualstack AAAA records first, the container has no IPv6 address, and this
host's IPv6 path is broken - the same edge that 404'd the Go tarball. Fixed
with a mounted gai.conf that prefers IPv4.
Verified: provider loads (configuredProviders=[signal]), a test message was
delivered and confirmed received, 84 endpoints carry alerts, 86 UP / 0 DOWN.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
141 lines
8 KiB
YAML
141 lines
8 KiB
YAML
---
|
|
# Domain expiry, DNS correctness, and public endpoint reachability.
|
|
#
|
|
# These are the first checks in the estate that PULL rather than push, and that
|
|
# is the right way round for them: all three are about how the outside world
|
|
# sees us, so they must be measured from outside. Gatus polls from the
|
|
# observability host and needs nothing installed anywhere else - there is no
|
|
# script, no timer and no token, because nothing is reporting in.
|
|
#
|
|
# That also means these have no heartbeat. A heartbeat answers "did the thing
|
|
# that was supposed to report in do so"; when Gatus does the checking itself,
|
|
# failure is immediate and self-evident.
|
|
|
|
- name: Register the public-facing checks with Gatus
|
|
hosts: observability
|
|
become: yes
|
|
|
|
vars:
|
|
# Expected A records, derived from inventory rather than written down again.
|
|
# The estate's recurring bug is an address recorded in a second place and
|
|
# then left behind when the machine moved, so the check asserts against
|
|
# ansible_host - if a box is renumbered, inventory is the one edit.
|
|
dns_records:
|
|
- {sub: "{{ subdomains.gatus }}", host: monitoring}
|
|
- {sub: "{{ subdomains.headscale }}", host: spacey}
|
|
- {sub: "{{ subdomains.vaultwarden }}", host: vipy}
|
|
- {sub: "{{ subdomains.forgejo }}", host: vipy}
|
|
- {sub: "{{ subdomains.lnbits }}", host: vipy}
|
|
- {sub: "{{ subdomains.ntfy_emergency_app }}", host: vipy}
|
|
- {sub: "{{ subdomains.personal_blog }}", host: vipy}
|
|
- {sub: "{{ subdomains.memos }}", host: vipy}
|
|
- {sub: "{{ subdomains.mempool }}", host: vipy}
|
|
- {sub: "{{ subdomains.datum_gateway }}", host: vipy}
|
|
|
|
# A public resolver on purpose: this must test what the internet sees, not
|
|
# what a local cache or the tailnet's MagicDNS happens to answer.
|
|
dns_resolver: "1.1.1.1"
|
|
|
|
# Expected status per site, checked live before being written down.
|
|
# 401 is the CORRECT answer for the two behind basic auth - asserting 200
|
|
# there would go green precisely when the auth broke.
|
|
public_sites:
|
|
- {name: gatus, sub: "{{ subdomains.gatus }}", path: "/", status: 401}
|
|
- {name: headscale, sub: "{{ subdomains.headscale }}", path: "/health", status: 200}
|
|
- {name: vaultwarden, sub: "{{ subdomains.vaultwarden }}", path: "/", status: 200}
|
|
- {name: forgejo, sub: "{{ subdomains.forgejo }}", path: "/", status: 200}
|
|
- {name: lnbits, sub: "{{ subdomains.lnbits }}", path: "/", status: 200}
|
|
- {name: avisame, sub: "{{ subdomains.ntfy_emergency_app }}", path: "/", status: 200}
|
|
- {name: blog, sub: "{{ subdomains.personal_blog }}", path: "/", status: 200}
|
|
- {name: memos, sub: "{{ subdomains.memos }}", path: "/", status: 200}
|
|
- {name: mempool, sub: "{{ subdomains.mempool }}", path: "/", status: 200}
|
|
- {name: datum, sub: "{{ subdomains.datum_gateway }}", path: "/", status: 401}
|
|
|
|
# Ports published from the edge host by socket_proxy.
|
|
public_tcp:
|
|
- {name: bitcoin-p2p, host: vipy, port: "{{ hostvars['knots_box_local'].bitcoin_p2p_port }}"}
|
|
- {name: fulcrum-ssl, host: vipy, port: "{{ hostvars['fulcrum_box_local'].fulcrum_ssl_port }}"}
|
|
- {name: datum-stratum, host: vipy, port: "{{ hostvars['knots_box_local'].datum_gateway_stratum_port }}"}
|
|
|
|
tasks:
|
|
# ── Domain expiry ────────────────────────────────────────────────────────
|
|
# Each domain needs a URL SCHEME: Gatus derives the endpoint type from the
|
|
# prefix (endpoint.Type()), so a bare "example.com" is UNKNOWN and the whole
|
|
# config is rejected. No status is asserted, only the WHOIS/RDAP expiry, so
|
|
# whatever the apex serves - a real site, or the registrar's parking page -
|
|
# is irrelevant.
|
|
#
|
|
# 24h, and upstream enforces a 5m minimum for DOMAIN_EXPIRATION anyway
|
|
# because it uses a free whois service that must not be hammered.
|
|
# 336h = 14 days of runway, because renewal is a manual act at the registrar.
|
|
- name: Build the domain endpoints
|
|
ansible.builtin.set_fact:
|
|
domain_endpoints: "{{ domain_endpoints | default([]) + [{
|
|
'name': item,
|
|
'group': 'domain',
|
|
'url': 'https://' ~ item,
|
|
'interval': '24h',
|
|
'conditions': ['[DOMAIN_EXPIRATION] > 336h'],
|
|
'alerts': [{'type': 'signal', 'failure-threshold': 1,
|
|
'success-threshold': 1, 'send-on-resolved': true,
|
|
'minimum-reminder-interval': '168h'}]}] }}"
|
|
loop: "{{ monitored_domains }}"
|
|
|
|
# ── DNS ──────────────────────────────────────────────────────────────────
|
|
# 6h, not daily: a DNS query is cheap and a wrong record is an outage. The
|
|
# domain check stays at 24h because it does a WHOIS/RDAP lookup against a
|
|
# free service. Alert on the first failure - at a 6h interval, waiting for
|
|
# three would be nearly a day.
|
|
- name: Build the DNS endpoints
|
|
ansible.builtin.set_fact:
|
|
dns_endpoints: "{{ dns_endpoints | default([]) + [{
|
|
'name': item.sub ~ '.' ~ root_domain,
|
|
'group': 'dns',
|
|
'url': dns_resolver,
|
|
'interval': '6h',
|
|
'dns': {'query-type': 'A', 'query-name': item.sub ~ '.' ~ root_domain},
|
|
'conditions': ['[DNS_RCODE] == NOERROR',
|
|
'[BODY] == ' ~ hostvars[item.host].ansible_host],
|
|
'alerts': [{'type': 'signal', 'failure-threshold': 1,
|
|
'success-threshold': 1, 'send-on-resolved': true,
|
|
'minimum-reminder-interval': '24h'}]}] }}"
|
|
loop: "{{ dns_records }}"
|
|
|
|
# ── Public HTTP ──────────────────────────────────────────────────────────
|
|
# failure-threshold 3 at a 5m interval = 15 minutes. A pulled endpoint has
|
|
# no heartbeat window, so unlike the push checks the tolerance has to live
|
|
# in the threshold - and one failed poll of a public site is usually a blip.
|
|
- name: Build the public HTTP endpoints
|
|
ansible.builtin.set_fact:
|
|
http_endpoints: "{{ http_endpoints | default([]) + [{
|
|
'name': item.name,
|
|
'group': 'public',
|
|
'url': 'https://' ~ item.sub ~ '.' ~ root_domain ~ item.path,
|
|
'interval': '5m',
|
|
'conditions': ['[STATUS] == ' ~ item.status,
|
|
'[CERTIFICATE_EXPIRATION] > 168h'],
|
|
'alerts': [{'type': 'signal', 'failure-threshold': 3,
|
|
'success-threshold': 2, 'send-on-resolved': true,
|
|
'minimum-reminder-interval': '6h'}]}] }}"
|
|
loop: "{{ public_sites }}"
|
|
|
|
# ── Public TCP ───────────────────────────────────────────────────────────
|
|
- name: Build the public TCP endpoints
|
|
ansible.builtin.set_fact:
|
|
tcp_endpoints: "{{ tcp_endpoints | default([]) + [{
|
|
'name': item.name,
|
|
'group': 'public',
|
|
'url': 'tcp://' ~ hostvars[item.host].ansible_host ~ ':' ~ item.port,
|
|
'interval': '5m',
|
|
'conditions': ['[CONNECTED] == true'],
|
|
'alerts': [{'type': 'signal', 'failure-threshold': 3,
|
|
'success-threshold': 2, 'send-on-resolved': true,
|
|
'minimum-reminder-interval': '6h'}]}] }}"
|
|
loop: "{{ public_tcp }}"
|
|
|
|
- name: Register the public-facing endpoints
|
|
ansible.builtin.include_role:
|
|
name: gatus_endpoint
|
|
vars:
|
|
gatus_endpoint_name: public
|
|
gatus_endpoint_pulled: "{{ domain_endpoints + dns_endpoints + http_endpoints + tcp_endpoints }}"
|