alerting: Signal via signal-cli-rest-api, and faster failure detection
Three changes: detection windows tightened, a Signal transport deployed, and
alerts attached to all 84 non-transport endpoints.
── Failing faster ──────────────────────────────────────────────────────────
The heartbeat is a TICKER, not a deadline: Gatus wakes every interval and asks
"did anything arrive in the last interval", so real detection is 1-2x the
window. And a window can only ever be as tight as the push frequency - which is
why two checks moved rather than just having their numbers changed.
liveness, cpu, ups, service-health, probes 16m -> 11m
disk, zfs daily/30h -> 6-hourly/7h
backup store + pull job daily/30h -> 6-hourly/7h
DNS records 24h -> 6h
backup dump 30h -> 26h
backup-dump stays slow because the dump genuinely is daily. The store-side check
catches the same fault within 6h by reading the source's dump timestamp out of
the artefact filename, so 26h is a backstop rather than the primary signal.
── Thresholds differ by check type, deliberately ───────────────────────────
`failure-threshold` counts consecutive failures, but "consecutive" is a
different amount of wall-clock time per check: a push endpoint produces one
failure per heartbeat window, a pulled one per interval. The default of 3 would
mean 33 minutes on an 11m heartbeat and over a day on a 7h one.
More importantly the heartbeat window ALREADY encodes the tolerance - an 11m
window on a 5-minute push is exactly "one missed push forgiven" - so stacking a
threshold of 3 triples a tolerance that was already chosen. Hence:
push/heartbeat endpoints failure-threshold 1
pulled, 5m (public) failure-threshold 3 (= 15 minutes)
pulled, 6h/24h (dns, domain) failure-threshold 1
── The Signal transport ────────────────────────────────────────────────────
roles/signal_api runs signal-cli-rest-api on the observability host, pinned by
digest, MODE=native.
It publishes NO PORTS. The API has no authentication of any kind - anything that
reaches it can send messages as you and read your Signal. Gatus talks to it over
a shared docker network by service name, which is also WHY the network exists:
Gatus runs in a container, so the host's loopback is unreachable from it and a
port published on 127.0.0.1 would not have worked.
MODE=native and not json-rpc because this VPS has 464MB of RAM and already runs
Gatus and Caddy. The json-rpc modes hold a resident JVM; native runs a binary
per request, and alerts are rare enough that startup cost per alert is the right
trade.
Monitored - Gatus polls /v1/health over the same network path the alerts take,
so it proves the delivery route rather than mere container liveness. Deliberately
NOT backed up: the data directory holds Signal private keys and the recovery
path is to link the device again from the phone.
Neither the signal-api endpoint nor Gatus's self-check carries a Signal alert.
If either is down, Signal is precisely what cannot deliver the alert.
── Four traps hit while linking, all now in the role README ────────────────
* /v1/qrcodelink is BROKEN in native mode - returns "no data to encode" while
the binary itself emits a perfectly good URI. Not worked around by switching
MODE, which would put a JVM in the path of every alert permanently.
* The data dir must be owned by uid 1000, not root. A root-owned 0700 dir
cannot be traversed by the container user, so linking silently never
completes and /v1/accounts returns "Failed to read local accounts list".
* `docker exec` runs as ROOT while the service runs as uid 1000, so without
--config the account is written to /root/... on the container's ephemeral
layer. It reports success and is destroyed on the next recreate.
* The phone reporting "network error" was IPv6: chat.signal.org resolves to
dualstack AAAA records first, the container has no IPv6 address, and this
host's IPv6 path is broken - the same edge that 404'd the Go tarball. Fixed
with a mounted gai.conf that prefers IPv4.
Verified: provider loads (configuredProviders=[signal]), a test message was
delivered and confirmed received, 84 endpoints carry alerts, 86 UP / 0 DOWN.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
parent
85040d5f67
commit
3a9e1d5851
20 changed files with 752 additions and 181 deletions
|
|
@ -34,6 +34,23 @@
|
|||
# slow apt run, a reboot - does not raise an alarm, but a check that has
|
||||
# genuinely stopped does.
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# Alerting thresholds, and why they differ by check type.
|
||||
#
|
||||
# `failure-threshold` counts CONSECUTIVE failures, but "consecutive" means a
|
||||
# different amount of wall-clock time per check:
|
||||
#
|
||||
# push/heartbeat endpoints a failure is produced once per heartbeat window
|
||||
# pulled endpoints a failure is produced once per interval
|
||||
#
|
||||
# So the default of 3 would mean 33 minutes on an 11m heartbeat and over a day
|
||||
# on a 7h one - and the heartbeat window ALREADY encodes the tolerance. An 11m
|
||||
# window on a 5-minute push is precisely "one missed push forgiven"; stacking a
|
||||
# threshold of 3 on top triples a tolerance that was already chosen.
|
||||
#
|
||||
# Hence: push endpoints alert on the FIRST heartbeat failure. Pulled endpoints
|
||||
# have no built-in tolerance, so the threshold is where it belongs for them.
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
- name: Register the host checks with Gatus
|
||||
hosts: observability
|
||||
become: yes
|
||||
|
|
@ -47,7 +64,7 @@
|
|||
'name': item,
|
||||
'group': 'liveness',
|
||||
'token': gatus_push_tokens[item],
|
||||
'heartbeat': '16m'}] }}"
|
||||
'heartbeat': '11m'}] }}"
|
||||
loop: "{{ monitored }}"
|
||||
|
||||
- name: Build the disk endpoint list
|
||||
|
|
@ -56,13 +73,20 @@
|
|||
'name': item,
|
||||
'group': 'disk',
|
||||
'token': gatus_push_tokens[item],
|
||||
'heartbeat': '30h'}] }}"
|
||||
'heartbeat': '7h'}] }}"
|
||||
loop: "{{ monitored }}"
|
||||
|
||||
- name: Register liveness endpoints
|
||||
ansible.builtin.include_role:
|
||||
name: gatus_endpoint
|
||||
vars:
|
||||
gatus_endpoint_default_alerts:
|
||||
- type: signal
|
||||
# 1, not 3: the heartbeat window is the tolerance. See the note above.
|
||||
failure-threshold: 1
|
||||
success-threshold: 2
|
||||
send-on-resolved: true
|
||||
minimum-reminder-interval: 6h
|
||||
gatus_endpoint_name: liveness
|
||||
gatus_endpoint_external: "{{ liveness_endpoints }}"
|
||||
|
||||
|
|
@ -70,6 +94,13 @@
|
|||
ansible.builtin.include_role:
|
||||
name: gatus_endpoint
|
||||
vars:
|
||||
gatus_endpoint_default_alerts:
|
||||
- type: signal
|
||||
# 1, not 3: the heartbeat window is the tolerance. See the note above.
|
||||
failure-threshold: 1
|
||||
success-threshold: 2
|
||||
send-on-resolved: true
|
||||
minimum-reminder-interval: 6h
|
||||
gatus_endpoint_name: disk
|
||||
gatus_endpoint_external: "{{ disk_endpoints }}"
|
||||
|
||||
|
|
@ -77,20 +108,27 @@
|
|||
ansible.builtin.include_role:
|
||||
name: gatus_endpoint
|
||||
vars:
|
||||
gatus_endpoint_default_alerts:
|
||||
- type: signal
|
||||
# 1, not 3: the heartbeat window is the tolerance. See the note above.
|
||||
failure-threshold: 1
|
||||
success-threshold: 2
|
||||
send-on-resolved: true
|
||||
minimum-reminder-interval: 6h
|
||||
gatus_endpoint_name: hypervisor
|
||||
gatus_endpoint_external:
|
||||
- name: cpu
|
||||
group: hypervisor
|
||||
token: "{{ gatus_push_tokens['nodito'] }}"
|
||||
heartbeat: "16m"
|
||||
heartbeat: "11m"
|
||||
- name: zfs
|
||||
group: hypervisor
|
||||
token: "{{ gatus_push_tokens['nodito'] }}"
|
||||
heartbeat: "30h"
|
||||
heartbeat: "7h"
|
||||
- name: ups
|
||||
group: hypervisor
|
||||
token: "{{ gatus_push_tokens['nodito'] }}"
|
||||
heartbeat: "16m"
|
||||
heartbeat: "11m"
|
||||
|
||||
- name: Deploy host liveness and disk checks
|
||||
hosts: managed
|
||||
|
|
@ -120,10 +158,13 @@
|
|||
healthcheck_name: disk-usage
|
||||
healthcheck_description: "Disk usage for {{ inventory_hostname }}"
|
||||
healthcheck_check: disk-usage
|
||||
# Daily. RandomizedDelaySec spreads twelve hosts across the hour rather
|
||||
# than having them all report in the same second.
|
||||
healthcheck_on_calendar: "*-*-* 07:00:00"
|
||||
healthcheck_randomized_delay: "3600"
|
||||
# Every 6h rather than daily. Disk usage itself moves slowly, but the
|
||||
# heartbeat can only be as tight as the push frequency - a daily push
|
||||
# forces a >24h window, and a stuck check then hides for a day and a
|
||||
# half. Six-hourly buys a 7h window. RandomizedDelaySec spreads the
|
||||
# hosts so twelve boxes do not all report in the same second.
|
||||
healthcheck_on_calendar: "*-*-* 00/6:00:00"
|
||||
healthcheck_randomized_delay: "900"
|
||||
healthcheck_boot_delay: "5min"
|
||||
healthcheck_push_url: "{{ gatus_api }}/disk_{{ host_key }}/external"
|
||||
healthcheck_push_token: "{{ host_token }}"
|
||||
|
|
@ -157,7 +198,7 @@
|
|||
healthcheck_check: zfs-health
|
||||
healthcheck_packages: [curl, jq]
|
||||
healthcheck_zfs_pool: "{{ zfs_pool_name }}"
|
||||
healthcheck_on_calendar: "*-*-* 07:20:00"
|
||||
healthcheck_on_calendar: "*-*-* 00/6:20:00"
|
||||
healthcheck_boot_delay: "10min"
|
||||
healthcheck_push_url: "{{ gatus_api }}/hypervisor_zfs/external"
|
||||
healthcheck_push_token: "{{ host_token }}"
|
||||
|
|
|
|||
|
|
@ -1,7 +1,7 @@
|
|||
---
|
||||
# Is each systemd-deployed service actually running?
|
||||
#
|
||||
# Every 5 minutes, with a 16-minute Gatus heartbeat - three missed runs before
|
||||
# Every 5 minutes, with an 11-minute Gatus heartbeat - one missed run before
|
||||
# it alarms, so a reboot or a slow check does not page anyone, but a host that
|
||||
# stops reporting does.
|
||||
#
|
||||
|
|
@ -28,6 +28,23 @@
|
|||
# Register one endpoint per unit. Runs first: Gatus reloads within 30s, and the
|
||||
# host play above takes minutes, so every endpoint exists before its first push.
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# Alerting thresholds, and why they differ by check type.
|
||||
#
|
||||
# `failure-threshold` counts CONSECUTIVE failures, but "consecutive" means a
|
||||
# different amount of wall-clock time per check:
|
||||
#
|
||||
# push/heartbeat endpoints a failure is produced once per heartbeat window
|
||||
# pulled endpoints a failure is produced once per interval
|
||||
#
|
||||
# So the default of 3 would mean 33 minutes on an 11m heartbeat and over a day
|
||||
# on a 7h one - and the heartbeat window ALREADY encodes the tolerance. An 11m
|
||||
# window on a 5-minute push is precisely "one missed push forgiven"; stacking a
|
||||
# threshold of 3 on top triples a tolerance that was already chosen.
|
||||
#
|
||||
# Hence: push endpoints alert on the FIRST heartbeat failure. Pulled endpoints
|
||||
# have no built-in tolerance, so the threshold is where it belongs for them.
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
- name: Register the service checks with Gatus
|
||||
hosts: observability
|
||||
become: yes
|
||||
|
|
@ -47,13 +64,20 @@
|
|||
'name': (item.0.host | lower | regex_replace('[/_.,# +&]', '-')) ~ '/' ~ item.1,
|
||||
'group': 'services',
|
||||
'token': gatus_push_tokens[item.0.host],
|
||||
'heartbeat': '16m'}] }}"
|
||||
'heartbeat': '11m'}] }}"
|
||||
loop: "{{ host_units | subelements('units') }}"
|
||||
|
||||
- name: Register the service endpoints
|
||||
ansible.builtin.include_role:
|
||||
name: gatus_endpoint
|
||||
vars:
|
||||
gatus_endpoint_default_alerts:
|
||||
- type: signal
|
||||
# 1, not 3: the heartbeat window is the tolerance. See the note above.
|
||||
failure-threshold: 1
|
||||
success-threshold: 2
|
||||
send-on-resolved: true
|
||||
minimum-reminder-interval: 6h
|
||||
gatus_endpoint_name: services
|
||||
gatus_endpoint_external: "{{ service_endpoints }}"
|
||||
|
||||
|
|
|
|||
|
|
@ -75,23 +75,36 @@
|
|||
'group': 'domain',
|
||||
'url': 'https://' ~ item,
|
||||
'interval': '24h',
|
||||
'conditions': ['[DOMAIN_EXPIRATION] > 336h']}] }}"
|
||||
'conditions': ['[DOMAIN_EXPIRATION] > 336h'],
|
||||
'alerts': [{'type': 'signal', 'failure-threshold': 1,
|
||||
'success-threshold': 1, 'send-on-resolved': true,
|
||||
'minimum-reminder-interval': '168h'}]}] }}"
|
||||
loop: "{{ monitored_domains }}"
|
||||
|
||||
# ── DNS ──────────────────────────────────────────────────────────────────
|
||||
# 6h, not daily: a DNS query is cheap and a wrong record is an outage. The
|
||||
# domain check stays at 24h because it does a WHOIS/RDAP lookup against a
|
||||
# free service. Alert on the first failure - at a 6h interval, waiting for
|
||||
# three would be nearly a day.
|
||||
- name: Build the DNS endpoints
|
||||
ansible.builtin.set_fact:
|
||||
dns_endpoints: "{{ dns_endpoints | default([]) + [{
|
||||
'name': item.sub ~ '.' ~ root_domain,
|
||||
'group': 'dns',
|
||||
'url': dns_resolver,
|
||||
'interval': '24h',
|
||||
'interval': '6h',
|
||||
'dns': {'query-type': 'A', 'query-name': item.sub ~ '.' ~ root_domain},
|
||||
'conditions': ['[DNS_RCODE] == NOERROR',
|
||||
'[BODY] == ' ~ hostvars[item.host].ansible_host]}] }}"
|
||||
'[BODY] == ' ~ hostvars[item.host].ansible_host],
|
||||
'alerts': [{'type': 'signal', 'failure-threshold': 1,
|
||||
'success-threshold': 1, 'send-on-resolved': true,
|
||||
'minimum-reminder-interval': '24h'}]}] }}"
|
||||
loop: "{{ dns_records }}"
|
||||
|
||||
# ── Public HTTP ──────────────────────────────────────────────────────────
|
||||
# failure-threshold 3 at a 5m interval = 15 minutes. A pulled endpoint has
|
||||
# no heartbeat window, so unlike the push checks the tolerance has to live
|
||||
# in the threshold - and one failed poll of a public site is usually a blip.
|
||||
- name: Build the public HTTP endpoints
|
||||
ansible.builtin.set_fact:
|
||||
http_endpoints: "{{ http_endpoints | default([]) + [{
|
||||
|
|
@ -100,7 +113,10 @@
|
|||
'url': 'https://' ~ item.sub ~ '.' ~ root_domain ~ item.path,
|
||||
'interval': '5m',
|
||||
'conditions': ['[STATUS] == ' ~ item.status,
|
||||
'[CERTIFICATE_EXPIRATION] > 168h']}] }}"
|
||||
'[CERTIFICATE_EXPIRATION] > 168h'],
|
||||
'alerts': [{'type': 'signal', 'failure-threshold': 3,
|
||||
'success-threshold': 2, 'send-on-resolved': true,
|
||||
'minimum-reminder-interval': '6h'}]}] }}"
|
||||
loop: "{{ public_sites }}"
|
||||
|
||||
# ── Public TCP ───────────────────────────────────────────────────────────
|
||||
|
|
@ -111,7 +127,10 @@
|
|||
'group': 'public',
|
||||
'url': 'tcp://' ~ hostvars[item.host].ansible_host ~ ':' ~ item.port,
|
||||
'interval': '5m',
|
||||
'conditions': ['[CONNECTED] == true']}] }}"
|
||||
'conditions': ['[CONNECTED] == true'],
|
||||
'alerts': [{'type': 'signal', 'failure-threshold': 3,
|
||||
'success-threshold': 2, 'send-on-resolved': true,
|
||||
'minimum-reminder-interval': '6h'}]}] }}"
|
||||
loop: "{{ public_tcp }}"
|
||||
|
||||
- name: Register the public-facing endpoints
|
||||
|
|
|
|||
|
|
@ -14,6 +14,23 @@
|
|||
# They used to push to Uptime Kuma. The scripts now POST with a bearer token
|
||||
# instead of GETting ?status=up, and each host uses its own token.
|
||||
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# Alerting thresholds, and why they differ by check type.
|
||||
#
|
||||
# `failure-threshold` counts CONSECUTIVE failures, but "consecutive" means a
|
||||
# different amount of wall-clock time per check:
|
||||
#
|
||||
# push/heartbeat endpoints a failure is produced once per heartbeat window
|
||||
# pulled endpoints a failure is produced once per interval
|
||||
#
|
||||
# So the default of 3 would mean 33 minutes on an 11m heartbeat and over a day
|
||||
# on a 7h one - and the heartbeat window ALREADY encodes the tolerance. An 11m
|
||||
# window on a 5-minute push is precisely "one missed push forgiven"; stacking a
|
||||
# threshold of 3 on top triples a tolerance that was already chosen.
|
||||
#
|
||||
# Hence: push endpoints alert on the FIRST heartbeat failure. Pulled endpoints
|
||||
# have no built-in tolerance, so the threshold is where it belongs for them.
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
- name: Register the per-service probes with Gatus
|
||||
hosts: observability
|
||||
become: yes
|
||||
|
|
@ -36,12 +53,19 @@
|
|||
'name': item.name,
|
||||
'group': 'probe',
|
||||
'token': gatus_push_tokens[item.host],
|
||||
'heartbeat': '16m'}] }}"
|
||||
'heartbeat': '11m'}] }}"
|
||||
loop: "{{ probes }}"
|
||||
|
||||
- name: Register the probe endpoints
|
||||
ansible.builtin.include_role:
|
||||
name: gatus_endpoint
|
||||
vars:
|
||||
gatus_endpoint_default_alerts:
|
||||
- type: signal
|
||||
# 1, not 3: the heartbeat window is the tolerance. See the note above.
|
||||
failure-threshold: 1
|
||||
success-threshold: 2
|
||||
send-on-resolved: true
|
||||
minimum-reminder-interval: 6h
|
||||
gatus_endpoint_name: probes
|
||||
gatus_endpoint_external: "{{ probe_endpoints }}"
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue