alerting: Signal via signal-cli-rest-api, and faster failure detection
Three changes: detection windows tightened, a Signal transport deployed, and
alerts attached to all 84 non-transport endpoints.
── Failing faster ──────────────────────────────────────────────────────────
The heartbeat is a TICKER, not a deadline: Gatus wakes every interval and asks
"did anything arrive in the last interval", so real detection is 1-2x the
window. And a window can only ever be as tight as the push frequency - which is
why two checks moved rather than just having their numbers changed.
liveness, cpu, ups, service-health, probes 16m -> 11m
disk, zfs daily/30h -> 6-hourly/7h
backup store + pull job daily/30h -> 6-hourly/7h
DNS records 24h -> 6h
backup dump 30h -> 26h
backup-dump stays slow because the dump genuinely is daily. The store-side check
catches the same fault within 6h by reading the source's dump timestamp out of
the artefact filename, so 26h is a backstop rather than the primary signal.
── Thresholds differ by check type, deliberately ───────────────────────────
`failure-threshold` counts consecutive failures, but "consecutive" is a
different amount of wall-clock time per check: a push endpoint produces one
failure per heartbeat window, a pulled one per interval. The default of 3 would
mean 33 minutes on an 11m heartbeat and over a day on a 7h one.
More importantly the heartbeat window ALREADY encodes the tolerance - an 11m
window on a 5-minute push is exactly "one missed push forgiven" - so stacking a
threshold of 3 triples a tolerance that was already chosen. Hence:
push/heartbeat endpoints failure-threshold 1
pulled, 5m (public) failure-threshold 3 (= 15 minutes)
pulled, 6h/24h (dns, domain) failure-threshold 1
── The Signal transport ────────────────────────────────────────────────────
roles/signal_api runs signal-cli-rest-api on the observability host, pinned by
digest, MODE=native.
It publishes NO PORTS. The API has no authentication of any kind - anything that
reaches it can send messages as you and read your Signal. Gatus talks to it over
a shared docker network by service name, which is also WHY the network exists:
Gatus runs in a container, so the host's loopback is unreachable from it and a
port published on 127.0.0.1 would not have worked.
MODE=native and not json-rpc because this VPS has 464MB of RAM and already runs
Gatus and Caddy. The json-rpc modes hold a resident JVM; native runs a binary
per request, and alerts are rare enough that startup cost per alert is the right
trade.
Monitored - Gatus polls /v1/health over the same network path the alerts take,
so it proves the delivery route rather than mere container liveness. Deliberately
NOT backed up: the data directory holds Signal private keys and the recovery
path is to link the device again from the phone.
Neither the signal-api endpoint nor Gatus's self-check carries a Signal alert.
If either is down, Signal is precisely what cannot deliver the alert.
── Four traps hit while linking, all now in the role README ────────────────
* /v1/qrcodelink is BROKEN in native mode - returns "no data to encode" while
the binary itself emits a perfectly good URI. Not worked around by switching
MODE, which would put a JVM in the path of every alert permanently.
* The data dir must be owned by uid 1000, not root. A root-owned 0700 dir
cannot be traversed by the container user, so linking silently never
completes and /v1/accounts returns "Failed to read local accounts list".
* `docker exec` runs as ROOT while the service runs as uid 1000, so without
--config the account is written to /root/... on the container's ephemeral
layer. It reports success and is destroyed on the next recreate.
* The phone reporting "network error" was IPv6: chat.signal.org resolves to
dualstack AAAA records first, the container has no IPv6 address, and this
host's IPv6 path is broken - the same edge that 404'd the Go tarball. Fixed
with a mounted gai.conf that prefers IPv4.
Verified: provider loads (configuredProviders=[signal]), a test message was
delivered and confirmed received, 84 endpoints carry alerts, 86 UP / 0 DOWN.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
parent
85040d5f67
commit
3a9e1d5851
20 changed files with 752 additions and 181 deletions
39
ansible/roles/signal_api/defaults/main.yml
Normal file
39
ansible/roles/signal_api/defaults/main.yml
Normal file
|
|
@ -0,0 +1,39 @@
|
|||
---
|
||||
# signal-cli-rest-api: the transport Gatus uses to send Signal messages.
|
||||
#
|
||||
# Gatus does not speak Signal. It POSTs JSON to this service, which holds the
|
||||
# actual Signal identity and does the protocol work.
|
||||
|
||||
# Pinned by digest for the same reason as Gatus: a tag is mutable.
|
||||
# Upstream publishes no versioned tags worth pinning to, so this pins the
|
||||
# DIGEST that `latest` resolved to when this was reviewed. `latest` is a moving
|
||||
# target; a digest is a content address, and `docker compose pull` either
|
||||
# fetches exactly this image or fails.
|
||||
signal_api_image_digest: "sha256:2399d449123cdad56c4d859277e3b9127e1a00c4d2ab4601c239882609286cf8"
|
||||
signal_api_image: "bbernhard/signal-cli-rest-api@{{ signal_api_image_digest }}"
|
||||
|
||||
signal_api_dir: /opt/signal-api
|
||||
signal_api_data_dir: "{{ signal_api_dir }}/data"
|
||||
|
||||
# MODE matters on this host. Upstream offers normal / native / json-rpc /
|
||||
# json-rpc-native. json-rpc keeps a resident JVM daemon and upstream describes it
|
||||
# as "increased memory" - this VPS has 464MB total and already runs Gatus and
|
||||
# Caddy, so a resident JVM is not affordable. `native` runs a precompiled
|
||||
# GraalVM binary per request: no daemon, no resident cost, and alerts are rare
|
||||
# enough that paying startup per alert is the right trade.
|
||||
signal_api_mode: native
|
||||
|
||||
# Port INSIDE the shared docker network. Never published to the host: this API
|
||||
# has NO AUTHENTICATION of any kind. Anyone who can reach it can send messages
|
||||
# as you and read your Signal.
|
||||
signal_api_port: 8080
|
||||
|
||||
# Both this and Gatus join this network so Gatus can reach the API by service
|
||||
# name. Gatus runs in a container, so the host's loopback is NOT reachable from
|
||||
# it - this is why a shared network is required rather than a published port.
|
||||
signal_api_network: monitoring
|
||||
signal_api_service_name: signal-api
|
||||
|
||||
# The uid the upstream image drops to (`setpriv --reuid=1000`). The data
|
||||
# directory must be owned by it or signal-cli cannot write the account.
|
||||
signal_api_uid: 1000
|
||||
Loading…
Add table
Add a link
Reference in a new issue