Three changes: detection windows tightened, a Signal transport deployed, and
alerts attached to all 84 non-transport endpoints.
── Failing faster ──────────────────────────────────────────────────────────
The heartbeat is a TICKER, not a deadline: Gatus wakes every interval and asks
"did anything arrive in the last interval", so real detection is 1-2x the
window. And a window can only ever be as tight as the push frequency - which is
why two checks moved rather than just having their numbers changed.
liveness, cpu, ups, service-health, probes 16m -> 11m
disk, zfs daily/30h -> 6-hourly/7h
backup store + pull job daily/30h -> 6-hourly/7h
DNS records 24h -> 6h
backup dump 30h -> 26h
backup-dump stays slow because the dump genuinely is daily. The store-side check
catches the same fault within 6h by reading the source's dump timestamp out of
the artefact filename, so 26h is a backstop rather than the primary signal.
── Thresholds differ by check type, deliberately ───────────────────────────
`failure-threshold` counts consecutive failures, but "consecutive" is a
different amount of wall-clock time per check: a push endpoint produces one
failure per heartbeat window, a pulled one per interval. The default of 3 would
mean 33 minutes on an 11m heartbeat and over a day on a 7h one.
More importantly the heartbeat window ALREADY encodes the tolerance - an 11m
window on a 5-minute push is exactly "one missed push forgiven" - so stacking a
threshold of 3 triples a tolerance that was already chosen. Hence:
push/heartbeat endpoints failure-threshold 1
pulled, 5m (public) failure-threshold 3 (= 15 minutes)
pulled, 6h/24h (dns, domain) failure-threshold 1
── The Signal transport ────────────────────────────────────────────────────
roles/signal_api runs signal-cli-rest-api on the observability host, pinned by
digest, MODE=native.
It publishes NO PORTS. The API has no authentication of any kind - anything that
reaches it can send messages as you and read your Signal. Gatus talks to it over
a shared docker network by service name, which is also WHY the network exists:
Gatus runs in a container, so the host's loopback is unreachable from it and a
port published on 127.0.0.1 would not have worked.
MODE=native and not json-rpc because this VPS has 464MB of RAM and already runs
Gatus and Caddy. The json-rpc modes hold a resident JVM; native runs a binary
per request, and alerts are rare enough that startup cost per alert is the right
trade.
Monitored - Gatus polls /v1/health over the same network path the alerts take,
so it proves the delivery route rather than mere container liveness. Deliberately
NOT backed up: the data directory holds Signal private keys and the recovery
path is to link the device again from the phone.
Neither the signal-api endpoint nor Gatus's self-check carries a Signal alert.
If either is down, Signal is precisely what cannot deliver the alert.
── Four traps hit while linking, all now in the role README ────────────────
* /v1/qrcodelink is BROKEN in native mode - returns "no data to encode" while
the binary itself emits a perfectly good URI. Not worked around by switching
MODE, which would put a JVM in the path of every alert permanently.
* The data dir must be owned by uid 1000, not root. A root-owned 0700 dir
cannot be traversed by the container user, so linking silently never
completes and /v1/accounts returns "Failed to read local accounts list".
* `docker exec` runs as ROOT while the service runs as uid 1000, so without
--config the account is written to /root/... on the container's ephemeral
layer. It reports success and is destroyed on the next recreate.
* The phone reporting "network error" was IPv6: chat.signal.org resolves to
dualstack AAAA records first, the container has no IPv6 address, and this
host's IPv6 path is broken - the same edge that 404'd the Go tarball. Fixed
with a mounted gai.conf that prefers IPv4.
Verified: provider loads (configuredProviders=[signal]), a test message was
delivered and confirmed received, 84 endpoints carry alerts, 86 UP / 0 DOWN.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
90 lines
5.4 KiB
YAML
90 lines
5.4 KiB
YAML
---
|
|
# Everything, in the order it has to happen.
|
|
#
|
|
# This file is a TABLE OF CONTENTS, not a second source of truth. It says what
|
|
# runs and in what order. It does NOT say which hosts get what — that stays on
|
|
# the `hosts:` line inside each playbook, exactly where it is today. Nothing
|
|
# moves; this file only makes the set readable in one place.
|
|
#
|
|
# What runs on a host? ansible-playbook site.yml --limit <host> --list-hosts
|
|
# Who gets thing Y? the `hosts:` line in Y's own playbook
|
|
# What is a host? ansible-inventory --graph
|
|
#
|
|
# Run a slice with --limit, or run one playbook directly as before. Nothing here
|
|
# changes how any individual playbook behaves.
|
|
|
|
# ── Baseline: every managed machine ─────────────────────────────────────────
|
|
- import_playbook: infra/01_user_and_access_setup_playbook.yml
|
|
- import_playbook: infra/02_firewall_and_fail2ban_playbook.yml
|
|
- import_playbook: infra/900_install_rsync.yml
|
|
- import_playbook: infra/920_join_headscale_mesh.yml
|
|
# Idempotent and kept permanently: guarantees a rebuilt or restored host cannot
|
|
# quietly bring the Uptime-Kuma-era monitoring back.
|
|
- import_playbook: infra/409_remove_legacy_monitoring.yml
|
|
|
|
# ── Monitoring ──────────────────────────────────────────────────────────────
|
|
# Gatus first: the three plays below register endpoints with it, and registering
|
|
# against a host that is not serving yet would simply fail.
|
|
- import_playbook: services/gatus/deploy_gatus_playbook.yml
|
|
# The Signal transport for Gatus alerts. Shares a docker network with Gatus and
|
|
# publishes no ports - the API has no authentication. Needs a one-time manual
|
|
# device link; see roles/signal_api/README.md.
|
|
- import_playbook: services/signal-api/deploy_signal_api_playbook.yml
|
|
- import_playbook: infra/400_host_monitoring.yml
|
|
- import_playbook: infra/401_service_monitoring.yml
|
|
- import_playbook: infra/402_public_monitoring.yml
|
|
# Registers where the per-service probes report. The probes themselves are
|
|
# deployed by each service's own playbook further down; the endpoints must exist
|
|
# before the first push arrives.
|
|
- import_playbook: infra/403_service_probe_registration.yml
|
|
# 910_docker says `hosts: managed`, but only 5 of 11 managed hosts have or need
|
|
# Docker. Left out until it has a [docker] group — see the note in PLAN_7.
|
|
|
|
# ── The hypervisor ──────────────────────────────────────────────────────────
|
|
- import_playbook: infra/nodito/31_proxmox_community_repos_playbook.yml
|
|
- import_playbook: infra/nodito/32_zfs_pool_setup_playbook.yml
|
|
- import_playbook: infra/nodito/34_nut_ups_setup_playbook.yml
|
|
|
|
# ── Reverse proxy, before anything that registers a vhost ───────────────────
|
|
- import_playbook: services/caddy_playbook.yml
|
|
|
|
# ── Services ────────────────────────────────────────────────────────────────
|
|
- import_playbook: services/bitcoin-knots/deploy_bitcoin_knots_playbook.yml
|
|
- import_playbook: services/fulcrum/deploy_fulcrum_playbook.yml
|
|
- import_playbook: services/datum-gateway/deploy_datum_gateway_playbook.yml
|
|
- import_playbook: services/mempool/deploy_mempool_playbook.yml
|
|
- import_playbook: services/memos/deploy_memos_playbook.yml
|
|
- import_playbook: services/forgejo-runner/deploy_forgejo_runner_playbook.yml
|
|
- import_playbook: services/phoenixd/deploy_phoenixd_playbook.yml
|
|
- import_playbook: services/headscale/deploy_headscale_playbook.yml
|
|
- import_playbook: services/vaultwarden/deploy_vaultwarden_playbook.yml
|
|
- import_playbook: services/forgejo/deploy_forgejo_playbook.yml
|
|
- import_playbook: services/lnbits/deploy_lnbits_playbook.yml
|
|
- import_playbook: services/ntfy-emergency-app/deploy_ntfy_emergency_app_playbook.yml
|
|
- import_playbook: services/personal-blog/deploy_personal_blog_playbook.yml
|
|
|
|
# ── Backups: each source dumps itself, the box pulls ────────────────────────
|
|
- import_playbook: services/headscale/setup_backup_headscale.yml
|
|
- import_playbook: services/vaultwarden/setup_backup_vaultwarden.yml
|
|
- import_playbook: services/forgejo/setup_backup_forgejo.yml
|
|
- import_playbook: services/lnbits/setup_backup_lnbits.yml
|
|
- import_playbook: services/memos/setup_backup_memos.yml
|
|
- import_playbook: playbooks/backups.yml
|
|
|
|
# Deliberately not here. Every playbook in the repo is either imported above or
|
|
# listed below, so this file accounts for all of them:
|
|
#
|
|
# infra/910_docker_playbook.yml says `hosts: managed`, but Docker is on 5
|
|
# of 11 managed hosts and those 5 are exactly
|
|
# the ones that need it. Running it would
|
|
# install Docker on the Bitcoin node and the
|
|
# hypervisor. Needs a [docker] group first.
|
|
#
|
|
# infra/nodito/30_proxmox_bootstrap one-shot: bare-metal bootstrap, run once
|
|
# infra/nodito/33_..._cloud_template one-shot: builds the VM template
|
|
#
|
|
#
|
|
# services/vaultwarden/disable_ deliberate manual actions, not convergence
|
|
# vaultwarden_sign_ups_playbook.yml
|
|
# services/personal-blog/setup_
|
|
# deploy_alias_lapy.yml
|