personal_infra/ansible/roles/signal_api/tasks/main.yml

94 lines
3.2 KiB
YAML
Raw Normal View History

alerting: Signal via signal-cli-rest-api, and faster failure detection Three changes: detection windows tightened, a Signal transport deployed, and alerts attached to all 84 non-transport endpoints. ── Failing faster ────────────────────────────────────────────────────────── The heartbeat is a TICKER, not a deadline: Gatus wakes every interval and asks "did anything arrive in the last interval", so real detection is 1-2x the window. And a window can only ever be as tight as the push frequency - which is why two checks moved rather than just having their numbers changed. liveness, cpu, ups, service-health, probes 16m -> 11m disk, zfs daily/30h -> 6-hourly/7h backup store + pull job daily/30h -> 6-hourly/7h DNS records 24h -> 6h backup dump 30h -> 26h backup-dump stays slow because the dump genuinely is daily. The store-side check catches the same fault within 6h by reading the source's dump timestamp out of the artefact filename, so 26h is a backstop rather than the primary signal. ── Thresholds differ by check type, deliberately ─────────────────────────── `failure-threshold` counts consecutive failures, but "consecutive" is a different amount of wall-clock time per check: a push endpoint produces one failure per heartbeat window, a pulled one per interval. The default of 3 would mean 33 minutes on an 11m heartbeat and over a day on a 7h one. More importantly the heartbeat window ALREADY encodes the tolerance - an 11m window on a 5-minute push is exactly "one missed push forgiven" - so stacking a threshold of 3 triples a tolerance that was already chosen. Hence: push/heartbeat endpoints failure-threshold 1 pulled, 5m (public) failure-threshold 3 (= 15 minutes) pulled, 6h/24h (dns, domain) failure-threshold 1 ── The Signal transport ──────────────────────────────────────────────────── roles/signal_api runs signal-cli-rest-api on the observability host, pinned by digest, MODE=native. It publishes NO PORTS. The API has no authentication of any kind - anything that reaches it can send messages as you and read your Signal. Gatus talks to it over a shared docker network by service name, which is also WHY the network exists: Gatus runs in a container, so the host's loopback is unreachable from it and a port published on 127.0.0.1 would not have worked. MODE=native and not json-rpc because this VPS has 464MB of RAM and already runs Gatus and Caddy. The json-rpc modes hold a resident JVM; native runs a binary per request, and alerts are rare enough that startup cost per alert is the right trade. Monitored - Gatus polls /v1/health over the same network path the alerts take, so it proves the delivery route rather than mere container liveness. Deliberately NOT backed up: the data directory holds Signal private keys and the recovery path is to link the device again from the phone. Neither the signal-api endpoint nor Gatus's self-check carries a Signal alert. If either is down, Signal is precisely what cannot deliver the alert. ── Four traps hit while linking, all now in the role README ──────────────── * /v1/qrcodelink is BROKEN in native mode - returns "no data to encode" while the binary itself emits a perfectly good URI. Not worked around by switching MODE, which would put a JVM in the path of every alert permanently. * The data dir must be owned by uid 1000, not root. A root-owned 0700 dir cannot be traversed by the container user, so linking silently never completes and /v1/accounts returns "Failed to read local accounts list". * `docker exec` runs as ROOT while the service runs as uid 1000, so without --config the account is written to /root/... on the container's ephemeral layer. It reports success and is destroyed on the next recreate. * The phone reporting "network error" was IPv6: chat.signal.org resolves to dualstack AAAA records first, the container has no IPv6 address, and this host's IPv6 path is broken - the same edge that 404'd the Go tarball. Fixed with a mounted gai.conf that prefers IPv4. Verified: provider loads (configuredProviders=[signal]), a test message was delivered and confirmed received, 84 endpoints carry alerts, 86 UP / 0 DOWN. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 21:40:37 +02:00
---
- name: Assert Docker is available
ansible.builtin.command: docker --version
register: signal_docker_check
changed_when: false
# Created explicitly rather than by either compose file, so neither stack has to
# be deployed before the other and neither owns it.
- name: Ensure the shared monitoring network exists
ansible.builtin.command: "docker network create {{ signal_api_network }}"
register: signal_net
changed_when: "'already exists' not in signal_net.stderr"
failed_when:
- signal_net.rc != 0
- "'already exists' not in signal_net.stderr"
- name: Create the signal-api directory
ansible.builtin.file:
path: "{{ signal_api_dir }}"
state: directory
owner: root
group: root
mode: "0755"
# Owned by the container's uid, NOT root.
#
# The image drops to uid 1000 (`setpriv --reuid=1000`), and a root-owned 0700
# directory cannot be traversed by uid 1000 - signal-cli then fails to write the
# account and linking silently never completes, leaving a 39-byte accounts.json
# with no accounts and the API returning "Failed to read local accounts list".
#
# 0700 on uid 1000 is still private: only that uid and root can read the Signal
# private keys, which is the property actually wanted.
- name: Create the signal-api data directory owned by the container user
ansible.builtin.file:
path: "{{ signal_api_data_dir }}"
state: directory
owner: "{{ signal_api_uid }}"
group: "{{ signal_api_uid }}"
mode: "0700"
- name: Write the IPv4-preference resolver config
ansible.builtin.template:
src: gai.conf.j2
dest: "{{ signal_api_dir }}/gai.conf"
owner: root
group: root
mode: "0644"
- name: Write the docker compose file
ansible.builtin.template:
src: docker-compose.yml.j2
dest: "{{ signal_api_dir }}/docker-compose.yml"
owner: root
group: root
mode: "0644"
- name: Pull the pinned signal-api image
ansible.builtin.command:
cmd: docker compose pull
chdir: "{{ signal_api_dir }}"
register: signal_pull
changed_when: "'Downloaded newer image' in signal_pull.stderr or 'Pull complete' in signal_pull.stderr"
- name: Start signal-api
ansible.builtin.command:
cmd: docker compose up -d --remove-orphans
chdir: "{{ signal_api_dir }}"
register: signal_up
changed_when: "'Started' in signal_up.stderr or 'Created' in signal_up.stderr or 'Recreated' in signal_up.stderr"
- name: Wait for the API to answer
ansible.builtin.command:
cmd: "docker exec {{ signal_api_service_name }} curl -fsS http://localhost:{{ signal_api_port }}/v1/health"
register: signal_health
until: signal_health.rc == 0
retries: 12
delay: 5
changed_when: false
- name: Report whether an account is linked yet
ansible.builtin.command:
cmd: "docker exec {{ signal_api_service_name }} curl -fsS http://localhost:{{ signal_api_port }}/v1/accounts"
register: signal_accounts
changed_when: false
failed_when: false
- name: Show the linking status
ansible.builtin.debug:
msg: >-
{{ 'Linked account(s): ' ~ signal_accounts.stdout
if (signal_accounts.stdout | default('[]') | trim) not in ['[]', '', 'null']
else 'NO ACCOUNT LINKED YET - this is a one-time manual step, see the role README.' }}