personal_infra/ansible/site.yml
counterweight 85040d5f67
watchtower: remove from the estate, and with it ntfy
watchtower is being destroyed. Removed from [vps], with its host_vars, its push
token, and the six Gatus endpoints that referenced it (liveness, disk, two
systemd services, the ntfy DNS record and the ntfy HTTP check).

ntfy went with it - it ran nowhere else - so services/ntfy is deleted,
subdomains.ntfy and ntfy_topic are gone from group_vars, and the ntfy playbook
is out of site.yml. ntfy_topic already had no readers: the three infra/4xx plays
that used it were deleted when their checks were superseded.

Two things this exposed.

services/ntfy/deploy_ntfy_playbook.yml was pointing at the WRONG MACHINE. It
said `hosts: observability`, which resolves to the host `monitoring`
(64.226.70.190) - but ntfy ran on watchtower, and ntfy.contrapeso.xyz pointed
there. Running it would have installed ntfy on the new VPS. Moot now, but it is
the same stale-identity failure as the rest: the group meant watchtower when the
play was written, and nobody revisited it when the group changed. Watchtower was
in [vps] and NO role group at all, while running caddy, ntfy and Uptime Kuma -
nothing in the repo managed any of it.

More seriously: ntfy-emergency-app on vipy (avisame.contrapeso.xyz) sends its
notifications to https://ntfy.contrapeso.xyz, topic "emergencia". Destroying
watchtower breaks it, and it is an EMERGENCY notifier - it would fail silently
at exactly the moment it matters. That is NOT resolved here, deliberately:
standing ntfy up elsewhere, pointing at ntfy.sh, or retiring the app are all
decisions, not cleanups.

What this change does is make the break impossible to miss. The URL was derived
from subdomains.ntfy, so deleting that would have turned it into an undefined
variable buried in a template. It is now an explicit ntfy_service_url in the
app's own vars, still holding the old value, with the three options written
above it. The ntfy credentials stay in the vault because that app still needs
them - the vault was restored from HEAD and only watchtower's push token
removed, rather than re-handling the plaintext.

Verified: no reference to watchtower or its IP anywhere in the repo; Gatus down
from 91 to 85 endpoints, 85 UP, 0 DOWN.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 11:26:47 +02:00

86 lines
5.2 KiB
YAML

---
# Everything, in the order it has to happen.
#
# This file is a TABLE OF CONTENTS, not a second source of truth. It says what
# runs and in what order. It does NOT say which hosts get what — that stays on
# the `hosts:` line inside each playbook, exactly where it is today. Nothing
# moves; this file only makes the set readable in one place.
#
# What runs on a host? ansible-playbook site.yml --limit <host> --list-hosts
# Who gets thing Y? the `hosts:` line in Y's own playbook
# What is a host? ansible-inventory --graph
#
# Run a slice with --limit, or run one playbook directly as before. Nothing here
# changes how any individual playbook behaves.
# ── Baseline: every managed machine ─────────────────────────────────────────
- import_playbook: infra/01_user_and_access_setup_playbook.yml
- import_playbook: infra/02_firewall_and_fail2ban_playbook.yml
- import_playbook: infra/900_install_rsync.yml
- import_playbook: infra/920_join_headscale_mesh.yml
# Idempotent and kept permanently: guarantees a rebuilt or restored host cannot
# quietly bring the Uptime-Kuma-era monitoring back.
- import_playbook: infra/409_remove_legacy_monitoring.yml
# ── Monitoring ──────────────────────────────────────────────────────────────
# Gatus first: the three plays below register endpoints with it, and registering
# against a host that is not serving yet would simply fail.
- import_playbook: services/gatus/deploy_gatus_playbook.yml
- import_playbook: infra/400_host_monitoring.yml
- import_playbook: infra/401_service_monitoring.yml
- import_playbook: infra/402_public_monitoring.yml
# Registers where the per-service probes report. The probes themselves are
# deployed by each service's own playbook further down; the endpoints must exist
# before the first push arrives.
- import_playbook: infra/403_service_probe_registration.yml
# 910_docker says `hosts: managed`, but only 5 of 11 managed hosts have or need
# Docker. Left out until it has a [docker] group — see the note in PLAN_7.
# ── The hypervisor ──────────────────────────────────────────────────────────
- import_playbook: infra/nodito/31_proxmox_community_repos_playbook.yml
- import_playbook: infra/nodito/32_zfs_pool_setup_playbook.yml
- import_playbook: infra/nodito/34_nut_ups_setup_playbook.yml
# ── Reverse proxy, before anything that registers a vhost ───────────────────
- import_playbook: services/caddy_playbook.yml
# ── Services ────────────────────────────────────────────────────────────────
- import_playbook: services/bitcoin-knots/deploy_bitcoin_knots_playbook.yml
- import_playbook: services/fulcrum/deploy_fulcrum_playbook.yml
- import_playbook: services/datum-gateway/deploy_datum_gateway_playbook.yml
- import_playbook: services/mempool/deploy_mempool_playbook.yml
- import_playbook: services/memos/deploy_memos_playbook.yml
- import_playbook: services/forgejo-runner/deploy_forgejo_runner_playbook.yml
- import_playbook: services/phoenixd/deploy_phoenixd_playbook.yml
- import_playbook: services/headscale/deploy_headscale_playbook.yml
- import_playbook: services/vaultwarden/deploy_vaultwarden_playbook.yml
- import_playbook: services/forgejo/deploy_forgejo_playbook.yml
- import_playbook: services/lnbits/deploy_lnbits_playbook.yml
- import_playbook: services/ntfy-emergency-app/deploy_ntfy_emergency_app_playbook.yml
- import_playbook: services/personal-blog/deploy_personal_blog_playbook.yml
# ── Backups: each source dumps itself, the box pulls ────────────────────────
- import_playbook: services/headscale/setup_backup_headscale.yml
- import_playbook: services/vaultwarden/setup_backup_vaultwarden.yml
- import_playbook: services/forgejo/setup_backup_forgejo.yml
- import_playbook: services/lnbits/setup_backup_lnbits.yml
- import_playbook: services/memos/setup_backup_memos.yml
- import_playbook: playbooks/backups.yml
# Deliberately not here. Every playbook in the repo is either imported above or
# listed below, so this file accounts for all of them:
#
# infra/910_docker_playbook.yml says `hosts: managed`, but Docker is on 5
# of 11 managed hosts and those 5 are exactly
# the ones that need it. Running it would
# install Docker on the Bitcoin node and the
# hypervisor. Needs a [docker] group first.
#
# infra/nodito/30_proxmox_bootstrap one-shot: bare-metal bootstrap, run once
# infra/nodito/33_..._cloud_template one-shot: builds the VM template
#
#
# services/vaultwarden/disable_ deliberate manual actions, not convergence
# vaultwarden_sign_ups_playbook.yml
# services/personal-blog/setup_
# deploy_alias_lapy.yml