watchtower is being destroyed. Removed from [vps], with its host_vars, its push
token, and the six Gatus endpoints that referenced it (liveness, disk, two
systemd services, the ntfy DNS record and the ntfy HTTP check).
ntfy went with it - it ran nowhere else - so services/ntfy is deleted,
subdomains.ntfy and ntfy_topic are gone from group_vars, and the ntfy playbook
is out of site.yml. ntfy_topic already had no readers: the three infra/4xx plays
that used it were deleted when their checks were superseded.
Two things this exposed.
services/ntfy/deploy_ntfy_playbook.yml was pointing at the WRONG MACHINE. It
said `hosts: observability`, which resolves to the host `monitoring`
(64.226.70.190) - but ntfy ran on watchtower, and ntfy.contrapeso.xyz pointed
there. Running it would have installed ntfy on the new VPS. Moot now, but it is
the same stale-identity failure as the rest: the group meant watchtower when the
play was written, and nobody revisited it when the group changed. Watchtower was
in [vps] and NO role group at all, while running caddy, ntfy and Uptime Kuma -
nothing in the repo managed any of it.
More seriously: ntfy-emergency-app on vipy (avisame.contrapeso.xyz) sends its
notifications to https://ntfy.contrapeso.xyz, topic "emergencia". Destroying
watchtower breaks it, and it is an EMERGENCY notifier - it would fail silently
at exactly the moment it matters. That is NOT resolved here, deliberately:
standing ntfy up elsewhere, pointing at ntfy.sh, or retiring the app are all
decisions, not cleanups.
What this change does is make the break impossible to miss. The URL was derived
from subdomains.ntfy, so deleting that would have turned it into an undefined
variable buried in a template. It is now an explicit ntfy_service_url in the
app's own vars, still holding the old value, with the three options written
above it. The ntfy credentials stay in the vault because that app still needs
them - the vault was restored from HEAD and only watchtower's push token
removed, rather than re-handling the plaintext.
Verified: no reference to watchtower or its IP anywhere in the repo; Gatus down
from 91 to 85 endpoints, 85 UP, 0 DOWN.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
site.yml is a TABLE OF CONTENTS, not a second source of truth. It is 25
import_playbook: lines and comments - no `hosts:`, no `roles:`. Which hosts get
what stays on the `hosts:` line inside each playbook, exactly where it already
was; nothing moved. Every role is already wrapped in a thin playbook carrying
its own `hosts:` line, so there is no roles-vs-playbooks split to reconcile:
from here everything is a playbook.
What it buys:
What runs on a host? ansible-playbook site.yml --limit <host> --list-hosts
Who gets thing Y? the `hosts:` line in Y's own playbook
What is a host? ansible-inventory --graph
Note --list-hosts, not --list-tasks: the latter prints every play regardless of
--limit, so it will happily show you the bitcoin play under memos-box.
Nine playbooks are deliberately excluded and the file names every one with a
reason, so it accounts for all of them: the three infra/4xx monitoring plays
(still assert on the removed Uptime Kuma credentials and fail immediately),
910_docker (says `hosts: managed`, but Docker is on 5 of 11 managed hosts and
those 5 are exactly the ones that need it - running it installs Docker on the
Bitcoin node and the hypervisor), two nodito one-shots, the Kuma notification
setup, and two deliberate manual actions.
Writing it surfaced an inventory collision. There is a HOST named `monitoring`
in [vps] AND a group [monitoring], so Ansible warned and resolved `hosts:
monitoring` to the host:
[WARNING]: Found both group and host with same name: monitoring
The group is renamed to [observability]; the host keeps its name. [caddy:children]
and the two ntfy playbooks follow. Behaviour is unchanged - `hosts: monitoring`
already resolved to the host - but the ambiguity is gone and the warning with it.
Verified: inventory graph is warning-free, site.yml passes --syntax-check, and
per-host play counts are identical before and after the rename.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
services/caddy_playbook.yml was the one play still targeting a location
group (vps) rather than a role group. The two coincide today — vps is
exactly vipy, watchtower and spacey, the three hosts with
/etc/caddy/sites-enabled — but adding a fourth VPS that does not run
Caddy would have silently pulled it into the play.
[caddy:children] is edge + monitoring + vpn_control. Verified the play
selects the same three machines before and after.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>