2026-09-11 22:17:13 +02:00
|
|
|
new_user: counterweight
|
|
|
|
|
ssh_port: 22
|
|
|
|
|
allow_ssh_from: "any"
|
|
|
|
|
root_domain: contrapeso.xyz
|
2026-09-11 22:43:55 +02:00
|
|
|
|
|
|
|
|
# Uptime Kuma was decommissioned on 2026-09-11. The monitoring blocks in the
|
|
|
|
|
# playbooks are kept deliberately — the check logic is meant to be rewired to
|
|
|
|
|
# whatever replaces it. This flag keeps them inert until then. See archive/uptime_kuma/.
|
|
|
|
|
uptime_kuma_enabled: false
|
2026-09-12 16:20:42 +02:00
|
|
|
|
|
|
|
|
# age recipient for all backup artefacts
|
|
|
|
|
age_backup_recipient: "age192wwdaseqej2ggwyp884gtm05c396anp7chr0vr8m47g50fahpyqr9fsza"
|
|
|
|
|
|
|
|
|
|
# Public key small-backups-box pulls with
|
|
|
|
|
# Authorised on each source host for an unprivileged, dedicated user only
|
|
|
|
|
backup_pull_public_key: "ssh-ed25519 AAAAC3NzaC1lZDI1NTE5AAAAIOfIixKMhA9z+Nvyx6ToZIniC8aEgyiInRiboaTTemgX offsite-backup-pull"
|
ansible: delete the duplicated vars files, move globals to group_vars/all
Three files existed only as second copies of things group_vars/all already
auto-loads, and 34 playbooks named them in vars_files: - which outranks
group_vars, so the copies won. The day someone edited one and not the other,
those plays would silently keep the stale value. infra_vars.yml was already
drifting: group_vars/all/main.yml had grown age_backup_recipient and
backup_pull_public_key that it lacked.
infra_vars.yml - a strict subset of group_vars/all/main.yml
infra_secrets.yml - decrypts byte-identical to group_vars/all/vault.yml
infra_secrets.yml.example - documented Uptime Kuma credentials as the reason
the file exists, which stopped being true
Deleted, along with 62 vars_files entries across 34 playbooks (12 of which
named ../../group_vars/all/main.yml directly - same defect, a vars_files entry
duplicating an auto-loaded file at higher precedence than the file itself).
Checked before touching anything: infra_secrets.yml was listed LAST in 10 plays,
after services_config.yml, so removal would flip precedence if the two shared a
key. They share none, and neither does services_config.yml with
group_vars/all/main.yml, so the removal is provably inert.
services_config.yml was the last one standing. It held four unrelated things:
caddy_sites_dir - an identical copy of roles/caddy_site/defaults/.
Deleted; the role default is now the only one.
*.tailscale_hostname (x3) - a THIRD copy of each box's identity, which
inventory.ini already holds as ansible_host.
Deleted. Edge plays now read
hostvars['<host>'].ansible_host - verified an
edge play resolves that with nothing loaded and
the other host in no play. Three copies of one
name is how bitcoin_rpc_host ended up labelled
"knots_box" while pointing at fulcrum-box.
subdomains, ntfy topic, - genuinely global: their readers span managed,
headscale namespace monitoring, vpn_control and edge, so no single
group covers them. Moved to group_vars/all/main.yml
where they auto-load. The ntfy_topic and
headscale_namespace indirection through
service_settings collapses to the global name.
the four cross-host ports - the only entries with a real justification.
Left in place; they move in the next commit.
Also dead, all Uptime Kuma residue or duplication:
phoenixd_monitor_name, forgejo_runner healthcheck_timeout_seconds/retries,
fulcrum_tailscale_hostname, and bitcoin_knots_version - the last being a
v-prefixed copy of bitcoin_knots_version_short that nothing read, two
hand-maintained copies of one version string.
Corrected a false comment: services_config.yml claimed the uptime_kuma subdomain
"no longer resolves to anything". It resolves to 164.92.239.72 and answers HTTP
302, and 11 playbooks still template it. Same wrong premise as PLAN_3.
Verification: all 37 playbooks' --list-tasks output is byte-identical before and
after. A probe resolving all 22 values services_config.yml used to supply returns
21 identical and one intended deletion (caddy_sites_dir, now role-only - confirmed
the role still resolves it: "Ensure Caddy sites-enabled directory exists" comes
back ok against the real path). memos check-diff identical before and after.
Syntax passes on every playbook.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 20:58:46 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
# ─────────────────────────────────────────────────────────────────────────────
|
|
|
|
|
# Subdomains. Global because the edge host proxies for services that live on
|
|
|
|
|
# other machines, so no single inventory group covers the readers. Combine with
|
|
|
|
|
# root_domain above to build an FQDN.
|
|
|
|
|
#
|
|
|
|
|
# Moved here from services_config.yml, which 30 plays had to remember to name in
|
|
|
|
|
# vars_files: - a file everyone must opt into is a file someone will forget.
|
|
|
|
|
# ─────────────────────────────────────────────────────────────────────────────
|
|
|
|
|
subdomains:
|
gatus: deploy on prd-monitoring, behind Caddy basic auth
First step of replacing Uptime Kuma and ntfy. Gatus runs on the new
observability VPS, fronted by Caddy at status.contrapeso.xyz.
Deployed as the upstream container image, not built from source. The role did
build from source first - their Dockerfile is a bare `CGO_ENABLED=0 go build`,
the Vue dashboard is compiled in via `//go:embed static` in web/static.go, and
CGO can stay off because the sqlite driver is pure-Go modernc.org/sqlite - but
that produces a binary upstream never ran, and it meant compiling the AWS SDK
and gRPC on the smallest box in the estate. That load was heavy enough that
unrelated Ansible tasks timed out while it ran. The cost of the container is a
daemon on the machine whose job is to notice when everything else breaks; that
trade is made deliberately and is written down in the role README.
Pinned by DIGEST, not tag. A tag is mutable - v5.36.0 can be repushed - so
pinning it alone is a weaker promise than it looks:
gatus_image: "ghcr.io/twin/gatus@sha256:c5f210d0..."
`docker compose pull` now either fetches exactly the reviewed image or fails.
gatus_version is kept beside it only so a human can read the release; the two
move together.
The image is FROM scratch, so it has no /etc/passwd and its default user is
root. The container runs as 10001:10001 with the host data dir owned to match,
plus read_only, cap_drop ALL, and no-new-privileges. NET_RAW is added back only
when gatus_allow_icmp, so the capability for icmp:// checks is a visible grant
rather than something inherited from running as root.
Config is a DIRECTORY, not a file. Gatus merges every *.yaml under
GATUS_CONFIG_PATH - maps deep-merge, lists append - so the role owns
00-base.yaml (web, storage, ui, alerting, security) and each service will drop
its own file into endpoints/, the same shape as caddy_site. A primitive defined
twice is ambiguous and upstream refuses it, so anything that is not a list lives
in the base file and nowhere else.
Two bugs the deploy caught:
* Gatus panics on a config with no endpoints ("configuration should contain at
least one endpoint or suite"), so "install now, add endpoints later" is not
a valid state. The role ships endpoints/00-self.yaml checking its own
/health. Less circular than it looks: it proves the directory merged, the
listener serves, and storage accepted a write.
* web.address was carried over from the systemd design as 127.0.0.1. Inside a
container that is the CONTAINER's loopback, which docker-proxy cannot reach
- gatus came up healthy, self-check passing, while every connection to the
published port was refused. It now always binds 0.0.0.0 inside the
container; the isolation comes from publishing to 127.0.0.1 on the host.
Auth is done at the edge, NOT with Gatus's own security.basic. Reading
api/api.go, that middleware protects exactly four routes - the statuses
endpoints. Everything else is registered on the unprotected router, including
/api/v1/config, every badge, and /api/v1/endpoints/:key/uptimes/:duration and
.../response-times/:duration/history, which return real data to anyone who can
guess a key ("<group>_<name>"). Verified against the live instance: all seven
routes returned 200 unauthenticated, and /uptimes/24h returned "1.000000".
So the vhost uses caddy_site_body with a path carve-out rather than
caddy_site_basic_auth, which has no way to exempt a path. The external-endpoint
push API must NOT sit behind basic auth: it authenticates with
`Authorization: Bearer <token>`, and basic auth wants the same header. It is not
unauthenticated - the handler 401s on a missing prefix, an empty token, or a
token that does not match that endpoint's own.
Verified end to end. All seven previously-open routes now 401. The push path
distinguishes cleanly: POST with no auth gets Gatus's own "invalid Authorization
header" with NO WWW-Authenticate; POST with a bogus Bearer gets 404 (key looked
up, no external endpoints yet); GET on the same path gets Caddy's 401 with
WWW-Authenticate: Basic, so the exemption is scoped to POST alone. The
self-check still passes because it polls localhost inside the container and
never traverses Caddy.
The host itself was rebuilt from scratch: 01 (ok=9 changed=8), 02 (ok=12
changed=6), 910_docker (--limit, since that playbook still wrongly claims all of
`managed` needs Docker), caddy (ok=13 changed=8), gatus (ok=18 changed=2).
Not done here: gatus_alerting is still {} - valid, and every condition is
evaluated and recorded, there is just nowhere to shout until a provider is
chosen to replace ntfy.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 22:29:03 +02:00
|
|
|
# Monitoring
|
|
|
|
|
gatus: status
|
ansible: delete the duplicated vars files, move globals to group_vars/all
Three files existed only as second copies of things group_vars/all already
auto-loads, and 34 playbooks named them in vars_files: - which outranks
group_vars, so the copies won. The day someone edited one and not the other,
those plays would silently keep the stale value. infra_vars.yml was already
drifting: group_vars/all/main.yml had grown age_backup_recipient and
backup_pull_public_key that it lacked.
infra_vars.yml - a strict subset of group_vars/all/main.yml
infra_secrets.yml - decrypts byte-identical to group_vars/all/vault.yml
infra_secrets.yml.example - documented Uptime Kuma credentials as the reason
the file exists, which stopped being true
Deleted, along with 62 vars_files entries across 34 playbooks (12 of which
named ../../group_vars/all/main.yml directly - same defect, a vars_files entry
duplicating an auto-loaded file at higher precedence than the file itself).
Checked before touching anything: infra_secrets.yml was listed LAST in 10 plays,
after services_config.yml, so removal would flip precedence if the two shared a
key. They share none, and neither does services_config.yml with
group_vars/all/main.yml, so the removal is provably inert.
services_config.yml was the last one standing. It held four unrelated things:
caddy_sites_dir - an identical copy of roles/caddy_site/defaults/.
Deleted; the role default is now the only one.
*.tailscale_hostname (x3) - a THIRD copy of each box's identity, which
inventory.ini already holds as ansible_host.
Deleted. Edge plays now read
hostvars['<host>'].ansible_host - verified an
edge play resolves that with nothing loaded and
the other host in no play. Three copies of one
name is how bitcoin_rpc_host ended up labelled
"knots_box" while pointing at fulcrum-box.
subdomains, ntfy topic, - genuinely global: their readers span managed,
headscale namespace monitoring, vpn_control and edge, so no single
group covers them. Moved to group_vars/all/main.yml
where they auto-load. The ntfy_topic and
headscale_namespace indirection through
service_settings collapses to the global name.
the four cross-host ports - the only entries with a real justification.
Left in place; they move in the next commit.
Also dead, all Uptime Kuma residue or duplication:
phoenixd_monitor_name, forgejo_runner healthcheck_timeout_seconds/retries,
fulcrum_tailscale_hostname, and bitcoin_knots_version - the last being a
v-prefixed copy of bitcoin_knots_version_short that nothing read, two
hand-maintained copies of one version string.
Corrected a false comment: services_config.yml claimed the uptime_kuma subdomain
"no longer resolves to anything". It resolves to 164.92.239.72 and answers HTTP
302, and 11 playbooks still template it. Same wrong premise as PLAN_3.
Verification: all 37 playbooks' --list-tasks output is byte-identical before and
after. A probe resolving all 22 values services_config.yml used to supply returns
21 identical and one intended deletion (caddy_sites_dir, now role-only - confirmed
the role still resolves it: "Ensure Caddy sites-enabled directory exists" comes
back ok against the real path). memos check-diff identical before and after.
Syntax passes on every playbook.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 20:58:46 +02:00
|
|
|
ntfy: ntfy
|
|
|
|
|
# Uptime Kuma IS still running and this subdomain DOES resolve
|
|
|
|
|
# (164.92.239.72, HTTP 302). Only the Ansible code and the vault credentials
|
|
|
|
|
# were retired. A comment here previously claimed the opposite.
|
|
|
|
|
uptime_kuma: uptime
|
|
|
|
|
|
|
|
|
|
# VPN infrastructure (spacey)
|
|
|
|
|
headscale: headscale
|
|
|
|
|
|
|
|
|
|
# Core services (vipy)
|
|
|
|
|
vaultwarden: vault
|
|
|
|
|
forgejo: forgejo
|
|
|
|
|
lnbits: wallet
|
|
|
|
|
|
|
|
|
|
# Secondary services (vipy)
|
|
|
|
|
ntfy_emergency_app: avisame
|
|
|
|
|
personal_blog: pablohere
|
|
|
|
|
|
|
|
|
|
# Memos (memos-box)
|
|
|
|
|
memos: memos
|
|
|
|
|
|
|
|
|
|
# Mempool block explorer (mempool-box, proxied via vipy)
|
|
|
|
|
mempool: mempool
|
|
|
|
|
|
|
|
|
|
# DATUM Gateway dashboard (knots-box, proxied via vipy)
|
|
|
|
|
datum_gateway: datum
|
|
|
|
|
|
|
|
|
|
# Read by plays targeting managed, monitoring, vpn_control and edge - no one
|
|
|
|
|
# group covers them, so these are global rather than group_vars/<group>.
|
|
|
|
|
ntfy_topic: alerts
|
|
|
|
|
headscale_namespace: counter-net
|
2026-09-14 09:53:15 +02:00
|
|
|
|
|
|
|
|
# ─────────────────────────────────────────────────────────────────────────────
|
|
|
|
|
# Domains whose registration expiry is monitored (infra/402_public_monitoring).
|
|
|
|
|
#
|
|
|
|
|
# Registration renewal is a manual act at the registrar, and losing a domain is
|
|
|
|
|
# not recoverable in the way losing a host is - so these are checked daily and
|
|
|
|
|
# alarm with two weeks of runway.
|
|
|
|
|
#
|
|
|
|
|
# root_domain is the estate's own domain; the rest are domains we own that are
|
|
|
|
|
# served from it or from a host in the inventory.
|
|
|
|
|
# ─────────────────────────────────────────────────────────────────────────────
|
|
|
|
|
monitored_domains:
|
|
|
|
|
- "{{ root_domain }}"
|
|
|
|
|
- arbret.com
|