personal_infra/ansible/group_vars/all/main.yml

61 lines
2.4 KiB
YAML
Raw Normal View History

2026-09-11 22:17:13 +02:00
new_user: counterweight
ssh_port: 22
allow_ssh_from: "any"
root_domain: contrapeso.xyz
# Uptime Kuma was decommissioned on 2026-09-11. The monitoring blocks in the
# playbooks are kept deliberately — the check logic is meant to be rewired to
# whatever replaces it. This flag keeps them inert until then. See archive/uptime_kuma/.
uptime_kuma_enabled: false
2026-09-12 16:20:42 +02:00
# age recipient for all backup artefacts
age_backup_recipient: "age192wwdaseqej2ggwyp884gtm05c396anp7chr0vr8m47g50fahpyqr9fsza"
# Public key small-backups-box pulls with
# Authorised on each source host for an unprivileged, dedicated user only
backup_pull_public_key: "ssh-ed25519 AAAAC3NzaC1lZDI1NTE5AAAAIOfIixKMhA9z+Nvyx6ToZIniC8aEgyiInRiboaTTemgX offsite-backup-pull"
ansible: delete the duplicated vars files, move globals to group_vars/all Three files existed only as second copies of things group_vars/all already auto-loads, and 34 playbooks named them in vars_files: - which outranks group_vars, so the copies won. The day someone edited one and not the other, those plays would silently keep the stale value. infra_vars.yml was already drifting: group_vars/all/main.yml had grown age_backup_recipient and backup_pull_public_key that it lacked. infra_vars.yml - a strict subset of group_vars/all/main.yml infra_secrets.yml - decrypts byte-identical to group_vars/all/vault.yml infra_secrets.yml.example - documented Uptime Kuma credentials as the reason the file exists, which stopped being true Deleted, along with 62 vars_files entries across 34 playbooks (12 of which named ../../group_vars/all/main.yml directly - same defect, a vars_files entry duplicating an auto-loaded file at higher precedence than the file itself). Checked before touching anything: infra_secrets.yml was listed LAST in 10 plays, after services_config.yml, so removal would flip precedence if the two shared a key. They share none, and neither does services_config.yml with group_vars/all/main.yml, so the removal is provably inert. services_config.yml was the last one standing. It held four unrelated things: caddy_sites_dir - an identical copy of roles/caddy_site/defaults/. Deleted; the role default is now the only one. *.tailscale_hostname (x3) - a THIRD copy of each box's identity, which inventory.ini already holds as ansible_host. Deleted. Edge plays now read hostvars['<host>'].ansible_host - verified an edge play resolves that with nothing loaded and the other host in no play. Three copies of one name is how bitcoin_rpc_host ended up labelled "knots_box" while pointing at fulcrum-box. subdomains, ntfy topic, - genuinely global: their readers span managed, headscale namespace monitoring, vpn_control and edge, so no single group covers them. Moved to group_vars/all/main.yml where they auto-load. The ntfy_topic and headscale_namespace indirection through service_settings collapses to the global name. the four cross-host ports - the only entries with a real justification. Left in place; they move in the next commit. Also dead, all Uptime Kuma residue or duplication: phoenixd_monitor_name, forgejo_runner healthcheck_timeout_seconds/retries, fulcrum_tailscale_hostname, and bitcoin_knots_version - the last being a v-prefixed copy of bitcoin_knots_version_short that nothing read, two hand-maintained copies of one version string. Corrected a false comment: services_config.yml claimed the uptime_kuma subdomain "no longer resolves to anything". It resolves to 164.92.239.72 and answers HTTP 302, and 11 playbooks still template it. Same wrong premise as PLAN_3. Verification: all 37 playbooks' --list-tasks output is byte-identical before and after. A probe resolving all 22 values services_config.yml used to supply returns 21 identical and one intended deletion (caddy_sites_dir, now role-only - confirmed the role still resolves it: "Ensure Caddy sites-enabled directory exists" comes back ok against the real path). memos check-diff identical before and after. Syntax passes on every playbook. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 20:58:46 +02:00
# ─────────────────────────────────────────────────────────────────────────────
# Subdomains. Global because the edge host proxies for services that live on
# other machines, so no single inventory group covers the readers. Combine with
# root_domain above to build an FQDN.
#
# Moved here from services_config.yml, which 30 plays had to remember to name in
# vars_files: - a file everyone must opt into is a file someone will forget.
# ─────────────────────────────────────────────────────────────────────────────
subdomains:
gatus: deploy on prd-monitoring, behind Caddy basic auth First step of replacing Uptime Kuma and ntfy. Gatus runs on the new observability VPS, fronted by Caddy at status.contrapeso.xyz. Deployed as the upstream container image, not built from source. The role did build from source first - their Dockerfile is a bare `CGO_ENABLED=0 go build`, the Vue dashboard is compiled in via `//go:embed static` in web/static.go, and CGO can stay off because the sqlite driver is pure-Go modernc.org/sqlite - but that produces a binary upstream never ran, and it meant compiling the AWS SDK and gRPC on the smallest box in the estate. That load was heavy enough that unrelated Ansible tasks timed out while it ran. The cost of the container is a daemon on the machine whose job is to notice when everything else breaks; that trade is made deliberately and is written down in the role README. Pinned by DIGEST, not tag. A tag is mutable - v5.36.0 can be repushed - so pinning it alone is a weaker promise than it looks: gatus_image: "ghcr.io/twin/gatus@sha256:c5f210d0..." `docker compose pull` now either fetches exactly the reviewed image or fails. gatus_version is kept beside it only so a human can read the release; the two move together. The image is FROM scratch, so it has no /etc/passwd and its default user is root. The container runs as 10001:10001 with the host data dir owned to match, plus read_only, cap_drop ALL, and no-new-privileges. NET_RAW is added back only when gatus_allow_icmp, so the capability for icmp:// checks is a visible grant rather than something inherited from running as root. Config is a DIRECTORY, not a file. Gatus merges every *.yaml under GATUS_CONFIG_PATH - maps deep-merge, lists append - so the role owns 00-base.yaml (web, storage, ui, alerting, security) and each service will drop its own file into endpoints/, the same shape as caddy_site. A primitive defined twice is ambiguous and upstream refuses it, so anything that is not a list lives in the base file and nowhere else. Two bugs the deploy caught: * Gatus panics on a config with no endpoints ("configuration should contain at least one endpoint or suite"), so "install now, add endpoints later" is not a valid state. The role ships endpoints/00-self.yaml checking its own /health. Less circular than it looks: it proves the directory merged, the listener serves, and storage accepted a write. * web.address was carried over from the systemd design as 127.0.0.1. Inside a container that is the CONTAINER's loopback, which docker-proxy cannot reach - gatus came up healthy, self-check passing, while every connection to the published port was refused. It now always binds 0.0.0.0 inside the container; the isolation comes from publishing to 127.0.0.1 on the host. Auth is done at the edge, NOT with Gatus's own security.basic. Reading api/api.go, that middleware protects exactly four routes - the statuses endpoints. Everything else is registered on the unprotected router, including /api/v1/config, every badge, and /api/v1/endpoints/:key/uptimes/:duration and .../response-times/:duration/history, which return real data to anyone who can guess a key ("<group>_<name>"). Verified against the live instance: all seven routes returned 200 unauthenticated, and /uptimes/24h returned "1.000000". So the vhost uses caddy_site_body with a path carve-out rather than caddy_site_basic_auth, which has no way to exempt a path. The external-endpoint push API must NOT sit behind basic auth: it authenticates with `Authorization: Bearer <token>`, and basic auth wants the same header. It is not unauthenticated - the handler 401s on a missing prefix, an empty token, or a token that does not match that endpoint's own. Verified end to end. All seven previously-open routes now 401. The push path distinguishes cleanly: POST with no auth gets Gatus's own "invalid Authorization header" with NO WWW-Authenticate; POST with a bogus Bearer gets 404 (key looked up, no external endpoints yet); GET on the same path gets Caddy's 401 with WWW-Authenticate: Basic, so the exemption is scoped to POST alone. The self-check still passes because it polls localhost inside the container and never traverses Caddy. The host itself was rebuilt from scratch: 01 (ok=9 changed=8), 02 (ok=12 changed=6), 910_docker (--limit, since that playbook still wrongly claims all of `managed` needs Docker), caddy (ok=13 changed=8), gatus (ok=18 changed=2). Not done here: gatus_alerting is still {} - valid, and every condition is evaluated and recorded, there is just nowhere to shout until a provider is chosen to replace ntfy. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 22:29:03 +02:00
# Monitoring
gatus: status
ansible: delete the duplicated vars files, move globals to group_vars/all Three files existed only as second copies of things group_vars/all already auto-loads, and 34 playbooks named them in vars_files: - which outranks group_vars, so the copies won. The day someone edited one and not the other, those plays would silently keep the stale value. infra_vars.yml was already drifting: group_vars/all/main.yml had grown age_backup_recipient and backup_pull_public_key that it lacked. infra_vars.yml - a strict subset of group_vars/all/main.yml infra_secrets.yml - decrypts byte-identical to group_vars/all/vault.yml infra_secrets.yml.example - documented Uptime Kuma credentials as the reason the file exists, which stopped being true Deleted, along with 62 vars_files entries across 34 playbooks (12 of which named ../../group_vars/all/main.yml directly - same defect, a vars_files entry duplicating an auto-loaded file at higher precedence than the file itself). Checked before touching anything: infra_secrets.yml was listed LAST in 10 plays, after services_config.yml, so removal would flip precedence if the two shared a key. They share none, and neither does services_config.yml with group_vars/all/main.yml, so the removal is provably inert. services_config.yml was the last one standing. It held four unrelated things: caddy_sites_dir - an identical copy of roles/caddy_site/defaults/. Deleted; the role default is now the only one. *.tailscale_hostname (x3) - a THIRD copy of each box's identity, which inventory.ini already holds as ansible_host. Deleted. Edge plays now read hostvars['<host>'].ansible_host - verified an edge play resolves that with nothing loaded and the other host in no play. Three copies of one name is how bitcoin_rpc_host ended up labelled "knots_box" while pointing at fulcrum-box. subdomains, ntfy topic, - genuinely global: their readers span managed, headscale namespace monitoring, vpn_control and edge, so no single group covers them. Moved to group_vars/all/main.yml where they auto-load. The ntfy_topic and headscale_namespace indirection through service_settings collapses to the global name. the four cross-host ports - the only entries with a real justification. Left in place; they move in the next commit. Also dead, all Uptime Kuma residue or duplication: phoenixd_monitor_name, forgejo_runner healthcheck_timeout_seconds/retries, fulcrum_tailscale_hostname, and bitcoin_knots_version - the last being a v-prefixed copy of bitcoin_knots_version_short that nothing read, two hand-maintained copies of one version string. Corrected a false comment: services_config.yml claimed the uptime_kuma subdomain "no longer resolves to anything". It resolves to 164.92.239.72 and answers HTTP 302, and 11 playbooks still template it. Same wrong premise as PLAN_3. Verification: all 37 playbooks' --list-tasks output is byte-identical before and after. A probe resolving all 22 values services_config.yml used to supply returns 21 identical and one intended deletion (caddy_sites_dir, now role-only - confirmed the role still resolves it: "Ensure Caddy sites-enabled directory exists" comes back ok against the real path). memos check-diff identical before and after. Syntax passes on every playbook. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 20:58:46 +02:00
ntfy: ntfy
# Uptime Kuma IS still running and this subdomain DOES resolve
# (164.92.239.72, HTTP 302). Only the Ansible code and the vault credentials
# were retired. A comment here previously claimed the opposite.
uptime_kuma: uptime
# VPN infrastructure (spacey)
headscale: headscale
# Core services (vipy)
vaultwarden: vault
forgejo: forgejo
lnbits: wallet
# Secondary services (vipy)
ntfy_emergency_app: avisame
personal_blog: pablohere
# Memos (memos-box)
memos: memos
# Mempool block explorer (mempool-box, proxied via vipy)
mempool: mempool
# DATUM Gateway dashboard (knots-box, proxied via vipy)
datum_gateway: datum
# Read by plays targeting managed, monitoring, vpn_control and edge - no one
# group covers them, so these are global rather than group_vars/<group>.
ntfy_topic: alerts
headscale_namespace: counter-net