Three files existed only as second copies of things group_vars/all already
auto-loads, and 34 playbooks named them in vars_files: - which outranks
group_vars, so the copies won. The day someone edited one and not the other,
those plays would silently keep the stale value. infra_vars.yml was already
drifting: group_vars/all/main.yml had grown age_backup_recipient and
backup_pull_public_key that it lacked.
infra_vars.yml - a strict subset of group_vars/all/main.yml
infra_secrets.yml - decrypts byte-identical to group_vars/all/vault.yml
infra_secrets.yml.example - documented Uptime Kuma credentials as the reason
the file exists, which stopped being true
Deleted, along with 62 vars_files entries across 34 playbooks (12 of which
named ../../group_vars/all/main.yml directly - same defect, a vars_files entry
duplicating an auto-loaded file at higher precedence than the file itself).
Checked before touching anything: infra_secrets.yml was listed LAST in 10 plays,
after services_config.yml, so removal would flip precedence if the two shared a
key. They share none, and neither does services_config.yml with
group_vars/all/main.yml, so the removal is provably inert.
services_config.yml was the last one standing. It held four unrelated things:
caddy_sites_dir - an identical copy of roles/caddy_site/defaults/.
Deleted; the role default is now the only one.
*.tailscale_hostname (x3) - a THIRD copy of each box's identity, which
inventory.ini already holds as ansible_host.
Deleted. Edge plays now read
hostvars['<host>'].ansible_host - verified an
edge play resolves that with nothing loaded and
the other host in no play. Three copies of one
name is how bitcoin_rpc_host ended up labelled
"knots_box" while pointing at fulcrum-box.
subdomains, ntfy topic, - genuinely global: their readers span managed,
headscale namespace monitoring, vpn_control and edge, so no single
group covers them. Moved to group_vars/all/main.yml
where they auto-load. The ntfy_topic and
headscale_namespace indirection through
service_settings collapses to the global name.
the four cross-host ports - the only entries with a real justification.
Left in place; they move in the next commit.
Also dead, all Uptime Kuma residue or duplication:
phoenixd_monitor_name, forgejo_runner healthcheck_timeout_seconds/retries,
fulcrum_tailscale_hostname, and bitcoin_knots_version - the last being a
v-prefixed copy of bitcoin_knots_version_short that nothing read, two
hand-maintained copies of one version string.
Corrected a false comment: services_config.yml claimed the uptime_kuma subdomain
"no longer resolves to anything". It resolves to 164.92.239.72 and answers HTTP
302, and 11 playbooks still template it. Same wrong premise as PLAN_3.
Verification: all 37 playbooks' --list-tasks output is byte-identical before and
after. A probe resolving all 22 values services_config.yml used to supply returns
21 identical and one intended deletion (caddy_sites_dir, now role-only - confirmed
the role still resolves it: "Ensure Caddy sites-enabled directory exists" comes
back ok against the real path). memos check-diff identical before and after.
Syntax passes on every playbook.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|---|---|---|
| .. | ||
| defaults | ||
| tasks | ||
| templates | ||
| README.md | ||
mempool
Deploys the Mempool block explorer as a three-container
Docker Compose stack — MariaDB, backend, frontend — on mempool-box, and keeps a
health check on each.
Converted from deploy_mempool_playbook.yml (745 lines) under Plan 6. The
playbook is now 37 lines: this role, plus a second play that publishes the
frontend through Caddy on the edge host.
Phases
docker.yml |
Docker engine: repo, key, packages, service |
deploy.yml |
directories, docker-compose.yml, pull, up, wait-for-healthy |
healthcheck.yml |
three check scripts, three services, three timers |
Three health checks, not one
Mempool is three moving parts and knowing which one is down is the point, so
each gets its own check, unit and timer, driven by the mempool_healthchecks
list:
| checks | |
|---|---|
mariadb |
docker inspect health status of mempool-db |
backend |
GET /api/v1/backend-info |
frontend |
GET / |
Each records its answer in its exit code, which systemd keeps:
systemctl is-failed mempool-backend-healthcheck.service. Reporting elsewhere
is one field per check, push_url, and is the plug-in point for whatever
monitoring exists. Empty means check, exit honestly, report nowhere. The URLs
are credentials, so callers pass them from the vault.
Nothing here is specific to a monitoring product. The embedded Python that
created monitors over the Uptime Kuma API, the /tmp credentials file, the
push-URL file read back and parsed, and three systemd Environment= rewrites
are gone.
MariaDB owns its own data directory
{{ mempool_mysql_dir }} is bind-mounted into the container, which runs as uid
999 and must create files there. The playbook this replaced declared
owner: "{{ ansible_user }}" (1000) on it, which had drifted from reality ever
since the containers were created — unnoticed, because the playbook had not been
run since.
That was not academic. The first real run of this role pulled a newer
mariadb:10.11 and recreated mempool-db; had the chown still been in place,
MariaDB would have come back to a directory it could not write. The role now
ensures the directory exists and leaves ownership to the container.
mempool_frontend_port lives in services_config.yml
Two hosts need it: this role deploys the frontend on mempool-box, and the Caddy
play proxies to it from the edge host. A role default is invisible to the second
play, so the value lives in service_settings.mempool.frontend_port and the role
default derives from it.
Expect changed=2 on a converged host
Pull Mempool images and Deploy Mempool containers with docker compose are
bare command: tasks with no changed_when, so they always report changed.
That is the idempotent floor, not drift. Everything else reports ok.
mariadb:10.11 is a moving tag, so a run can pull a newer patch release and
recreate the database container. Pin it if that is not what you want.