Three files existed only as second copies of things group_vars/all already
auto-loads, and 34 playbooks named them in vars_files: - which outranks
group_vars, so the copies won. The day someone edited one and not the other,
those plays would silently keep the stale value. infra_vars.yml was already
drifting: group_vars/all/main.yml had grown age_backup_recipient and
backup_pull_public_key that it lacked.
infra_vars.yml - a strict subset of group_vars/all/main.yml
infra_secrets.yml - decrypts byte-identical to group_vars/all/vault.yml
infra_secrets.yml.example - documented Uptime Kuma credentials as the reason
the file exists, which stopped being true
Deleted, along with 62 vars_files entries across 34 playbooks (12 of which
named ../../group_vars/all/main.yml directly - same defect, a vars_files entry
duplicating an auto-loaded file at higher precedence than the file itself).
Checked before touching anything: infra_secrets.yml was listed LAST in 10 plays,
after services_config.yml, so removal would flip precedence if the two shared a
key. They share none, and neither does services_config.yml with
group_vars/all/main.yml, so the removal is provably inert.
services_config.yml was the last one standing. It held four unrelated things:
caddy_sites_dir - an identical copy of roles/caddy_site/defaults/.
Deleted; the role default is now the only one.
*.tailscale_hostname (x3) - a THIRD copy of each box's identity, which
inventory.ini already holds as ansible_host.
Deleted. Edge plays now read
hostvars['<host>'].ansible_host - verified an
edge play resolves that with nothing loaded and
the other host in no play. Three copies of one
name is how bitcoin_rpc_host ended up labelled
"knots_box" while pointing at fulcrum-box.
subdomains, ntfy topic, - genuinely global: their readers span managed,
headscale namespace monitoring, vpn_control and edge, so no single
group covers them. Moved to group_vars/all/main.yml
where they auto-load. The ntfy_topic and
headscale_namespace indirection through
service_settings collapses to the global name.
the four cross-host ports - the only entries with a real justification.
Left in place; they move in the next commit.
Also dead, all Uptime Kuma residue or duplication:
phoenixd_monitor_name, forgejo_runner healthcheck_timeout_seconds/retries,
fulcrum_tailscale_hostname, and bitcoin_knots_version - the last being a
v-prefixed copy of bitcoin_knots_version_short that nothing read, two
hand-maintained copies of one version string.
Corrected a false comment: services_config.yml claimed the uptime_kuma subdomain
"no longer resolves to anything". It resolves to 164.92.239.72 and answers HTTP
302, and 11 playbooks still template it. Same wrong premise as PLAN_3.
Verification: all 37 playbooks' --list-tasks output is byte-identical before and
after. A probe resolving all 22 values services_config.yml used to supply returns
21 identical and one intended deletion (caddy_sites_dir, now role-only - confirmed
the role still resolves it: "Ensure Caddy sites-enabled directory exists" comes
back ok against the real path). memos check-diff identical before and after.
Syntax passes on every playbook.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|---|---|---|
| .. | ||
| defaults | ||
| handlers | ||
| tasks | ||
| templates | ||
| README.md | ||
datum_gateway
Builds and runs DATUM Gateway, the
solo/pooled mining gateway, on knots-box. The calling playbook adds two more
plays on the edge host: the dashboard via caddy_site, and the public Stratum
port via socket_proxy.
Converted from deploy_datum_gateway_playbook.yml (802 lines) under Plan 6. The
playbook is now 68 lines and keeps all three plays.
⚠ This is half of a system
The Bitcoin Knots node on the same host feeds this gateway through
blocknotify=killall -USR1 datum_gateway in bitcoin.conf — see
roles/bitcoin_knots/README.md, where that line was found to be missing from the
template entirely. Changing either config means thinking about both.
Interrupting Stratum costs mining shares. Check before any run that restarts it:
ss -tn state established '( sport = :23334 )'
Two pieces of drift where the node was right
The repo and the node had diverged on values that matter, and the deployment would have applied the repo's:
| node (correct) | repo said | |
|---|---|---|
datum_mining_address |
bc1qvrj3g84… |
bc1qdse9dsg… |
pool_pass_workers / _full_users |
false |
true |
The address is the one that would have hurt: it is where block rewards are
paid, and unlike fulcrum and bitcoin-knots the Restart datum-gateway handler
here was never gated, so the change would have applied immediately rather than
sitting inert. Both corrected in the vault and defaults, with notes.
Verify semantics rather than text when touching config.json — render it and
compare parsed JSON, because the live file is single-line and the template is
pretty-printed, so a textual diff is all noise:
json.load(open('live.json')) == json.load(open('rendered.json'))
config.json holds real secrets — diff is suppressed
The file carries bitcoind.rpcpassword and api.admin_password. --diff
prints rendered content, so the task sets diff: false by default; pass
-e datum_reveal_config=true to opt in.
Note pool_pass_workers / pool_pass_full_users are booleans, not
passwords, despite the names — they control DATUM's pool-password passthrough.
mining.pool_address is a Bitcoin address and public by nature.
Expect changed on the compile every run
Configure cmake build and Compile datum_gateway are bare command: tasks
with no changed_when, so they always report changed and always re-run. The
build is reproducible — Install datum_gateway binary sees identical content and
does not replace it, so the installed binary keeps its original timestamp — but
the compile itself is wasted work on every run. That is the idempotent floor, not
drift.
Monitoring: one variable, no product knowledge
The check tests the gateway API and records the answer in its exit code, which
systemd keeps: systemctl is-failed datum-gateway-healthcheck.service. Set
healthcheck_push_url to report anywhere accepting an HTTP ping.
Unlike the other services here, only the health-check timer handler was gated
by uptime_kuma_enabled; the main deployment restart worked throughout.