Nothing in the repo pushes to, authenticates against, or is gated by Uptime
Kuma any more.
── The sixth instance of the banner bug ────────────────────────────────────
memos had `Restart memos` guarded by `uptime_kuma_enabled`, because the
deprecation banner was placed immediately above it and swept it in. It is a
HANDLER, so every memos config change since 2026-09-11 applied to disk and
silently never restarted the service. Ungated.
That is the same failure found in forgejo-runner's self-assert, phoenixd's timer
enable, mempool's three timer enables, fulcrum's restart handler and bitcoind's
restart handler. Every guard was read and asked "monitoring or deployment?"
before being deleted, which is the only reason this was caught.
── What was removed ────────────────────────────────────────────────────────
30 uptime_kuma_enabled guards across 7 unconverted service playbooks, and
the 29 Kuma monitor-creation tasks they gated (embedded Python that drove
the Kuma API, temp credential files, cleanup)
7 dead uptime_kuma_api_url definitions
7 stale DEPRECATED banners
uptime_kuma_enabled and subdomains.uptime_kuma from group_vars/all
healthcheck_push_urls from the vault - 30 push tokens
services/ntfy/setup_ntfy_uptime_kuma_notification.yml -> archive/
The explanatory comments in the six converted roles are KEPT on purpose. They
record why a handler is ungated, and deleting the explanation invites someone
to helpfully re-add the guard.
── The probes moved rather than died ───────────────────────────────────────
Eight per-service health checks were still pushing to Kuma. They are not
superseded by infra/401: that answers "is the unit running", these answer "does
the service actually respond" - an RPC call to bitcoind, a TCP connect to
Fulcrum's Electrum port, an HTTP fetch from Mempool's backend. A process can be
perfectly `active` and useless.
So they were repointed, not deleted. Gatus external endpoints take a POST with
a bearer token and success=true|false where Kuma took a GET with ?status=up, so
report() now maps up/down to true/false internally and no call site changed.
Registered by infra/403 as the `probe` group, one token per host.
Two bugs fixed while in there:
* forgejo-runner's check only ever reported SUCCESS - it exited before pushing
when the runner was down, so a failure was invisible until the heartbeat
window expired. Reporting the failure is the entire point of a check.
* All six healthcheck .service units were mode 0644 and now carry a bearer
token. They are 0600.
Verified: 91 endpoints, 91 UP, 0 DOWN. Every probe triggered by hand and
confirmed arriving. Zero Kuma URLs left in the vault, zero live references in
any playbook or role.
Still standing, deliberately: the Kuma container on watchtower, its Caddy vhost,
and the uptime.contrapeso.xyz DNS record. Turning the service off is a separate
decision from removing the code that talked to it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|---|---|---|
| .. | ||
| defaults | ||
| handlers | ||
| tasks | ||
| templates | ||
| README.md | ||
bitcoin_knots
Builds Bitcoin Knots from source with PGP + SHA256 verification of the release
tarball, runs it as a full node on knots-box, and keeps a health check on a
systemd timer. The second play in the calling playbook publishes the P2P port
from the edge host via socket_proxy.
Converted from deploy_bitcoin_knots_playbook.yml (892 lines) under Plan 6. The
playbook is now 40 lines.
The build is guarded; the chain is never touched
build.yml is 32 tasks, every one carrying
when: not bitcoind_binary_exists.stat.exists. On a host that already has the
binary the whole download / verify / 30-60 minute compile skips — including the
two state: absent deletions, which target /opt/bitcoin-knots/source and the
extracted build directory.
The chain lives elsewhere and nothing here touches it:
bitcoin_knots_dir |
/opt/bitcoin-knots — build tree, safe to delete |
bitcoin_data_dir |
/var/lib/bitcoin — config, logs, wallets |
bitcoin_large_data_dir |
/mnt/knots_data — ~875 GB of blockchain |
The signature-verification tasks are the security control of this role. They are copied verbatim; do not "simplify" them.
⚠ This node is half of the mining setup
bitcoin.conf carries a DATUM Gateway section that was hand-added on the node
and was missing from the playbook's template:
blockmaxsize=3985000
blockmaxweight=3985000
blocknotify=killall -USR1 datum_gateway
maxmempool=1000
blockreconstructionextratxn=1000000
blocknotify is how datum_gateway learns a new block landed. Running the old
playbook would have deleted all of it, and solo mining would have carried on
grinding against a stale template — a silent failure that costs money rather
than raising an error. The template now carries it behind
bitcoin_datum_gateway_enabled.
bitcoin-knots and datum-gateway are one system, not two services. Changing either config means thinking about both.
The restart handler, and why exactness matters now
The hand-written Restart bitcoind handler carried
when: uptime_kuma_enabled | default(false), so it had been inert since the
decommissioning: bitcoin.conf and the systemd unit both notify it and neither
could restart anything. A config change applied to disk, reported success, and
never took effect.
It is ungated here — which raises the bar for the template. Any residual
difference between the template and the live file, down to a trailing newline,
means the task reports changed and restarts a Bitcoin node on every run. It
took four rounds of --check --diff to reach changed=0: the DATUM section, an
explanatory comment that was rendering into the deployed file (now a {# #}
Jinja comment), a # Pruning (optional) comment the live file had, and one
trailing blank line.
dbcache
Computed as 90% of RAM unless bitcoin_dbcache_mb_override is set. The live node
was hand-tuned to 200 MB; the calculation produces 3528. As with fulcrum,
set_fact outranks role defaults, so the calculation honours the override — a
value pinned only in defaults/ is silently ignored.
Monitoring: one variable, no product knowledge
The check tests bitcoind's RPC and records the answer in its exit code, which
systemd keeps: systemctl is-failed bitcoin-knots-healthcheck.service. Set
healthcheck_push_url to report anywhere that accepts an HTTP ping.
The timer had last fired 2026-08-09 while still reporting active and
enabled — the same OnBootSec + OnUnitActiveSec dead chain as fulcrum, where
nothing re-arms it if the service does not run in a given boot. The role runs the
check once after enabling, which both smoke-tests it and supplies the reference
the timer schedules from.