Commit graph

36 commits

Author SHA1 Message Date
efa9eb55ca
monitoring: systemd services, domain expiry, DNS correctness, public endpoints
Four more check types, 42 endpoints, taking the estate from 39 to 81.

── systemd services (infra/401) ────────────────────────────────────────────
Every unit we deploy, checked every 5 minutes with a 16-minute heartbeat.

This closes the gap that let a real bug run unnoticed earlier today: a backup
script left forgejo, lnbits, headscale and memos stopped, and NOTHING caught
it. The dumps exited 0, the artefacts were correct, the deploy said failed=0,
and liveness only proves the HOST is up - not that anything on it serves.

One endpoint PER UNIT but only ONE timer per host: the check iterates that
host's units and pushes a result for each, the way check-backups.sh reports per
source. Four units on vipy would otherwise mean four scripts, services and
timers. A single host-level red light would also say "something on vipy is
down" without saying which, which is not the question anyone has.

Keys are host-qualified because unit names collide - caddy runs on four
machines. Which units a host runs lives in host_vars/<host>/monitored_services,
because "what runs here" is a property of the machine, the same reasoning as
the cross-host ports.

nut-driver-enumerator is deliberately excluded: it is a oneshot generator that
is `enabled` but always `inactive`, so it would report down forever. Checked
live before excluding it.

── domain, DNS and public endpoints (infra/402) ────────────────────────────
The first checks that PULL rather than push, and that is the right way round:
all three are about how the outside world sees us, so they must be measured
from outside. Nothing is installed anywhere - no script, no timer, no token.
They also have no heartbeat, because a heartbeat answers "did the reporter
report"; when Gatus does the checking itself, failure is immediate.

  domain   1 endpoint,  24h, [DOMAIN_EXPIRATION] > 336h (two weeks)
  dns      11 endpoints, 24h, [DNS_RCODE] == NOERROR and [BODY] == the IP
  public   14 endpoints, 5m,  11 HTTPS + 3 TCP

Expected A records are derived from inventory (hostvars[host].ansible_host),
not written down again. The estate's recurring bug is an address recorded in a
second place and left behind when the machine moved; asserting against
inventory means a renumbered box is one edit, not two.

Expected HTTP status was checked live per site rather than assumed. Two return
401 - the Gatus dashboard and the DATUM dashboard, both behind basic auth - and
that is what is asserted: expecting 200 there would go green precisely when the
auth broke. headscale asserts /health rather than /, which is a 404 by design.
Every HTTPS check also carries [CERTIFICATE_EXPIRATION] > 168h, which is free
on an endpoint already being polled and catches a renewal that silently stops.

Two things learned the hard way, both now in comments:

  * A domain-expiry endpoint needs a URL SCHEME. Gatus derives the endpoint type
    from the prefix (endpoint.Type()), so a bare "contrapeso.xyz" is UNKNOWN and
    the whole config is rejected. It is https:// plus a DOMAIN_EXPIRATION
    condition and no status assertion, so the registrar's parking page at the
    apex is irrelevant. Upstream also enforces a 5m minimum interval for that
    placeholder, because it uses a free whois service.

  * That rejection proved skip-invalid-config-update was worth adding. Gatus
    logged "the configuration file was updated, but it is not valid, the old
    configuration will continue being used" and kept running. Without it the
    reload path calls panic() and one malformed contributed file takes the
    monitor down.

Verified: 15/15 service endpoints UP; 26/27 public-facing UP. The single DOWN is
public_memos, correctly - memos-box is powered off, and the condition result
reads [STATUS] (502) == 200. memos-box and arbret-staging-box were shut down
deliberately from the Proxmox UI to test the liveness endpoints; their endpoints
stay registered and will go green when the VMs return.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 09:33:57 +02:00
ede407ebe4
monitoring: recover host checks for the whole estate, reported to Gatus
Five checks, 27 endpoints, replacing what Uptime Kuma used to watch:

  is it up      every 5min, all hosts
  is disk full  daily,      all hosts
  is CPU hot    every 5min, nodito
  is ZFS broken daily,      nodito
  is UPS online every 5min, nodito

Two roles, kept separate so neither knows about the other - they meet at a URL
and a token, the same way caddy_site and each service meet at a vhost:

  roles/gatus_endpoint  runs on the observability host, writes ONE file into
                        /opt/gatus/config/endpoints/. Gatus merges every *.yaml
                        there and appends lists, so callers compose without
                        coordinating.
  roles/healthcheck     runs on the monitored host: a check script, a systemd
                        service, a timer, and an optional push. Ships a library
                        of check bodies under templates/checks/.

Everything PUSHES. Gatus never reaches out, which matters because nodito and its
VMs are behind NAT, and because four of the five checks are internal state with
no pollable surface at all. Liveness pushes too, deliberately: a heartbeat proves
the host is running AND can reach the internet, where an ICMP probe from one
vantage point only proves it answers pings from there. And since Gatus alerts
when a heartbeat window expires, a check that stops running raises the alarm by
itself - a dead timer looks exactly like a dead host, which is the correct
reading.

One bearer token per host, generated straight into the vault and never printed.
A token only writes results for its own host's endpoints, so a compromised host
can lie about itself, which it could do anyway.

Three things learned from the source that shaped this:

  * Gatus polls its own config every 30s and reloads
    (main.listenToConfigurationFileChanges), so gatus_endpoint needs no restart
    handler - writing the file IS the deploy.
  * ...but on a reload it panics if the new config fails to parse, unless
    skip-invalid-config-update is set. Endpoint files are contributed by other
    playbooks, so one malformed file would take the monitor down at the worst
    possible moment. Now set.
  * The push URL uses a key Gatus computes, not the name you write:
    sanitize(group) + "_" + sanitize(name), lowercased with / _ . , space # + &
    replaced by "-" (config/key/key.go). So knots_box_local is
    knots-box-local in the URL. The playbook derives it rather than hand-writing.

storage: maximum-number-of-results 900, up from upstream's 100. Gatus bounds the
database by COUNT and trims inline on insert, so there is no retention job and
no way to fill a disk - but history depth is then a function of check frequency,
and 100 results at a 5-minute interval is 8 hours. 900 is ~3 days of liveness
and ~2.5 years of the daily disk check. The uptime table is separate and its
30-day retention is hard-coded upstream.

A bug worth recording: the first deploy shipped five scripts that all died with
"syntax error: unexpected end of file". Jinja strips an included template's
trailing newline and trim_blocks then eats the newline after {% endif %}, so the
closing brace of check() landed on the same line as the body's last statement -
`return 0}`. Every check was broken and the deploy still reported failed=0,
because the role's "run once" task has failed_when: false and reports the result
as a debug message nobody read. The blank line that fixes it is now load-bearing
and commented as such.

Verified by triggering every unit by hand rather than waiting on timers: all
checks exit 0 on all hosts, and Gatus shows 26 UP / 1 DOWN. The one DOWN is
liveness_watchtower, which is honest - that host currently refuses SSH (TCP
connects, no banner exchange) and is excluded from this deploy. It is also the
box still running Uptime Kuma.

Known waste, not yet fixed: healthcheck installs its dependencies per CHECK
rather than per HOST, so apt runs 29 times estate-wide for a curl that is
already present, and daemon_reload runs 4x per host.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 08:54:10 +02:00
3711421af5
ansible: move the cross-host ports to host_vars, delete services_config.yml
The four ports were the only entries in services_config.yml with a real
justification: each is read twice, by the role that deploys the service on its
own box AND by a socket-proxy or Caddy play that runs on the EDGE host and
publishes it. A role default is invisible to that second play.

But the shape was wrong in two ways. The file had to be named in vars_files: by
30 plays - opt-in configuration that someone will eventually forget - and five
role defaults silently interpolated service_settings.*, so bitcoin_knots,
fulcrum, datum_gateway and mempool were not self-contained: using any of them
without that one vars_file entry broke it.

Each port now lives in host_vars/<owning box>/main.yml:

  host_vars/knots_box_local/main.yml    bitcoin_p2p_port, datum_gateway_api_port,
                                        datum_gateway_stratum_port
  host_vars/fulcrum_box_local/main.yml  fulcrum_ssl_port
  host_vars/mempool_box_local/main.yml  mempool_frontend_port

host_vars auto-loads and outranks role defaults, so the owning role picks the
value up with no vars_files at all, and the edge play reads the same single
definition as hostvars['<host>'].<name>. The role defaults keep the protocol
standard (8333, 50002, ...) so each role still works standalone, with the live
deployment's value in host_vars winning.

Also fixed a fourth copy of an inventory identity: the mempool Caddy play had
"mempool-box:{{ ... }}" hardcoded in the upstream. It now derives the host from
hostvars['mempool_box_local'].ansible_host, so inventory is the only place any
box's name is written down.

services_config.yml is deleted, with 25 more vars_files entries across 19
playbooks. Between this and the previous commit, 87 vars_files entries are gone
and every variable in the repo now comes from group_vars/all, host_vars,
inventory, a role default, or that service's own *_vars.yml.

Verification: an edge-host probe resolves all eight ports and hostnames to
byte-identical values to the ones services_config.yml used to supply. Each
owning host resolves its own port through host_vars. All 37 playbooks'
--list-tasks output is unchanged. The four edge plays that consume these values
all check-diff changed=0 - the socket-proxy and Caddy units on vipy are
byte-identical, which is the direct proof the rewiring landed on the same
values. fulcrum and datum-gateway check-diff exactly as before (ok=28/changed=1
and ok=15/changed=1, both the known timer re-arm).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 21:02:57 +02:00
954b683c71
ansible: delete the duplicated vars files, move globals to group_vars/all
Three files existed only as second copies of things group_vars/all already
auto-loads, and 34 playbooks named them in vars_files: - which outranks
group_vars, so the copies won. The day someone edited one and not the other,
those plays would silently keep the stale value. infra_vars.yml was already
drifting: group_vars/all/main.yml had grown age_backup_recipient and
backup_pull_public_key that it lacked.

  infra_vars.yml          - a strict subset of group_vars/all/main.yml
  infra_secrets.yml       - decrypts byte-identical to group_vars/all/vault.yml
  infra_secrets.yml.example - documented Uptime Kuma credentials as the reason
                              the file exists, which stopped being true

Deleted, along with 62 vars_files entries across 34 playbooks (12 of which
named ../../group_vars/all/main.yml directly - same defect, a vars_files entry
duplicating an auto-loaded file at higher precedence than the file itself).

Checked before touching anything: infra_secrets.yml was listed LAST in 10 plays,
after services_config.yml, so removal would flip precedence if the two shared a
key. They share none, and neither does services_config.yml with
group_vars/all/main.yml, so the removal is provably inert.

services_config.yml was the last one standing. It held four unrelated things:

  caddy_sites_dir            - an identical copy of roles/caddy_site/defaults/.
                               Deleted; the role default is now the only one.
  *.tailscale_hostname (x3)  - a THIRD copy of each box's identity, which
                               inventory.ini already holds as ansible_host.
                               Deleted. Edge plays now read
                               hostvars['<host>'].ansible_host - verified an
                               edge play resolves that with nothing loaded and
                               the other host in no play. Three copies of one
                               name is how bitcoin_rpc_host ended up labelled
                               "knots_box" while pointing at fulcrum-box.
  subdomains, ntfy topic,    - genuinely global: their readers span managed,
  headscale namespace          monitoring, vpn_control and edge, so no single
                               group covers them. Moved to group_vars/all/main.yml
                               where they auto-load. The ntfy_topic and
                               headscale_namespace indirection through
                               service_settings collapses to the global name.
  the four cross-host ports  - the only entries with a real justification.
                               Left in place; they move in the next commit.

Also dead, all Uptime Kuma residue or duplication:
  phoenixd_monitor_name, forgejo_runner healthcheck_timeout_seconds/retries,
  fulcrum_tailscale_hostname, and bitcoin_knots_version - the last being a
  v-prefixed copy of bitcoin_knots_version_short that nothing read, two
  hand-maintained copies of one version string.

Corrected a false comment: services_config.yml claimed the uptime_kuma subdomain
"no longer resolves to anything". It resolves to 164.92.239.72 and answers HTTP
302, and 11 playbooks still template it. Same wrong premise as PLAN_3.

Verification: all 37 playbooks' --list-tasks output is byte-identical before and
after. A probe resolving all 22 values services_config.yml used to supply returns
21 identical and one intended deletion (caddy_sites_dir, now role-only - confirmed
the role still resolves it: "Ensure Caddy sites-enabled directory exists" comes
back ok against the real path). memos check-diff identical before and after.
Syntax passes on every playbook.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 20:58:46 +02:00
0f03c503c8
nodito: de-Uptime-Kuma the ZFS and NUT playbooks, extract templates
These two are host-specific by design - nodito is a pet, not cattle - so they
stay playbooks rather than becoming roles. But they had rotted.

De-Kuma, following the pattern of the six service roles:

  Both plays opened with an assert on uptime_kuma_username/password, which were
  removed from the vault, so both failed before doing anything. Dropped that,
  the two embedded Python monitor-creation scripts, and their /tmp cleanup.
  Kept every check, threshold and systemd timer - those are the durable part.

  Reporting is now generic: `healthcheck_push_url` goes into the unit as
  Environment=HEALTHCHECK_PUSH_URL and the script reads ${HEALTHCHECK_PUSH_URL:-},
  treating empty as normal rather than an error. The exit code is the real
  answer; systemd keeps it. Both scripts now also report status=down on failure
  instead of only going silent. Live push URLs harvested into the vault so
  nothing observable changes for ZFS.

Three live bugs found while check-diffing:

  1. 32_zfs would have DE-REGISTERED the Proxmox storage. `pvesm remove` was
     gated on the storage existing and `pvesm add` on it NOT existing - mutually
     exclusive - so a real run removed the proxmox-tank-1 entry backing every VM
     and never put it back. It would also have dropped `mountpoint /var/lib/vz`,
     which the live entry has and `pvesm add` does not set. Registration is now
     add-only.

  2. zfs_disk_1 named ata-...WX11TN0Z, a disk no longer in the machine. The live
     mirror is WX120LHQ + WX11TN2P; a leg was replaced and the repo never caught
     up. Inert behind `when: zfs_pool_exists.rc != 0`, but wrong on any
     disaster-recovery run. This is the seventh instance of an identifier
     written down once whose hardware later moved.

  3. 34_nut has NEVER been applied to nodito - no /etc/nut file carries the
     "Managed by Ansible" marker; they were written by hand in January 2026. The
     vault held the literal CHANGE_ME_TO_SECURE_PASSWORD, so applying it would
     have overwritten a working upsd/upsmon auth pair with a placeholder and
     restarted NUT, leaving the hypervisor's UPS unable to trigger a clean
     shutdown on mains loss. The Kuma assert was the only thing stopping that,
     so removing it without a replacement would have armed the gun: there is now
     an explicit assert that refuses to run on the placeholder. The real
     password is in the vault and `Configure upsd users` check-diffs clean.

  Templates reconciled with the live files first, so applying 34_nut is close to
  a no-op: added `maxretry = 3` to ups.conf and OFFDURATION / RBWARNTIME /
  NOCOMMWARNTIME / FINALDELAY plus quoted POWERDOWNFLAG to upsmon.conf. The only
  substantive additions left are the NOTIFYMSG/NOTIFYFLAG syslog lines.

  /usr/local/bin/ups-heartbeat.sh on the box is an orphan - mode 0644, not
  executable, referenced by no unit and no cron entry - but its push token
  belongs to a monitor that still exists and answers, so that monitor has had no
  heartbeat since January. Harvested as healthcheck_push_urls.ups; applying this
  play is what will finally feed it.

Finish the host_vars migration:

  infra/nodito/nodito_vars.yml was byte-identical to host_vars/nodito/main.yml.
  Deleted it, moved nodito_secrets.yml to host_vars/nodito/vault.yml, and
  stripped the dead vars_files entries from all three playbooks.

Extract the 13 inline `content: |` blocks to infra/nodito/templates/, pulled via
a YAML load rather than retyped. 681+582 lines become 313+344 plus templates.
Ownership parity against HEAD checked mechanically: no owner/group/mode drift on
any surviving task.

Neither playbook has been applied. ZFS play 2 check-runs failed=0 with two
changes: the script rewrite (the deployed one is a hand-edited DEBUG VERSION
with `set -x`) and the Environment line.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 19:07:50 +02:00
b5b028c356
uptime-kuma: deprecation banners on the monitoring-only plays
These five plus the ntfy notification playbook assert on the credentials, so
they now fail immediately instead of running — deliberately, before anything is
installed. The banner says so and points at archive/uptime_kuma/.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-11 22:43:55 +02:00
80c9e6f3e3
uptime-kuma: remove credentials from the vaults
Drops uptime_kuma_username/password from infra_secrets.yml and its group_vars
copy and example, plus a dead push token in nodito_secrets.yml that nothing
referenced. They remain in git history — rotation is what actually retires them.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-11 22:43:15 +02:00
ceb5c4ad27
infra/nodito: target hypervisor group instead of nodito_host/nodito
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-11 21:56:24 +02:00
28a0bb2806
new groups, stop using all 2026-09-11 21:51:50 +02:00
6029317f3b
ansible: vault encrypt secrets 2026-09-11 17:53:12 +02:00
0736f29f79
little thingies 2026-03-22 21:25:08 +01:00
c6e1a01167
thingies 2026-02-08 18:22:31 +01:00
08281ce349
ups playbook 2026-01-11 22:43:27 +01:00
fe321050c1
monitor zfs 2026-01-04 23:19:19 +01:00
8863f800bf
bitcoin node stuff 2025-12-14 18:52:36 +01:00
2893bb77cd
improve tailscale add 2025-12-13 18:54:46 +01:00
0b578ee738
stuff 2025-12-08 10:34:04 +01:00
c14d61d090
stuff 2025-12-07 19:02:50 +01:00
6a43132bc8
too much stuff 2025-12-01 11:16:47 +01:00
fbbeb59c0e
stuff 2025-11-14 23:36:00 +01:00
c8754e1bdc
lots of stuff man 2025-11-06 23:09:44 +01:00
0bafa6ba2c
stuff 2025-11-03 16:51:38 +01:00
9d43c19189
hostname works 2025-11-02 01:48:07 +01:00
d4782d00cc
fix cloud init template 2025-11-02 01:24:39 +01:00
6f42e43efb
qemu works 2025-10-31 00:17:42 +01:00
102fad268c
qemu agent for vm template 2025-10-30 23:11:14 +01:00
e03c21e853
debian template 2025-10-30 11:21:48 +01:00
0c34e25502
packages, zfs pool 2025-10-29 00:13:15 +01:00
4a4c61308a
temp monitor 2025-10-26 23:39:02 +01:00
85012f8ba5
first steps with proxmox 2025-10-26 22:33:01 +01:00
13537aa984
Separate watchtower from vipy 2025-07-21 09:39:36 +02:00
8766af831c
a few things 2025-07-09 00:32:51 +02:00
Pablo Martin
3d3d65575b lots of stuff 2025-07-03 17:21:31 +02:00
Pablo Martin
dac4a98f79 uptime kuma backups work 2025-07-02 17:17:56 +02:00
Pablo Martin
eddde5e53a uptime kuma works 2025-07-01 17:02:28 +02:00
Pablo Martin
3343de2dc0 thingies 2025-07-01 16:14:44 +02:00