409-line playbook becomes a 16-line playbook plus a 318-line role with phases
split across tasks/{prerequisites,install,configure,service,healthcheck}.yml and
four templates. forgejo_runner_vars.yml is deleted; its content is the role's
defaults.
Applies the Plan 6 Stage 0 decision: keep whatever determines whether the
service is healthy, drop the Uptime Kuma specifics, make the reporting point
pluggable. Gone from the role: the embedded Python that created monitors over
the Kuma API, the /tmp credentials file, token extraction, the systemd
Environment= rewrite, and 8 `when: uptime_kuma_enabled` guards. What remains is
the check itself, its log, the systemd unit and timer, and an honest exit code -
`systemctl is-failed forgejo-runner-healthcheck.service` now answers the
question with no monitoring system involved at all.
Reporting is one variable, healthcheck_push_url, empty by default. Any endpoint
that accepts an HTTP ping plugs in there. A pull-based monitor wants it left
empty and reads unit state instead.
PREMISE CORRECTION: Uptime Kuma is NOT dead. Plan 3 recorded "48 push timers
curling an endpoint that no longer answers" and Plan 6 said the check had
"nowhere to report to". Both wrong - 24+ push scripts across 11 hosts are
pushing successfully right now (HTTP 200). Only the Ansible code and the vault
credentials were decommissioned; the service never stopped. So the existing push
URLs were harvested into a vaulted healthcheck_push_urls dict and are preserved,
keeping this refactor behaviour-neutral. Retiring Kuma stays a deliberate act
rather than a side effect. PLAN_3 and PLAN_6 are corrected.
Verified:
- task-list diff vs the old playbook shows ONLY the five Kuma tasks removed,
everything else identical and in the same order
- first run ok=22 changed=1 (the rewritten health script); both systemd units
and forgejo-runner.service came back ok, so the templates reproduce the
previous files byte-for-byte
- second run ok=22 changed=0, fully idempotent
- still reports "Ping sent successfully (HTTP 200)" from a script containing
zero Uptime Kuma references
- the 4 skipped tasks are genuine already-configured guards, checked not assumed
Two things for the next service:
- import_tasks, not include_tasks. Dynamic includes are opaque to --list-tasks,
which is the primary verification tool here; the first attempt produced a
useless diff.
- `Assert runner is running` was guarded by uptime_kuma_enabled and so had not
run since the decommissioning. It is not monitoring, it is the deployment
checking its own work - the deprecation banner swept it up with the Kuma
plumbing, and a runner that failed to start was deploying "successfully" in
silence. Ungated now. The banner was applied to contiguous blocks, so read
every uptime_kuma_enabled guard and ask whether it is monitoring or deployment.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
43 lines
1.5 KiB
Django/Jinja
43 lines
1.5 KiB
Django/Jinja
#!/bin/bash
|
|
# Forgejo Runner healthcheck — managed by Ansible (roles/forgejo_runner)
|
|
#
|
|
# Answers "is forgejo-runner healthy" and records it two ways: this log, and the
|
|
# exit code. The exit code is the durable artefact — systemd keeps it, so
|
|
# systemctl is-failed {{ healthcheck_service_name }}.service
|
|
# answers the question with no monitoring system involved.
|
|
#
|
|
# Reporting is optional and generic: if a push URL is configured it also pings
|
|
# it. Nothing here knows or cares which monitoring product is on the other end.
|
|
|
|
LOG_FILE="{{ healthcheck_log_file }}"
|
|
PUSH_URL="{{ healthcheck_push_url }}"
|
|
|
|
log_message() {
|
|
echo "$(date '+%Y-%m-%d %H:%M:%S') - $1" >> "$LOG_FILE"
|
|
}
|
|
|
|
main() {
|
|
if ! systemctl is-active --quiet forgejo-runner; then
|
|
log_message "ERROR: forgejo-runner is not active"
|
|
exit 1
|
|
fi
|
|
|
|
if [ -z "$PUSH_URL" ]; then
|
|
# Healthy, and nothing to report to. Not an error: the exit code below
|
|
# is still a complete answer for anything reading unit state.
|
|
log_message "forgejo-runner is active (no push URL configured)"
|
|
exit 0
|
|
fi
|
|
|
|
log_message "forgejo-runner is active, sending ping"
|
|
response=$(curl -s -w "\n%{http_code}" "$PUSH_URL?status=up&msg=forgejo-runner%20is%20active" 2>&1)
|
|
http_code=$(echo "$response" | tail -n1)
|
|
if [ "$http_code" = "200" ] || [ "$http_code" = "201" ]; then
|
|
log_message "Ping sent successfully (HTTP $http_code)"
|
|
else
|
|
log_message "ERROR: Failed to send ping (HTTP $http_code)"
|
|
exit 1
|
|
fi
|
|
}
|
|
|
|
main
|