forgejo-runner: convert to a role, de-Uptime-Kuma the health check
409-line playbook becomes a 16-line playbook plus a 318-line role with phases
split across tasks/{prerequisites,install,configure,service,healthcheck}.yml and
four templates. forgejo_runner_vars.yml is deleted; its content is the role's
defaults.
Applies the Plan 6 Stage 0 decision: keep whatever determines whether the
service is healthy, drop the Uptime Kuma specifics, make the reporting point
pluggable. Gone from the role: the embedded Python that created monitors over
the Kuma API, the /tmp credentials file, token extraction, the systemd
Environment= rewrite, and 8 `when: uptime_kuma_enabled` guards. What remains is
the check itself, its log, the systemd unit and timer, and an honest exit code -
`systemctl is-failed forgejo-runner-healthcheck.service` now answers the
question with no monitoring system involved at all.
Reporting is one variable, healthcheck_push_url, empty by default. Any endpoint
that accepts an HTTP ping plugs in there. A pull-based monitor wants it left
empty and reads unit state instead.
PREMISE CORRECTION: Uptime Kuma is NOT dead. Plan 3 recorded "48 push timers
curling an endpoint that no longer answers" and Plan 6 said the check had
"nowhere to report to". Both wrong - 24+ push scripts across 11 hosts are
pushing successfully right now (HTTP 200). Only the Ansible code and the vault
credentials were decommissioned; the service never stopped. So the existing push
URLs were harvested into a vaulted healthcheck_push_urls dict and are preserved,
keeping this refactor behaviour-neutral. Retiring Kuma stays a deliberate act
rather than a side effect. PLAN_3 and PLAN_6 are corrected.
Verified:
- task-list diff vs the old playbook shows ONLY the five Kuma tasks removed,
everything else identical and in the same order
- first run ok=22 changed=1 (the rewritten health script); both systemd units
and forgejo-runner.service came back ok, so the templates reproduce the
previous files byte-for-byte
- second run ok=22 changed=0, fully idempotent
- still reports "Ping sent successfully (HTTP 200)" from a script containing
zero Uptime Kuma references
- the 4 skipped tasks are genuine already-configured guards, checked not assumed
Two things for the next service:
- import_tasks, not include_tasks. Dynamic includes are opaque to --list-tasks,
which is the primary verification tool here; the first attempt produced a
useless diff.
- `Assert runner is running` was guarded by uptime_kuma_enabled and so had not
run since the decommissioning. It is not monitoring, it is the deployment
checking its own work - the deprecation banner swept it up with the Kuma
plumbing, and a runner that failed to start was deploying "successfully" in
silence. Ungated now. The banner was applied to contiguous blocks, so read
every uptime_kuma_enabled guard and ask whether it is monitoring or deployment.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
parent
2ebb2f9a64
commit
73340d5fbe
16 changed files with 710 additions and 528 deletions
40
ansible/roles/forgejo_runner/defaults/main.yml
Normal file
40
ansible/roles/forgejo_runner/defaults/main.yml
Normal file
|
|
@ -0,0 +1,40 @@
|
|||
---
|
||||
# Binary
|
||||
forgejo_runner_version: "6.3.1"
|
||||
forgejo_runner_arch: "linux-amd64"
|
||||
forgejo_runner_url: "https://code.forgejo.org/forgejo/runner/releases/download/v{{ forgejo_runner_version }}/forgejo-runner-{{ forgejo_runner_version }}-{{ forgejo_runner_arch }}"
|
||||
forgejo_runner_bin_path: "/usr/local/bin/forgejo-runner"
|
||||
|
||||
# Runtime
|
||||
forgejo_runner_user: "runner"
|
||||
forgejo_runner_dir: "/opt/forgejo-runner"
|
||||
forgejo_runner_config_path: "{{ forgejo_runner_dir }}/config.yml"
|
||||
forgejo_runner_labels: "docker:docker://node:20-bookworm,ubuntu-latest:docker://node:20-bookworm,ubuntu-22.04:docker://node:20-bookworm,ubuntu-24.04:docker://node:20-bookworm"
|
||||
|
||||
# The Forgejo instance this runner registers with.
|
||||
forgejo_instance_url: "https://forgejo.contrapeso.xyz"
|
||||
# forgejo_runner_registration_token comes from the vault.
|
||||
|
||||
# --- Health check -----------------------------------------------------------
|
||||
# The check answers "is this service healthy" and records the answer two ways:
|
||||
# a log file, and its own exit code. The exit code is the durable artefact —
|
||||
# systemd stores it, so `systemctl is-failed forgejo-runner-healthcheck.service`
|
||||
# answers the question with no monitoring system involved at all.
|
||||
healthcheck_interval_seconds: 60
|
||||
healthcheck_timeout_seconds: 90
|
||||
healthcheck_retries: 1
|
||||
healthcheck_script_dir: /opt/forgejo-runner-healthcheck
|
||||
healthcheck_script_path: "{{ healthcheck_script_dir }}/forgejo_runner_healthcheck.sh"
|
||||
healthcheck_log_file: "{{ healthcheck_script_dir }}/forgejo_runner_healthcheck.log"
|
||||
healthcheck_service_name: forgejo-runner-healthcheck
|
||||
|
||||
# WHERE TO REPORT HEALTH — the one place to plug in monitoring.
|
||||
#
|
||||
# Empty means "check, log, exit honestly, report nowhere". Set it to any URL
|
||||
# that accepts an HTTP ping and the check will report there. Nothing in this
|
||||
# role is specific to a particular monitoring product: the Uptime Kuma API
|
||||
# calls, monitor creation and token handling that used to live here are gone.
|
||||
#
|
||||
# A pull-based monitor (Prometheus node_exporter textfile, say) needs this left
|
||||
# empty — it reads the systemd unit state instead.
|
||||
healthcheck_push_url: ""
|
||||
Loading…
Add table
Add a link
Reference in a new issue