personal_infra/ansible/roles/forgejo_runner/README.md
counterweight 73340d5fbe
forgejo-runner: convert to a role, de-Uptime-Kuma the health check
409-line playbook becomes a 16-line playbook plus a 318-line role with phases
split across tasks/{prerequisites,install,configure,service,healthcheck}.yml and
four templates. forgejo_runner_vars.yml is deleted; its content is the role's
defaults.

Applies the Plan 6 Stage 0 decision: keep whatever determines whether the
service is healthy, drop the Uptime Kuma specifics, make the reporting point
pluggable. Gone from the role: the embedded Python that created monitors over
the Kuma API, the /tmp credentials file, token extraction, the systemd
Environment= rewrite, and 8 `when: uptime_kuma_enabled` guards. What remains is
the check itself, its log, the systemd unit and timer, and an honest exit code -
`systemctl is-failed forgejo-runner-healthcheck.service` now answers the
question with no monitoring system involved at all.

Reporting is one variable, healthcheck_push_url, empty by default. Any endpoint
that accepts an HTTP ping plugs in there. A pull-based monitor wants it left
empty and reads unit state instead.

PREMISE CORRECTION: Uptime Kuma is NOT dead. Plan 3 recorded "48 push timers
curling an endpoint that no longer answers" and Plan 6 said the check had
"nowhere to report to". Both wrong - 24+ push scripts across 11 hosts are
pushing successfully right now (HTTP 200). Only the Ansible code and the vault
credentials were decommissioned; the service never stopped. So the existing push
URLs were harvested into a vaulted healthcheck_push_urls dict and are preserved,
keeping this refactor behaviour-neutral. Retiring Kuma stays a deliberate act
rather than a side effect. PLAN_3 and PLAN_6 are corrected.

Verified:
  - task-list diff vs the old playbook shows ONLY the five Kuma tasks removed,
    everything else identical and in the same order
  - first run ok=22 changed=1 (the rewritten health script); both systemd units
    and forgejo-runner.service came back ok, so the templates reproduce the
    previous files byte-for-byte
  - second run ok=22 changed=0, fully idempotent
  - still reports "Ping sent successfully (HTTP 200)" from a script containing
    zero Uptime Kuma references
  - the 4 skipped tasks are genuine already-configured guards, checked not assumed

Two things for the next service:

- import_tasks, not include_tasks. Dynamic includes are opaque to --list-tasks,
  which is the primary verification tool here; the first attempt produced a
  useless diff.
- `Assert runner is running` was guarded by uptime_kuma_enabled and so had not
  run since the decommissioning. It is not monitoring, it is the deployment
  checking its own work - the deprecation banner swept it up with the Kuma
  plumbing, and a runner that failed to start was deploying "successfully" in
  silence. Ungated now. The banner was applied to contiguous blocks, so read
  every uptime_kuma_enabled guard and ask whether it is monitoring or deployment.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-12 18:15:29 +02:00

59 lines
2.3 KiB
Markdown

# `forgejo_runner`
Installs and runs a Forgejo Actions runner, registers it with the Forgejo
instance, and keeps a health check on a systemd timer.
Converted from `deploy_forgejo_runner_playbook.yml` (409 lines) under Plan 6.
The playbook is now 16 lines.
## Phases
`tasks/main.yml` imports five files in order:
| | |
|---|---|
| `prerequisites.yml` | Docker must be present |
| `install.yml` | binary, system user, working directory |
| `configure.yml` | config file, registration with the instance |
| `service.yml` | systemd unit, start, assert it came up |
| `healthcheck.yml` | check script, unit, timer |
`import_tasks`, not `include_tasks` — static imports are visible to
`--list-tasks`, which is how the conversion was verified against the playbook it
replaced.
## Monitoring: one variable, no product knowledge
This role contains **nothing specific to any monitoring system**. What used to
be here — an ~80-line embedded Python script creating monitors over the Uptime
Kuma API, a `/tmp` credentials file, token extraction, a systemd `Environment=`
rewrite, and 8 `when: uptime_kuma_enabled` guards — is gone.
What remains answers the actual question, *is this service healthy*, and records
it two ways:
- **the exit code**, which systemd keeps: `systemctl is-failed
forgejo-runner-healthcheck.service` is a complete answer with no monitoring
system involved at all;
- **a log file** at `{{ healthcheck_log_file }}`.
To report health somewhere, set one variable:
```yaml
healthcheck_push_url: "https://example/api/push/TOKEN"
```
Any endpoint accepting an HTTP ping works. Empty (the default) means check, log,
exit honestly, report nowhere — which is also the right setting for a *pull*-based
monitor like Prometheus' textfile collector, since that reads unit state instead.
The push URL is a credential (anyone holding it can forge an "up"), so callers
pass it from the vault rather than committing it.
## One behaviour change, deliberate
`Assert runner is running` used to be guarded by `uptime_kuma_enabled`, so it
never ran. It is not a monitoring task — it is the deployment checking its own
work — and the deprecation banner swept it up by mistake. It is ungated here,
which means a runner that fails to start now fails the play instead of
deploying "successfully" in silence.