60 lines
2.3 KiB
Markdown
60 lines
2.3 KiB
Markdown
|
|
# `forgejo_runner`
|
||
|
|
|
||
|
|
Installs and runs a Forgejo Actions runner, registers it with the Forgejo
|
||
|
|
instance, and keeps a health check on a systemd timer.
|
||
|
|
|
||
|
|
Converted from `deploy_forgejo_runner_playbook.yml` (409 lines) under Plan 6.
|
||
|
|
The playbook is now 16 lines.
|
||
|
|
|
||
|
|
## Phases
|
||
|
|
|
||
|
|
`tasks/main.yml` imports five files in order:
|
||
|
|
|
||
|
|
| | |
|
||
|
|
|---|---|
|
||
|
|
| `prerequisites.yml` | Docker must be present |
|
||
|
|
| `install.yml` | binary, system user, working directory |
|
||
|
|
| `configure.yml` | config file, registration with the instance |
|
||
|
|
| `service.yml` | systemd unit, start, assert it came up |
|
||
|
|
| `healthcheck.yml` | check script, unit, timer |
|
||
|
|
|
||
|
|
`import_tasks`, not `include_tasks` — static imports are visible to
|
||
|
|
`--list-tasks`, which is how the conversion was verified against the playbook it
|
||
|
|
replaced.
|
||
|
|
|
||
|
|
## Monitoring: one variable, no product knowledge
|
||
|
|
|
||
|
|
This role contains **nothing specific to any monitoring system**. What used to
|
||
|
|
be here — an ~80-line embedded Python script creating monitors over the Uptime
|
||
|
|
Kuma API, a `/tmp` credentials file, token extraction, a systemd `Environment=`
|
||
|
|
rewrite, and 8 `when: uptime_kuma_enabled` guards — is gone.
|
||
|
|
|
||
|
|
What remains answers the actual question, *is this service healthy*, and records
|
||
|
|
it two ways:
|
||
|
|
|
||
|
|
- **the exit code**, which systemd keeps: `systemctl is-failed
|
||
|
|
forgejo-runner-healthcheck.service` is a complete answer with no monitoring
|
||
|
|
system involved at all;
|
||
|
|
- **a log file** at `{{ healthcheck_log_file }}`.
|
||
|
|
|
||
|
|
To report health somewhere, set one variable:
|
||
|
|
|
||
|
|
```yaml
|
||
|
|
healthcheck_push_url: "https://example/api/push/TOKEN"
|
||
|
|
```
|
||
|
|
|
||
|
|
Any endpoint accepting an HTTP ping works. Empty (the default) means check, log,
|
||
|
|
exit honestly, report nowhere — which is also the right setting for a *pull*-based
|
||
|
|
monitor like Prometheus' textfile collector, since that reads unit state instead.
|
||
|
|
|
||
|
|
The push URL is a credential (anyone holding it can forge an "up"), so callers
|
||
|
|
pass it from the vault rather than committing it.
|
||
|
|
|
||
|
|
## One behaviour change, deliberate
|
||
|
|
|
||
|
|
`Assert runner is running` used to be guarded by `uptime_kuma_enabled`, so it
|
||
|
|
never ran. It is not a monitoring task — it is the deployment checking its own
|
||
|
|
work — and the deprecation banner swept it up by mistake. It is ungated here,
|
||
|
|
which means a runner that fails to start now fails the play instead of
|
||
|
|
deploying "successfully" in silence.
|