forgejo-runner: convert to a role, de-Uptime-Kuma the health check

409-line playbook becomes a 16-line playbook plus a 318-line role with phases
split across tasks/{prerequisites,install,configure,service,healthcheck}.yml and
four templates. forgejo_runner_vars.yml is deleted; its content is the role's
defaults.

Applies the Plan 6 Stage 0 decision: keep whatever determines whether the
service is healthy, drop the Uptime Kuma specifics, make the reporting point
pluggable. Gone from the role: the embedded Python that created monitors over
the Kuma API, the /tmp credentials file, token extraction, the systemd
Environment= rewrite, and 8 `when: uptime_kuma_enabled` guards. What remains is
the check itself, its log, the systemd unit and timer, and an honest exit code -
`systemctl is-failed forgejo-runner-healthcheck.service` now answers the
question with no monitoring system involved at all.

Reporting is one variable, healthcheck_push_url, empty by default. Any endpoint
that accepts an HTTP ping plugs in there. A pull-based monitor wants it left
empty and reads unit state instead.

PREMISE CORRECTION: Uptime Kuma is NOT dead. Plan 3 recorded "48 push timers
curling an endpoint that no longer answers" and Plan 6 said the check had
"nowhere to report to". Both wrong - 24+ push scripts across 11 hosts are
pushing successfully right now (HTTP 200). Only the Ansible code and the vault
credentials were decommissioned; the service never stopped. So the existing push
URLs were harvested into a vaulted healthcheck_push_urls dict and are preserved,
keeping this refactor behaviour-neutral. Retiring Kuma stays a deliberate act
rather than a side effect. PLAN_3 and PLAN_6 are corrected.

Verified:
  - task-list diff vs the old playbook shows ONLY the five Kuma tasks removed,
    everything else identical and in the same order
  - first run ok=22 changed=1 (the rewritten health script); both systemd units
    and forgejo-runner.service came back ok, so the templates reproduce the
    previous files byte-for-byte
  - second run ok=22 changed=0, fully idempotent
  - still reports "Ping sent successfully (HTTP 200)" from a script containing
    zero Uptime Kuma references
  - the 4 skipped tasks are genuine already-configured guards, checked not assumed

Two things for the next service:

- import_tasks, not include_tasks. Dynamic includes are opaque to --list-tasks,
  which is the primary verification tool here; the first attempt produced a
  useless diff.
- `Assert runner is running` was guarded by uptime_kuma_enabled and so had not
  run since the decommissioning. It is not monitoring, it is the deployment
  checking its own work - the deprecation banner swept it up with the Kuma
  plumbing, and a runner that failed to start was deploying "successfully" in
  silence. Ungated now. The banner was applied to contiguous blocks, so read
  every uptime_kuma_enabled guard and ask whether it is monitoring or deployment.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
counterweight 2026-09-12 18:15:29 +02:00
parent 2ebb2f9a64
commit 73340d5fbe
Signed by: counterweight
GPG key ID: 883EDBAA726BD96C
16 changed files with 710 additions and 528 deletions

View file

@ -0,0 +1,59 @@
# `forgejo_runner`
Installs and runs a Forgejo Actions runner, registers it with the Forgejo
instance, and keeps a health check on a systemd timer.
Converted from `deploy_forgejo_runner_playbook.yml` (409 lines) under Plan 6.
The playbook is now 16 lines.
## Phases
`tasks/main.yml` imports five files in order:
| | |
|---|---|
| `prerequisites.yml` | Docker must be present |
| `install.yml` | binary, system user, working directory |
| `configure.yml` | config file, registration with the instance |
| `service.yml` | systemd unit, start, assert it came up |
| `healthcheck.yml` | check script, unit, timer |
`import_tasks`, not `include_tasks` — static imports are visible to
`--list-tasks`, which is how the conversion was verified against the playbook it
replaced.
## Monitoring: one variable, no product knowledge
This role contains **nothing specific to any monitoring system**. What used to
be here — an ~80-line embedded Python script creating monitors over the Uptime
Kuma API, a `/tmp` credentials file, token extraction, a systemd `Environment=`
rewrite, and 8 `when: uptime_kuma_enabled` guards — is gone.
What remains answers the actual question, *is this service healthy*, and records
it two ways:
- **the exit code**, which systemd keeps: `systemctl is-failed
forgejo-runner-healthcheck.service` is a complete answer with no monitoring
system involved at all;
- **a log file** at `{{ healthcheck_log_file }}`.
To report health somewhere, set one variable:
```yaml
healthcheck_push_url: "https://example/api/push/TOKEN"
```
Any endpoint accepting an HTTP ping works. Empty (the default) means check, log,
exit honestly, report nowhere — which is also the right setting for a *pull*-based
monitor like Prometheus' textfile collector, since that reads unit state instead.
The push URL is a credential (anyone holding it can forge an "up"), so callers
pass it from the vault rather than committing it.
## One behaviour change, deliberate
`Assert runner is running` used to be guarded by `uptime_kuma_enabled`, so it
never ran. It is not a monitoring task — it is the deployment checking its own
work — and the deprecation banner swept it up by mistake. It is ungated here,
which means a runner that fails to start now fails the play instead of
deploying "successfully" in silence.

View file

@ -0,0 +1,40 @@
---
# Binary
forgejo_runner_version: "6.3.1"
forgejo_runner_arch: "linux-amd64"
forgejo_runner_url: "https://code.forgejo.org/forgejo/runner/releases/download/v{{ forgejo_runner_version }}/forgejo-runner-{{ forgejo_runner_version }}-{{ forgejo_runner_arch }}"
forgejo_runner_bin_path: "/usr/local/bin/forgejo-runner"
# Runtime
forgejo_runner_user: "runner"
forgejo_runner_dir: "/opt/forgejo-runner"
forgejo_runner_config_path: "{{ forgejo_runner_dir }}/config.yml"
forgejo_runner_labels: "docker:docker://node:20-bookworm,ubuntu-latest:docker://node:20-bookworm,ubuntu-22.04:docker://node:20-bookworm,ubuntu-24.04:docker://node:20-bookworm"
# The Forgejo instance this runner registers with.
forgejo_instance_url: "https://forgejo.contrapeso.xyz"
# forgejo_runner_registration_token comes from the vault.
# --- Health check -----------------------------------------------------------
# The check answers "is this service healthy" and records the answer two ways:
# a log file, and its own exit code. The exit code is the durable artefact —
# systemd stores it, so `systemctl is-failed forgejo-runner-healthcheck.service`
# answers the question with no monitoring system involved at all.
healthcheck_interval_seconds: 60
healthcheck_timeout_seconds: 90
healthcheck_retries: 1
healthcheck_script_dir: /opt/forgejo-runner-healthcheck
healthcheck_script_path: "{{ healthcheck_script_dir }}/forgejo_runner_healthcheck.sh"
healthcheck_log_file: "{{ healthcheck_script_dir }}/forgejo_runner_healthcheck.log"
healthcheck_service_name: forgejo-runner-healthcheck
# WHERE TO REPORT HEALTH — the one place to plug in monitoring.
#
# Empty means "check, log, exit honestly, report nowhere". Set it to any URL
# that accepts an HTTP ping and the check will report there. Nothing in this
# role is specific to a particular monitoring product: the Uptime Kuma API
# calls, monitor creation and token handling that used to live here are gone.
#
# A pull-based monitor (Prometheus node_exporter textfile, say) needs this left
# empty — it reads the systemd unit state instead.
healthcheck_push_url: ""

View file

@ -0,0 +1,43 @@
---
- name: Check if config already exists
stat:
path: "{{ forgejo_runner_config_path }}"
register: config_stat
- name: Generate default config
shell: "{{ forgejo_runner_bin_path }} generate-config > {{ forgejo_runner_config_path }}"
args:
chdir: "{{ forgejo_runner_dir }}"
when: not config_stat.stat.exists
- name: Set config file ownership
file:
path: "{{ forgejo_runner_config_path }}"
owner: "{{ forgejo_runner_user }}"
group: "{{ forgejo_runner_user }}"
when: not config_stat.stat.exists
# ── 6. Register runner ─────────────────────────────────────────────
- name: Check if runner is already registered
stat:
path: "{{ forgejo_runner_dir }}/.runner"
register: runner_stat
- name: Register runner with Forgejo instance
command: >
{{ forgejo_runner_bin_path }} register --no-interactive
--instance {{ forgejo_instance_url }}
--token {{ forgejo_runner_registration_token }}
--name forgejo-runner-box
--labels "{{ forgejo_runner_labels }}"
args:
chdir: "{{ forgejo_runner_dir }}"
when: not runner_stat.stat.exists
- name: Set runner registration file ownership
file:
path: "{{ forgejo_runner_dir }}/.runner"
owner: "{{ forgejo_runner_user }}"
group: "{{ forgejo_runner_user }}"
when: not runner_stat.stat.exists

View file

@ -0,0 +1,73 @@
---
# Everything here answers "is the service healthy" and records the answer.
# The Uptime Kuma specifics that used to surround it — an embedded Python script
# that created monitors over the API, a /tmp credentials file, token extraction,
# and a systemd Environment= rewrite — are gone. What reports where is now one
# variable, healthcheck_push_url. See the role README.
- name: Create healthcheck script directory
ansible.builtin.file:
path: "{{ healthcheck_script_dir }}"
state: directory
owner: root
group: root
mode: '0755'
- name: Create forgejo-runner healthcheck script
ansible.builtin.template:
src: healthcheck.sh.j2
dest: "{{ healthcheck_script_path }}"
owner: root
group: root
mode: '0755'
validate: "bash -n %s"
- name: Create healthcheck systemd service
ansible.builtin.template:
src: healthcheck.service.j2
dest: "/etc/systemd/system/{{ healthcheck_service_name }}.service"
owner: root
group: root
mode: '0644'
- name: Create healthcheck systemd timer
ansible.builtin.template:
src: healthcheck.timer.j2
dest: "/etc/systemd/system/{{ healthcheck_service_name }}.timer"
owner: root
group: root
mode: '0644'
- name: Reload systemd for healthcheck units
systemd:
daemon_reload: yes
- name: Enable and start healthcheck timer
systemd:
name: "{{ healthcheck_service_name }}.timer"
enabled: yes
state: started
- name: Test healthcheck script
command: "{{ healthcheck_script_path }}"
register: healthcheck_test
changed_when: false
- name: Verify healthcheck script works
assert:
that:
- healthcheck_test.rc == 0
fail_msg: "Healthcheck script failed to execute properly"
- name: Display deployment summary
debug:
msg: |
Forgejo Runner deployed successfully!
Runner Name: forgejo-runner-box
Instance: {{ forgejo_instance_url }}
Working Directory: {{ forgejo_runner_dir }}
Service: forgejo-runner.service ({{ runner_active.stdout }})
Healthcheck Monitor: {{ healthcheck_service_name }}
Healthcheck Interval: Every {{ healthcheck_interval_seconds }}s
Reporting to: {{ healthcheck_push_url | default('', true) | regex_replace('/api/push/.*', '/api/push/***') | default('(nowhere - set healthcheck_push_url)', true) }}

View file

@ -0,0 +1,27 @@
---
- name: Download forgejo-runner binary
get_url:
url: "{{ forgejo_runner_url }}"
dest: "{{ forgejo_runner_bin_path }}"
mode: '0755'
# ── 3. Create runner system user ───────────────────────────────────
- name: Create runner system user
user:
name: "{{ forgejo_runner_user }}"
system: yes
shell: /usr/sbin/nologin
home: "{{ forgejo_runner_dir }}"
create_home: no
groups: docker
append: yes
comment: 'Forgejo Runner'
# ── 4. Create working directory ────────────────────────────────────
- name: Create forgejo-runner working directory
file:
path: "{{ forgejo_runner_dir }}"
state: directory
owner: "{{ forgejo_runner_user }}"
group: "{{ forgejo_runner_user }}"
mode: '0750'

View file

@ -0,0 +1,9 @@
---
# import_tasks, not include_tasks: these are unconditional phases, and static
# imports are visible to `--list-tasks`. That matters because the task list is
# how this refactor was verified against the playbook it replaced.
- ansible.builtin.import_tasks: prerequisites.yml
- ansible.builtin.import_tasks: install.yml
- ansible.builtin.import_tasks: configure.yml
- ansible.builtin.import_tasks: service.yml
- ansible.builtin.import_tasks: healthcheck.yml

View file

@ -0,0 +1,12 @@
---
- name: Check if Docker is installed
command: docker --version
register: docker_check
changed_when: false
failed_when: docker_check.rc != 0
- name: Fail if Docker is not available
assert:
that:
- docker_check.rc == 0
fail_msg: "Docker is required for forgejo-runner but is not installed"

View file

@ -0,0 +1,33 @@
---
- name: Create forgejo-runner systemd service
ansible.builtin.template:
src: forgejo-runner.service.j2
dest: /etc/systemd/system/forgejo-runner.service
owner: root
group: root
mode: '0644'
- name: Reload systemd
systemd:
daemon_reload: yes
- name: Enable and start forgejo-runner service
systemd:
name: forgejo-runner
enabled: yes
state: started
- name: Verify forgejo-runner is active
command: systemctl is-active forgejo-runner
register: runner_active
changed_when: false
# Ungated on purpose. This was previously guarded by `uptime_kuma_enabled`, but
# it is not a monitoring task — it is the deployment asserting its own success.
# The deprecation banner swept it up along with the Kuma plumbing, which meant a
# broken runner deployed "successfully" and silently.
- name: Assert runner is running
assert:
that:
- runner_active.stdout == "active"
fail_msg: "forgejo-runner service is not active: {{ runner_active.stdout }}"

View file

@ -0,0 +1,17 @@
[Unit]
Description=Forgejo Runner
Documentation=https://forgejo.org/docs/latest/admin/actions/
After=docker.service
Requires=docker.service
[Service]
Type=simple
User={{ forgejo_runner_user }}
Group={{ forgejo_runner_user }}
WorkingDirectory={{ forgejo_runner_dir }}
ExecStart={{ forgejo_runner_bin_path }} daemon --config {{ forgejo_runner_config_path }}
Restart=on-failure
RestartSec=10
[Install]
WantedBy=multi-user.target

View file

@ -0,0 +1,13 @@
[Unit]
Description=Forgejo Runner Healthcheck
After=network.target
[Service]
Type=oneshot
ExecStart={{ healthcheck_script_path }}
User=root
StandardOutput=journal
StandardError=journal
[Install]
WantedBy=multi-user.target

View file

@ -0,0 +1,43 @@
#!/bin/bash
# Forgejo Runner healthcheck — managed by Ansible (roles/forgejo_runner)
#
# Answers "is forgejo-runner healthy" and records it two ways: this log, and the
# exit code. The exit code is the durable artefact — systemd keeps it, so
# systemctl is-failed {{ healthcheck_service_name }}.service
# answers the question with no monitoring system involved.
#
# Reporting is optional and generic: if a push URL is configured it also pings
# it. Nothing here knows or cares which monitoring product is on the other end.
LOG_FILE="{{ healthcheck_log_file }}"
PUSH_URL="{{ healthcheck_push_url }}"
log_message() {
echo "$(date '+%Y-%m-%d %H:%M:%S') - $1" >> "$LOG_FILE"
}
main() {
if ! systemctl is-active --quiet forgejo-runner; then
log_message "ERROR: forgejo-runner is not active"
exit 1
fi
if [ -z "$PUSH_URL" ]; then
# Healthy, and nothing to report to. Not an error: the exit code below
# is still a complete answer for anything reading unit state.
log_message "forgejo-runner is active (no push URL configured)"
exit 0
fi
log_message "forgejo-runner is active, sending ping"
response=$(curl -s -w "\n%{http_code}" "$PUSH_URL?status=up&msg=forgejo-runner%20is%20active" 2>&1)
http_code=$(echo "$response" | tail -n1)
if [ "$http_code" = "200" ] || [ "$http_code" = "201" ]; then
log_message "Ping sent successfully (HTTP $http_code)"
else
log_message "ERROR: Failed to send ping (HTTP $http_code)"
exit 1
fi
}
main

View file

@ -0,0 +1,11 @@
[Unit]
Description=Run Forgejo Runner Healthcheck every minute
Requires={{ healthcheck_service_name }}.service
[Timer]
OnBootSec=30sec
OnUnitActiveSec={{ healthcheck_interval_seconds }}sec
Persistent=true
[Install]
WantedBy=timers.target