personal_infra/ansible/roles/forgejo_runner/tasks/healthcheck.yml

74 lines
2.3 KiB
YAML
Raw Normal View History

forgejo-runner: convert to a role, de-Uptime-Kuma the health check 409-line playbook becomes a 16-line playbook plus a 318-line role with phases split across tasks/{prerequisites,install,configure,service,healthcheck}.yml and four templates. forgejo_runner_vars.yml is deleted; its content is the role's defaults. Applies the Plan 6 Stage 0 decision: keep whatever determines whether the service is healthy, drop the Uptime Kuma specifics, make the reporting point pluggable. Gone from the role: the embedded Python that created monitors over the Kuma API, the /tmp credentials file, token extraction, the systemd Environment= rewrite, and 8 `when: uptime_kuma_enabled` guards. What remains is the check itself, its log, the systemd unit and timer, and an honest exit code - `systemctl is-failed forgejo-runner-healthcheck.service` now answers the question with no monitoring system involved at all. Reporting is one variable, healthcheck_push_url, empty by default. Any endpoint that accepts an HTTP ping plugs in there. A pull-based monitor wants it left empty and reads unit state instead. PREMISE CORRECTION: Uptime Kuma is NOT dead. Plan 3 recorded "48 push timers curling an endpoint that no longer answers" and Plan 6 said the check had "nowhere to report to". Both wrong - 24+ push scripts across 11 hosts are pushing successfully right now (HTTP 200). Only the Ansible code and the vault credentials were decommissioned; the service never stopped. So the existing push URLs were harvested into a vaulted healthcheck_push_urls dict and are preserved, keeping this refactor behaviour-neutral. Retiring Kuma stays a deliberate act rather than a side effect. PLAN_3 and PLAN_6 are corrected. Verified: - task-list diff vs the old playbook shows ONLY the five Kuma tasks removed, everything else identical and in the same order - first run ok=22 changed=1 (the rewritten health script); both systemd units and forgejo-runner.service came back ok, so the templates reproduce the previous files byte-for-byte - second run ok=22 changed=0, fully idempotent - still reports "Ping sent successfully (HTTP 200)" from a script containing zero Uptime Kuma references - the 4 skipped tasks are genuine already-configured guards, checked not assumed Two things for the next service: - import_tasks, not include_tasks. Dynamic includes are opaque to --list-tasks, which is the primary verification tool here; the first attempt produced a useless diff. - `Assert runner is running` was guarded by uptime_kuma_enabled and so had not run since the decommissioning. It is not monitoring, it is the deployment checking its own work - the deprecation banner swept it up with the Kuma plumbing, and a runner that failed to start was deploying "successfully" in silence. Ungated now. The banner was applied to contiguous blocks, so read every uptime_kuma_enabled guard and ask whether it is monitoring or deployment. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-12 18:15:29 +02:00
---
# Everything here answers "is the service healthy" and records the answer.
# The Uptime Kuma specifics that used to surround it — an embedded Python script
# that created monitors over the API, a /tmp credentials file, token extraction,
# and a systemd Environment= rewrite — are gone. What reports where is now one
# variable, healthcheck_push_url. See the role README.
- name: Create healthcheck script directory
ansible.builtin.file:
path: "{{ healthcheck_script_dir }}"
state: directory
owner: root
group: root
mode: '0755'
- name: Create forgejo-runner healthcheck script
ansible.builtin.template:
src: healthcheck.sh.j2
dest: "{{ healthcheck_script_path }}"
owner: root
group: root
mode: '0755'
validate: "bash -n %s"
- name: Create healthcheck systemd service
ansible.builtin.template:
src: healthcheck.service.j2
dest: "/etc/systemd/system/{{ healthcheck_service_name }}.service"
owner: root
group: root
mode: '0644'
- name: Create healthcheck systemd timer
ansible.builtin.template:
src: healthcheck.timer.j2
dest: "/etc/systemd/system/{{ healthcheck_service_name }}.timer"
owner: root
group: root
mode: '0644'
- name: Reload systemd for healthcheck units
systemd:
daemon_reload: yes
- name: Enable and start healthcheck timer
systemd:
name: "{{ healthcheck_service_name }}.timer"
enabled: yes
state: started
- name: Test healthcheck script
command: "{{ healthcheck_script_path }}"
register: healthcheck_test
changed_when: false
- name: Verify healthcheck script works
assert:
that:
- healthcheck_test.rc == 0
fail_msg: "Healthcheck script failed to execute properly"
- name: Display deployment summary
debug:
msg: |
Forgejo Runner deployed successfully!
Runner Name: forgejo-runner-box
Instance: {{ forgejo_instance_url }}
Working Directory: {{ forgejo_runner_dir }}
Service: forgejo-runner.service ({{ runner_active.stdout }})
Healthcheck Monitor: {{ healthcheck_service_name }}
Healthcheck Interval: Every {{ healthcheck_interval_seconds }}s
Reporting to: {{ healthcheck_push_url | default('', true) | regex_replace('/api/push/.*', '/api/push/***') | default('(nowhere - set healthcheck_push_url)', true) }}