archive: record Uptime Kuma monitors and setup before decommissioning
Captured from the live instance rather than the repo: the playbooks created 17 monitors, the server had 75. The rest existed only in the UI. Push tokens are excluded deliberately — they are live credentials. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
parent
b0f14ea365
commit
79942525b1
6 changed files with 1405 additions and 0 deletions
152
archive/uptime_kuma/MONITORS.md
Normal file
152
archive/uptime_kuma/MONITORS.md
Normal file
|
|
@ -0,0 +1,152 @@
|
|||
# Uptime Kuma — monitor inventory (archived)
|
||||
|
||||
Captured from the live instance at `https://uptime.contrapeso.xyz` on 2026-09-11,
|
||||
immediately before decommissioning. This is the **authoritative** record: most of
|
||||
these monitors existed only in the Uptime Kuma UI and were never described by any
|
||||
playbook in this repo.
|
||||
|
||||
**75 monitors total** — 16 group, 8 http, 3 port, 48 push. All were active.
|
||||
|
||||
Push tokens are deliberately **not** recorded here: they are live credentials, and
|
||||
anything holding one could report a false 'up'. They die with the server.
|
||||
|
||||
---
|
||||
|
||||
## arbret - production *(7 monitors)*
|
||||
|
||||
| Monitor | Type | Target | Interval | Notes |
|
||||
|---|---|---|---|---|
|
||||
| arbret.com - arbret-analytics | push | — | 120s | Healthy when timer is scheduled and last run succeeded |
|
||||
| arbret.com - arbret-backup | push | — | 120s | Healthy when timer is scheduled and last run succeeded |
|
||||
| arbret.com - arbret-server | push | — | 120s | Healthy when arbret-server.service is active |
|
||||
| arbret.com - arbret-worker | push | — | 120s | Healthy when arbret-worker.service is active |
|
||||
| arbret.com - health | push | — | 120s | Healthy when GET /api/health returns status ok |
|
||||
| arbret.com - https | push | — | 120s | Healthy when HTTPS front-door returns 200 |
|
||||
| arbret.com - postgresql | push | — | 120s | Healthy when postgresql.service is active |
|
||||
|
||||
## arbret - staging *(7 monitors)*
|
||||
|
||||
| Monitor | Type | Target | Interval | Notes |
|
||||
|---|---|---|---|---|
|
||||
| arbretstaging.contrapeso.xyz - arbret-analytics | push | — | 120s | Healthy when timer is scheduled and last run succeeded |
|
||||
| arbretstaging.contrapeso.xyz - arbret-backup | push | — | 120s | Healthy when timer is scheduled and last run succeeded |
|
||||
| arbretstaging.contrapeso.xyz - arbret-server | push | — | 120s | Healthy when arbret-server.service is active |
|
||||
| arbretstaging.contrapeso.xyz - arbret-worker | push | — | 120s | Healthy when arbret-worker.service is active |
|
||||
| arbretstaging.contrapeso.xyz - health | push | — | 120s | Healthy when GET /api/health returns status ok |
|
||||
| arbretstaging.contrapeso.xyz - https | push | — | 120s | Healthy when HTTPS front-door returns 200 |
|
||||
| arbretstaging.contrapeso.xyz - postgresql | push | — | 120s | Healthy when postgresql.service is active |
|
||||
|
||||
## arbret-staging-box - infra *(2 monitors)*
|
||||
|
||||
| Monitor | Type | Target | Interval | Notes |
|
||||
|---|---|---|---|---|
|
||||
| disk-usage-arbret-staging-box-root | push | — | 960s | upside-down, Disk Usage: arbret-staging-box (/) - Alerts when usage excee |
|
||||
| system-healthcheck-arbret-staging-box | push | — | 90s | System Healthcheck: arbret-staging-box - Regular healthcheck |
|
||||
|
||||
## forgejo-runner-box - infra *(2 monitors)*
|
||||
|
||||
| Monitor | Type | Target | Interval | Notes |
|
||||
|---|---|---|---|---|
|
||||
| disk-usage-forgejo-runner-box-root | push | — | 960s | upside-down, Disk Usage: forgejo-runner-box (/) - Alerts when usage excee |
|
||||
| system-healthcheck-forgejo-runner-box | push | — | 90s | System Healthcheck: forgejo-runner-box - Regular healthcheck |
|
||||
|
||||
## fulcrum-box - infra *(2 monitors)*
|
||||
|
||||
| Monitor | Type | Target | Interval | Notes |
|
||||
|---|---|---|---|---|
|
||||
| disk-usage-fulcrum-box-root | push | — | 960s | upside-down, Disk Usage: fulcrum-box (/) - Alerts when usage exceeds 80% |
|
||||
| system-healthcheck-fulcrum-box | push | — | 90s | System Healthcheck: fulcrum-box - Regular healthcheck ping e |
|
||||
|
||||
## knots-box - infra *(2 monitors)*
|
||||
|
||||
| Monitor | Type | Target | Interval | Notes |
|
||||
|---|---|---|---|---|
|
||||
| disk-usage-knots-box-root | push | — | 960s | upside-down, Disk Usage: knots-box (/) - Alerts when usage exceeds 80% |
|
||||
| system-healthcheck-knots-box | push | — | 90s | System Healthcheck: knots-box - Regular healthcheck ping eve |
|
||||
|
||||
## memos-box - infra *(2 monitors)*
|
||||
|
||||
| Monitor | Type | Target | Interval | Notes |
|
||||
|---|---|---|---|---|
|
||||
| disk-usage-memos-box-root | push | — | 960s | upside-down, Disk Usage: memos-box (/) - Alerts when usage exceeds 80% |
|
||||
| system-healthcheck-memos-box | push | — | 90s | System Healthcheck: memos-box - Regular healthcheck ping eve |
|
||||
|
||||
## mempool-box - infra *(2 monitors)*
|
||||
|
||||
| Monitor | Type | Target | Interval | Notes |
|
||||
|---|---|---|---|---|
|
||||
| disk-usage-mempool-box-root | push | — | 960s | upside-down, Disk Usage: mempool-box (/) - Alerts when usage exceeds 80% |
|
||||
| system-healthcheck-mempool-box | push | — | 90s | System Healthcheck: mempool-box - Regular healthcheck ping e |
|
||||
|
||||
## nodito - infra *(4 monitors)*
|
||||
|
||||
| Monitor | Type | Target | Interval | Notes |
|
||||
|---|---|---|---|---|
|
||||
| UPS ONLINE | push | https:// | 90s | — |
|
||||
| cpu-temp-nodito | push | — | 120s | upside-down, CPU Temperature: nodito - Alerts when temperature exceeds 80 |
|
||||
| system-healthcheck-nodito | push | — | 300s | System Healthcheck: nodito - Regular healthcheck ping every |
|
||||
| zfs-health-nodito | push | — | 90000s | ZFS Pool Health: nodito - Daily health check for pool proxmo |
|
||||
|
||||
## nonkeiwaisi-box - infra *(0 monitors)*
|
||||
|
||||
_(empty)_
|
||||
|
||||
## prd-arbret - infra *(2 monitors)*
|
||||
|
||||
| Monitor | Type | Target | Interval | Notes |
|
||||
|---|---|---|---|---|
|
||||
| disk-usage-prd-arbret-root | push | — | 960s | upside-down, Disk Usage: prd-arbret (/) - Alerts when usage exceeds 80% |
|
||||
| system-healthcheck-prd-arbret | push | — | 90s | System Healthcheck: prd-arbret - Regular healthcheck ping ev |
|
||||
|
||||
## prd-spacey - infra *(2 monitors)*
|
||||
|
||||
| Monitor | Type | Target | Interval | Notes |
|
||||
|---|---|---|---|---|
|
||||
| disk-usage-prd-spacey-root | push | — | 960s | upside-down, Disk Usage: prd-spacey (/) - Alerts when usage exceeds 80% |
|
||||
| system-healthcheck-prd-spacey | push | — | 90s | System Healthcheck: prd-spacey - Regular healthcheck ping ev |
|
||||
|
||||
## prd-vipy - infra *(2 monitors)*
|
||||
|
||||
| Monitor | Type | Target | Interval | Notes |
|
||||
|---|---|---|---|---|
|
||||
| disk-usage-prd-vipy-root | push | — | 960s | upside-down, Disk Usage: prd-vipy (/) - Alerts when usage exceeds 80% |
|
||||
| system-healthcheck-prd-vipy | push | — | 90s | System Healthcheck: prd-vipy - Regular healthcheck ping ever |
|
||||
|
||||
## prd-watchtower - infra *(2 monitors)*
|
||||
|
||||
| Monitor | Type | Target | Interval | Notes |
|
||||
|---|---|---|---|---|
|
||||
| disk-usage-prd-watchtower-root | push | — | 960s | upside-down, Disk Usage: prd-watchtower (/) - Alerts when usage exceeds 8 |
|
||||
| system-healthcheck-prd-watchtower | push | — | 90s | System Healthcheck: prd-watchtower - Regular healthcheck pin |
|
||||
|
||||
## services *(19 monitors)*
|
||||
|
||||
| Monitor | Type | Target | Interval | Notes |
|
||||
|---|---|---|---|---|
|
||||
| Forgejo | http | https://forgejo.contrapeso.xyz/api/healthz | 90s | — |
|
||||
| Headscale | http | https://headscale.contrapeso.xyz/health | 60s | — |
|
||||
| LNBits | http | https://wallet.contrapeso.xyz/api/v1/health | 60s | — |
|
||||
| Memos | http | https://memos.contrapeso.xyz/healthz | 60s | — |
|
||||
| Mempool Public Access | http | https://mempool.contrapeso.xyz | 60s | — |
|
||||
| Personal Blog | http | https://pablohere.contrapeso.xyz | 60s | — |
|
||||
| Vaultwarden | http | https://vault.contrapeso.xyz/alive | 60s | — |
|
||||
| ntfy-emergency-app | http | https://avisame.contrapeso.xyz | 60s | — |
|
||||
| Bitcoin Knots P2P Public | port | 167.172.107.33:8333 | 60s | — |
|
||||
| DATUM Stratum (public) | port | 167.172.107.33:23334 | 60s | — |
|
||||
| Fulcrum SSL Public | port | 167.172.107.33:50002 | 60s | — |
|
||||
| Bitcoin Knots | push | — | 90s | — |
|
||||
| DATUM Gateway | push | — | 90s | — |
|
||||
| Fulcrum | push | — | 90s | — |
|
||||
| Mempool Backend | push | — | 180s | — |
|
||||
| Mempool Frontend | push | — | 90s | — |
|
||||
| Mempool MariaDB | push | — | 90s | — |
|
||||
| Phoenixd | push | — | 90s | — |
|
||||
| forgejo-runner-healthcheck | push | — | 90s | Forgejo Runner healthcheck - ping every 60s |
|
||||
|
||||
## small-backups-box - infra *(2 monitors)*
|
||||
|
||||
| Monitor | Type | Target | Interval | Notes |
|
||||
|---|---|---|---|---|
|
||||
| disk-usage-small-backups-box-root | push | — | 960s | upside-down, Disk Usage: small-backups-box (/) - Alerts when usage exceed |
|
||||
| system-healthcheck-small-backups-box | push | — | 90s | System Healthcheck: small-backups-box - Regular healthcheck |
|
||||
|
||||
51
archive/uptime_kuma/README.md
Normal file
51
archive/uptime_kuma/README.md
Normal file
|
|
@ -0,0 +1,51 @@
|
|||
# Uptime Kuma — archived
|
||||
|
||||
Uptime Kuma was the monitoring stack for this infrastructure until **2026-09-11**, when
|
||||
it was decommissioned. Everything that referenced it has been removed from the live
|
||||
playbooks; this folder is the record of what it was, kept so the setup can be understood
|
||||
later without digging through git history.
|
||||
|
||||
## Contents
|
||||
|
||||
| File | What it is |
|
||||
|---|---|
|
||||
| `MONITORS.md` | Every monitor that existed, grouped as it was in the UI. The authoritative record. |
|
||||
| `monitors.json` | The same data, machine-readable, as returned by the API. |
|
||||
| `deploy_uptime_kuma_playbook.yml` | How the server itself was deployed (Docker Compose on `monitoring`, behind Caddy). |
|
||||
| `uptime_kuma_vars.yml` | Variables the deploy playbook needed — it is unreadable without these. |
|
||||
| `setup_backup_uptime_kuma_to_lapy.yml` | How its data was backed up, and to where. Useful when disposing of the old volumes. |
|
||||
|
||||
## Why the inventory was captured from the live server, not the repo
|
||||
|
||||
The playbooks only ever created **17** monitors. The live instance had **75**. The
|
||||
difference was created by hand in the UI and existed nowhere else — so a repo-derived
|
||||
list would have silently lost two thirds of the picture. `MONITORS.md` is a snapshot of
|
||||
the real thing, taken immediately before removal.
|
||||
|
||||
Push tokens are deliberately excluded. They are live credentials — anything holding one
|
||||
can report a false "up" — and they become meaningless once the server is gone.
|
||||
|
||||
## What removal did NOT do
|
||||
|
||||
Removing the playbook code does not touch the machines. **48 push monitors** were driven
|
||||
by scripts and systemd timers installed *on the hosts*, which keep firing on their
|
||||
schedule and curling an endpoint that no longer answers. They are harmless but they are
|
||||
still there, writing logs and failing quietly.
|
||||
|
||||
Left behind, per host:
|
||||
|
||||
- `/opt/disk-monitoring` + `disk-usage-monitor.{service,timer}` — on all 12 `managed` hosts
|
||||
- `/opt/system-healthcheck` + `system-healthcheck.{service,timer}` — on all 12 `managed` hosts
|
||||
- `/opt/nodito-monitoring` + `nodito-cpu-temp-monitor.{service,timer}` — `nodito`
|
||||
- `/opt/zfs-monitoring` — `nodito`
|
||||
- `/opt/ups-monitoring` — `nodito`
|
||||
- `bitcoin-knots-healthcheck.{service,timer}` — `bitcoin`
|
||||
- `datum-gateway-healthcheck.{service,timer}` — `bitcoin`
|
||||
- `fulcrum-healthcheck.{service,timer}` — `electrum`
|
||||
- `mempool-{backend,frontend,mariadb}-healthcheck.service` — `mempool`
|
||||
- phoenixd and forgejo-runner healthcheck units — `edge`, `ci_runner`
|
||||
|
||||
Note `nut-monitor.service` on `nodito` is **NUT's own daemon**, not a monitoring
|
||||
leftover — do not remove it with the rest.
|
||||
|
||||
Cleaning these up is a separate decommissioning pass and was not part of the removal.
|
||||
74
archive/uptime_kuma/deploy_uptime_kuma_playbook.yml
Normal file
74
archive/uptime_kuma/deploy_uptime_kuma_playbook.yml
Normal file
|
|
@ -0,0 +1,74 @@
|
|||
- name: Deploy Uptime Kuma with Docker Compose and configure Caddy reverse proxy
|
||||
hosts: monitoring
|
||||
become: yes
|
||||
vars_files:
|
||||
- ../../infra_vars.yml
|
||||
- ../../services_config.yml
|
||||
- ./uptime_kuma_vars.yml
|
||||
vars:
|
||||
uptime_kuma_subdomain: "{{ subdomains.uptime_kuma }}"
|
||||
caddy_sites_dir: "{{ caddy_sites_dir }}"
|
||||
uptime_kuma_domain: "{{ uptime_kuma_subdomain }}.{{ root_domain }}"
|
||||
|
||||
tasks:
|
||||
- name: Create uptime kuma directory
|
||||
file:
|
||||
path: "{{ uptime_kuma_dir }}"
|
||||
state: directory
|
||||
owner: "{{ ansible_user }}"
|
||||
group: "{{ ansible_user }}"
|
||||
mode: '0755'
|
||||
|
||||
- name: Create docker-compose.yml for uptime kuma
|
||||
copy:
|
||||
dest: "{{ uptime_kuma_dir }}/docker-compose.yml"
|
||||
content: |
|
||||
version: "3"
|
||||
services:
|
||||
uptime-kuma:
|
||||
image: louislam/uptime-kuma:latest
|
||||
container_name: uptime-kuma
|
||||
restart: unless-stopped
|
||||
ports:
|
||||
- "{{ uptime_kuma_port }}:3001"
|
||||
volumes:
|
||||
- ./data:/app/data
|
||||
dns:
|
||||
- 1.1.1.1
|
||||
- 9.9.9.9
|
||||
- 8.8.8.8
|
||||
|
||||
- name: Deploy uptime kuma container with docker compose
|
||||
command: docker compose up -d
|
||||
args:
|
||||
chdir: "{{ uptime_kuma_dir }}"
|
||||
|
||||
- name: Ensure Caddy sites-enabled directory exists
|
||||
file:
|
||||
path: /etc/caddy/sites-enabled
|
||||
state: directory
|
||||
owner: root
|
||||
group: root
|
||||
mode: '0755'
|
||||
|
||||
- name: Ensure Caddyfile includes import directive for sites-enabled
|
||||
lineinfile:
|
||||
path: /etc/caddy/Caddyfile
|
||||
line: 'import sites-enabled/*'
|
||||
insertafter: EOF
|
||||
state: present
|
||||
backup: yes
|
||||
|
||||
- name: Create Caddy reverse proxy configuration for uptime kuma
|
||||
copy:
|
||||
dest: "{{ caddy_sites_dir }}/uptime-kuma.conf"
|
||||
content: |
|
||||
{{ uptime_kuma_domain }} {
|
||||
reverse_proxy localhost:{{ uptime_kuma_port }}
|
||||
}
|
||||
owner: root
|
||||
group: root
|
||||
mode: '0644'
|
||||
|
||||
- name: Reload Caddy to apply new config
|
||||
command: systemctl reload caddy
|
||||
1202
archive/uptime_kuma/monitors.json
Normal file
1202
archive/uptime_kuma/monitors.json
Normal file
File diff suppressed because it is too large
Load diff
107
archive/uptime_kuma/setup_backup_uptime_kuma_to_lapy.yml
Normal file
107
archive/uptime_kuma/setup_backup_uptime_kuma_to_lapy.yml
Normal file
|
|
@ -0,0 +1,107 @@
|
|||
- name: Configure local backup for Uptime Kuma from remote
|
||||
hosts: control
|
||||
gather_facts: no
|
||||
vars_files:
|
||||
- ../../infra_vars.yml
|
||||
- ./uptime_kuma_vars.yml
|
||||
vars:
|
||||
remote_data_path: "{{ uptime_kuma_data_dir }}"
|
||||
local_backup_dir: "{{ lookup('env', 'HOME') }}/uptime-kuma-backups"
|
||||
backup_script_path: "{{ lookup('env', 'HOME') }}/.local/bin/uptime_kuma_backup.sh"
|
||||
|
||||
tasks:
|
||||
- name: Debug remote backup vars
|
||||
debug:
|
||||
msg:
|
||||
- "remote_host={{ remote_host }}"
|
||||
- "remote_user={{ remote_user }}"
|
||||
- "remote_data_path='{{ remote_data_path }}'"
|
||||
- "local_backup_dir={{ local_backup_dir }}"
|
||||
|
||||
- name: Ensure local backup directory exists
|
||||
file:
|
||||
path: "{{ local_backup_dir }}"
|
||||
state: directory
|
||||
mode: '0755'
|
||||
|
||||
- name: Ensure ~/.local/bin exists
|
||||
file:
|
||||
path: "{{ lookup('env', 'HOME') }}/.local/bin"
|
||||
state: directory
|
||||
mode: '0755'
|
||||
|
||||
- name: Create backup script
|
||||
copy:
|
||||
dest: "{{ backup_script_path }}"
|
||||
mode: '0750'
|
||||
content: |
|
||||
#!/bin/bash
|
||||
set -euo pipefail
|
||||
|
||||
TIMESTAMP=$(date +'%Y-%m-%d')
|
||||
BACKUP_DIR="{{ local_backup_dir }}/$TIMESTAMP"
|
||||
mkdir -p "$BACKUP_DIR"
|
||||
|
||||
{% if remote_key_file %}
|
||||
SSH_CMD="ssh -i {{ remote_key_file }} -p {{ remote_port }}"
|
||||
{% else %}
|
||||
SSH_CMD="ssh -p {{ remote_port }}"
|
||||
{% endif %}
|
||||
|
||||
rsync -az -e "$SSH_CMD" --delete {{ remote_user }}@{{ remote_host }}:{{ remote_data_path }}/ "$BACKUP_DIR/"
|
||||
|
||||
# Rotate old backups (keep 14 days)
|
||||
# Calculate cutoff date (14 days ago) and delete backups older than that
|
||||
CUTOFF_DATE=$(date -d '14 days ago' +'%Y-%m-%d')
|
||||
for dir in "{{ local_backup_dir }}"/20*; do
|
||||
if [ -d "$dir" ]; then
|
||||
dir_date=$(basename "$dir")
|
||||
if [ "$dir_date" != "$TIMESTAMP" ] && [ "$dir_date" \< "$CUTOFF_DATE" ]; then
|
||||
rm -rf "$dir"
|
||||
fi
|
||||
fi
|
||||
done
|
||||
|
||||
- name: Ensure cronjob for backup exists
|
||||
cron:
|
||||
name: "Uptime Kuma backup"
|
||||
user: "{{ lookup('env', 'USER') }}"
|
||||
job: "{{ backup_script_path }}"
|
||||
minute: 0
|
||||
hour: "9,12,15,18"
|
||||
|
||||
- name: Run the backup script to make the first backup
|
||||
command: "{{ backup_script_path }}"
|
||||
|
||||
- name: Verify backup was created
|
||||
block:
|
||||
- name: Get today's date
|
||||
command: date +'%Y-%m-%d'
|
||||
register: today_date
|
||||
changed_when: false
|
||||
|
||||
- name: Check backup directory exists and contains files
|
||||
stat:
|
||||
path: "{{ local_backup_dir }}/{{ today_date.stdout }}"
|
||||
register: backup_dir_stat
|
||||
|
||||
- name: Verify backup directory exists
|
||||
assert:
|
||||
that:
|
||||
- backup_dir_stat.stat.exists
|
||||
- backup_dir_stat.stat.isdir
|
||||
fail_msg: "Backup directory {{ local_backup_dir }}/{{ today_date.stdout }} was not created"
|
||||
success_msg: "Backup directory {{ local_backup_dir }}/{{ today_date.stdout }} exists"
|
||||
|
||||
- name: Check if backup directory contains files
|
||||
find:
|
||||
paths: "{{ local_backup_dir }}/{{ today_date.stdout }}"
|
||||
recurse: yes
|
||||
register: backup_files
|
||||
|
||||
- name: Verify backup directory is not empty
|
||||
assert:
|
||||
that:
|
||||
- backup_files.files | length > 0
|
||||
fail_msg: "Backup directory {{ local_backup_dir }}/{{ today_date.stdout }} exists but is empty"
|
||||
success_msg: "Backup directory contains {{ backup_files.files | length }} file(s)"
|
||||
15
archive/uptime_kuma/uptime_kuma_vars.yml
Normal file
15
archive/uptime_kuma/uptime_kuma_vars.yml
Normal file
|
|
@ -0,0 +1,15 @@
|
|||
# General
|
||||
uptime_kuma_dir: /opt/uptime-kuma
|
||||
uptime_kuma_data_dir: "{{ uptime_kuma_dir }}/data"
|
||||
uptime_kuma_port: 3001
|
||||
|
||||
# Remote access
|
||||
remote_host_name: "{{ groups['monitoring'] | first }}"
|
||||
remote_host: "{{ hostvars.get(remote_host_name, {}).get('ansible_host', remote_host_name) }}"
|
||||
remote_user: "{{ hostvars.get(remote_host_name, {}).get('ansible_user', 'counterweight') }}"
|
||||
remote_key_file: "{{ hostvars.get(remote_host_name, {}).get('ansible_ssh_private_key_file', '') }}"
|
||||
remote_port: "{{ hostvars.get(remote_host_name, {}).get('ansible_port', 22) }}"
|
||||
|
||||
# Local backup
|
||||
local_backup_dir: "{{ lookup('env', 'HOME') }}/uptime-kuma-backups"
|
||||
backup_script_path: "{{ lookup('env', 'HOME') }}/.local/bin/uptime_kuma_backup.sh"
|
||||
Loading…
Add table
Add a link
Reference in a new issue