gatus: deploy on prd-monitoring, behind Caddy basic auth
First step of replacing Uptime Kuma and ntfy. Gatus runs on the new
observability VPS, fronted by Caddy at status.contrapeso.xyz.
Deployed as the upstream container image, not built from source. The role did
build from source first - their Dockerfile is a bare `CGO_ENABLED=0 go build`,
the Vue dashboard is compiled in via `//go:embed static` in web/static.go, and
CGO can stay off because the sqlite driver is pure-Go modernc.org/sqlite - but
that produces a binary upstream never ran, and it meant compiling the AWS SDK
and gRPC on the smallest box in the estate. That load was heavy enough that
unrelated Ansible tasks timed out while it ran. The cost of the container is a
daemon on the machine whose job is to notice when everything else breaks; that
trade is made deliberately and is written down in the role README.
Pinned by DIGEST, not tag. A tag is mutable - v5.36.0 can be repushed - so
pinning it alone is a weaker promise than it looks:
gatus_image: "ghcr.io/twin/gatus@sha256:c5f210d0..."
`docker compose pull` now either fetches exactly the reviewed image or fails.
gatus_version is kept beside it only so a human can read the release; the two
move together.
The image is FROM scratch, so it has no /etc/passwd and its default user is
root. The container runs as 10001:10001 with the host data dir owned to match,
plus read_only, cap_drop ALL, and no-new-privileges. NET_RAW is added back only
when gatus_allow_icmp, so the capability for icmp:// checks is a visible grant
rather than something inherited from running as root.
Config is a DIRECTORY, not a file. Gatus merges every *.yaml under
GATUS_CONFIG_PATH - maps deep-merge, lists append - so the role owns
00-base.yaml (web, storage, ui, alerting, security) and each service will drop
its own file into endpoints/, the same shape as caddy_site. A primitive defined
twice is ambiguous and upstream refuses it, so anything that is not a list lives
in the base file and nowhere else.
Two bugs the deploy caught:
* Gatus panics on a config with no endpoints ("configuration should contain at
least one endpoint or suite"), so "install now, add endpoints later" is not
a valid state. The role ships endpoints/00-self.yaml checking its own
/health. Less circular than it looks: it proves the directory merged, the
listener serves, and storage accepted a write.
* web.address was carried over from the systemd design as 127.0.0.1. Inside a
container that is the CONTAINER's loopback, which docker-proxy cannot reach
- gatus came up healthy, self-check passing, while every connection to the
published port was refused. It now always binds 0.0.0.0 inside the
container; the isolation comes from publishing to 127.0.0.1 on the host.
Auth is done at the edge, NOT with Gatus's own security.basic. Reading
api/api.go, that middleware protects exactly four routes - the statuses
endpoints. Everything else is registered on the unprotected router, including
/api/v1/config, every badge, and /api/v1/endpoints/:key/uptimes/:duration and
.../response-times/:duration/history, which return real data to anyone who can
guess a key ("<group>_<name>"). Verified against the live instance: all seven
routes returned 200 unauthenticated, and /uptimes/24h returned "1.000000".
So the vhost uses caddy_site_body with a path carve-out rather than
caddy_site_basic_auth, which has no way to exempt a path. The external-endpoint
push API must NOT sit behind basic auth: it authenticates with
`Authorization: Bearer <token>`, and basic auth wants the same header. It is not
unauthenticated - the handler 401s on a missing prefix, an empty token, or a
token that does not match that endpoint's own.
Verified end to end. All seven previously-open routes now 401. The push path
distinguishes cleanly: POST with no auth gets Gatus's own "invalid Authorization
header" with NO WWW-Authenticate; POST with a bogus Bearer gets 404 (key looked
up, no external endpoints yet); GET on the same path gets Caddy's 401 with
WWW-Authenticate: Basic, so the exemption is scoped to POST alone. The
self-check still passes because it polls localhost inside the container and
never traverses Caddy.
The host itself was rebuilt from scratch: 01 (ok=9 changed=8), 02 (ok=12
changed=6), 910_docker (--limit, since that playbook still wrongly claims all of
`managed` needs Docker), caddy (ok=13 changed=8), gatus (ok=18 changed=2).
Not done here: gatus_alerting is still {} - valid, and every condition is
evaluated and recorded, there is just nowhere to shout until a provider is
chosen to replace ntfy.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 22:29:03 +02:00
|
|
|
---
|
|
|
|
|
# Gatus, deployed the way upstream distributes it: the container image.
|
|
|
|
|
#
|
|
|
|
|
# Upstream publishes NO binary release assets - the image is the only artefact
|
|
|
|
|
# they ship, and therefore the only artefact they test. Building from source is
|
|
|
|
|
# possible (`go build` alone is enough; the Vue dashboard is compiled in via
|
|
|
|
|
# `//go:embed static`, and CGO_ENABLED=0 works because the sqlite driver is
|
|
|
|
|
# pure-Go modernc.org/sqlite) but it produces a binary upstream never ran, and
|
|
|
|
|
# it means a full compile of the AWS SDK and gRPC on the smallest box in the
|
|
|
|
|
# estate.
|
|
|
|
|
|
|
|
|
|
# ── Image ────────────────────────────────────────────────────────────────────
|
|
|
|
|
gatus_version: "v5.36.0"
|
|
|
|
|
# Pinned by DIGEST, not by tag. A tag is mutable - `v5.36.0` can be repushed -
|
|
|
|
|
# so pinning the tag alone is a weaker promise than it looks. The digest is the
|
|
|
|
|
# content address: if it resolves, it is byte-for-byte the image reviewed here.
|
|
|
|
|
# Both must be updated together; the tag is kept only so humans can read it.
|
|
|
|
|
gatus_image_digest: "sha256:c5f210d095fa78e6efaa20ffeb14803f2ba4f10615e16a6d12087697149617f0"
|
|
|
|
|
gatus_image: "ghcr.io/twin/gatus@{{ gatus_image_digest }}"
|
|
|
|
|
|
|
|
|
|
# ── Paths (host side) ────────────────────────────────────────────────────────
|
|
|
|
|
gatus_dir: /opt/gatus
|
|
|
|
|
gatus_config_dir: "{{ gatus_dir }}/config"
|
|
|
|
|
gatus_data_dir: "{{ gatus_dir }}/data"
|
|
|
|
|
|
|
|
|
|
# Gatus merges every *.yaml under GATUS_CONFIG_PATH and its subdirectories:
|
|
|
|
|
# maps deep-merge, lists append. That is why this role ships a config DIRECTORY
|
|
|
|
|
# rather than one file - each service contributes its own endpoint file, the
|
|
|
|
|
# same way each service contributes a vhost through `caddy_site`.
|
|
|
|
|
#
|
|
|
|
|
# Primitives must be defined exactly once across all files or the merge is
|
|
|
|
|
# ambiguous, so everything that is not a list lives in the base file and
|
|
|
|
|
# nowhere else.
|
|
|
|
|
gatus_base_config_file: "00-base.yaml"
|
|
|
|
|
gatus_endpoints_dir: "{{ gatus_config_dir }}/endpoints"
|
|
|
|
|
|
|
|
|
|
# ── Identity ─────────────────────────────────────────────────────────────────
|
|
|
|
|
# The image is FROM scratch, so it has no /etc/passwd and no user to drop to by
|
|
|
|
|
# name. Run it by numeric uid/gid instead, and own the data volume to match.
|
|
|
|
|
gatus_uid: 10001
|
|
|
|
|
gatus_gid: 10001
|
|
|
|
|
|
|
|
|
|
# ── Web ──────────────────────────────────────────────────────────────────────
|
|
|
|
|
gatus_port: 8080
|
|
|
|
|
# HOST-side address the container's port is published on. Gatus itself always
|
|
|
|
|
# binds 0.0.0.0 inside the container - see the note in config.yaml.j2. Never
|
|
|
|
|
# publish this on 0.0.0.0: the external-endpoint push API shares the dashboard's
|
|
|
|
|
# listener, and Caddy is what should be in front of both.
|
|
|
|
|
gatus_bind_address: "127.0.0.1"
|
|
|
|
|
gatus_ui_title: "Status"
|
|
|
|
|
gatus_ui_header: "Status"
|
|
|
|
|
|
|
|
|
|
# ── Storage ──────────────────────────────────────────────────────────────────
|
|
|
|
|
# sqlite, not memory: history has to survive a restart, or the dashboard lies
|
|
|
|
|
# about uptime after every deploy. Path is INSIDE the container.
|
|
|
|
|
gatus_storage_type: sqlite
|
|
|
|
|
gatus_storage_path: "/data/gatus.db"
|
|
|
|
|
gatus_storage_caching: true
|
|
|
|
|
|
monitoring: recover host checks for the whole estate, reported to Gatus
Five checks, 27 endpoints, replacing what Uptime Kuma used to watch:
is it up every 5min, all hosts
is disk full daily, all hosts
is CPU hot every 5min, nodito
is ZFS broken daily, nodito
is UPS online every 5min, nodito
Two roles, kept separate so neither knows about the other - they meet at a URL
and a token, the same way caddy_site and each service meet at a vhost:
roles/gatus_endpoint runs on the observability host, writes ONE file into
/opt/gatus/config/endpoints/. Gatus merges every *.yaml
there and appends lists, so callers compose without
coordinating.
roles/healthcheck runs on the monitored host: a check script, a systemd
service, a timer, and an optional push. Ships a library
of check bodies under templates/checks/.
Everything PUSHES. Gatus never reaches out, which matters because nodito and its
VMs are behind NAT, and because four of the five checks are internal state with
no pollable surface at all. Liveness pushes too, deliberately: a heartbeat proves
the host is running AND can reach the internet, where an ICMP probe from one
vantage point only proves it answers pings from there. And since Gatus alerts
when a heartbeat window expires, a check that stops running raises the alarm by
itself - a dead timer looks exactly like a dead host, which is the correct
reading.
One bearer token per host, generated straight into the vault and never printed.
A token only writes results for its own host's endpoints, so a compromised host
can lie about itself, which it could do anyway.
Three things learned from the source that shaped this:
* Gatus polls its own config every 30s and reloads
(main.listenToConfigurationFileChanges), so gatus_endpoint needs no restart
handler - writing the file IS the deploy.
* ...but on a reload it panics if the new config fails to parse, unless
skip-invalid-config-update is set. Endpoint files are contributed by other
playbooks, so one malformed file would take the monitor down at the worst
possible moment. Now set.
* The push URL uses a key Gatus computes, not the name you write:
sanitize(group) + "_" + sanitize(name), lowercased with / _ . , space # + &
replaced by "-" (config/key/key.go). So knots_box_local is
knots-box-local in the URL. The playbook derives it rather than hand-writing.
storage: maximum-number-of-results 900, up from upstream's 100. Gatus bounds the
database by COUNT and trims inline on insert, so there is no retention job and
no way to fill a disk - but history depth is then a function of check frequency,
and 100 results at a 5-minute interval is 8 hours. 900 is ~3 days of liveness
and ~2.5 years of the daily disk check. The uptime table is separate and its
30-day retention is hard-coded upstream.
A bug worth recording: the first deploy shipped five scripts that all died with
"syntax error: unexpected end of file". Jinja strips an included template's
trailing newline and trim_blocks then eats the newline after {% endif %}, so the
closing brace of check() landed on the same line as the body's last statement -
`return 0}`. Every check was broken and the deploy still reported failed=0,
because the role's "run once" task has failed_when: false and reports the result
as a debug message nobody read. The blank line that fixes it is now load-bearing
and commented as such.
Verified by triggering every unit by hand rather than waiting on timers: all
checks exit 0 on all hosts, and Gatus shows 26 UP / 1 DOWN. The one DOWN is
liveness_watchtower, which is honest - that host currently refuses SSH (TCP
connects, no banner exchange) and is excluded from this deploy. It is also the
box still running Uptime Kuma.
Known waste, not yet fixed: healthcheck installs its dependencies per CHECK
rather than per HOST, so apt runs 29 times estate-wide for a curl that is
already present, and daemon_reload runs 4x per host.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 08:54:10 +02:00
|
|
|
# Per-endpoint row caps. Gatus bounds the database by COUNT, not by time, and
|
|
|
|
|
# trims inline on insert (storage/store/sql/sql.go, InsertEndpointResult) - so
|
|
|
|
|
# there is no retention job to write and no way for this to fill a disk.
|
|
|
|
|
#
|
|
|
|
|
# History depth is therefore a function of check frequency, not of days:
|
|
|
|
|
# 900 results is ~3 days of a 5-minute liveness check, and ~2.5 years of a daily
|
|
|
|
|
# disk check. Upstream's default is 100, which would have been 8 hours of
|
|
|
|
|
# liveness - not enough to still see a weekend incident on Monday.
|
|
|
|
|
#
|
|
|
|
|
# Note the uptime table is separate and its 30-day retention is hard-coded
|
|
|
|
|
# upstream (uptimeRetention), so uptime percentages top out at 30 days whatever
|
|
|
|
|
# this is set to.
|
|
|
|
|
gatus_storage_max_results: 900
|
|
|
|
|
gatus_storage_max_events: 50
|
|
|
|
|
|
gatus: deploy on prd-monitoring, behind Caddy basic auth
First step of replacing Uptime Kuma and ntfy. Gatus runs on the new
observability VPS, fronted by Caddy at status.contrapeso.xyz.
Deployed as the upstream container image, not built from source. The role did
build from source first - their Dockerfile is a bare `CGO_ENABLED=0 go build`,
the Vue dashboard is compiled in via `//go:embed static` in web/static.go, and
CGO can stay off because the sqlite driver is pure-Go modernc.org/sqlite - but
that produces a binary upstream never ran, and it meant compiling the AWS SDK
and gRPC on the smallest box in the estate. That load was heavy enough that
unrelated Ansible tasks timed out while it ran. The cost of the container is a
daemon on the machine whose job is to notice when everything else breaks; that
trade is made deliberately and is written down in the role README.
Pinned by DIGEST, not tag. A tag is mutable - v5.36.0 can be repushed - so
pinning it alone is a weaker promise than it looks:
gatus_image: "ghcr.io/twin/gatus@sha256:c5f210d0..."
`docker compose pull` now either fetches exactly the reviewed image or fails.
gatus_version is kept beside it only so a human can read the release; the two
move together.
The image is FROM scratch, so it has no /etc/passwd and its default user is
root. The container runs as 10001:10001 with the host data dir owned to match,
plus read_only, cap_drop ALL, and no-new-privileges. NET_RAW is added back only
when gatus_allow_icmp, so the capability for icmp:// checks is a visible grant
rather than something inherited from running as root.
Config is a DIRECTORY, not a file. Gatus merges every *.yaml under
GATUS_CONFIG_PATH - maps deep-merge, lists append - so the role owns
00-base.yaml (web, storage, ui, alerting, security) and each service will drop
its own file into endpoints/, the same shape as caddy_site. A primitive defined
twice is ambiguous and upstream refuses it, so anything that is not a list lives
in the base file and nowhere else.
Two bugs the deploy caught:
* Gatus panics on a config with no endpoints ("configuration should contain at
least one endpoint or suite"), so "install now, add endpoints later" is not
a valid state. The role ships endpoints/00-self.yaml checking its own
/health. Less circular than it looks: it proves the directory merged, the
listener serves, and storage accepted a write.
* web.address was carried over from the systemd design as 127.0.0.1. Inside a
container that is the CONTAINER's loopback, which docker-proxy cannot reach
- gatus came up healthy, self-check passing, while every connection to the
published port was refused. It now always binds 0.0.0.0 inside the
container; the isolation comes from publishing to 127.0.0.1 on the host.
Auth is done at the edge, NOT with Gatus's own security.basic. Reading
api/api.go, that middleware protects exactly four routes - the statuses
endpoints. Everything else is registered on the unprotected router, including
/api/v1/config, every badge, and /api/v1/endpoints/:key/uptimes/:duration and
.../response-times/:duration/history, which return real data to anyone who can
guess a key ("<group>_<name>"). Verified against the live instance: all seven
routes returned 200 unauthenticated, and /uptimes/24h returned "1.000000".
So the vhost uses caddy_site_body with a path carve-out rather than
caddy_site_basic_auth, which has no way to exempt a path. The external-endpoint
push API must NOT sit behind basic auth: it authenticates with
`Authorization: Bearer <token>`, and basic auth wants the same header. It is not
unauthenticated - the handler 401s on a missing prefix, an empty token, or a
token that does not match that endpoint's own.
Verified end to end. All seven previously-open routes now 401. The push path
distinguishes cleanly: POST with no auth gets Gatus's own "invalid Authorization
header" with NO WWW-Authenticate; POST with a bogus Bearer gets 404 (key looked
up, no external endpoints yet); GET on the same path gets Caddy's 401 with
WWW-Authenticate: Basic, so the exemption is scoped to POST alone. The
self-check still passes because it polls localhost inside the container and
never traverses Caddy.
The host itself was rebuilt from scratch: 01 (ok=9 changed=8), 02 (ok=12
changed=6), 910_docker (--limit, since that playbook still wrongly claims all of
`managed` needs Docker), caddy (ok=13 changed=8), gatus (ok=18 changed=2).
Not done here: gatus_alerting is still {} - valid, and every condition is
evaluated and recorded, there is just nowhere to shout until a provider is
chosen to replace ntfy.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 22:29:03 +02:00
|
|
|
# ── Alerting ─────────────────────────────────────────────────────────────────
|
|
|
|
|
# Pass-through: rendered verbatim under `alerting:`, so any provider Gatus
|
|
|
|
|
# supports works without touching this role. Empty means "check and record,
|
|
|
|
|
# alert nowhere" - valid, and the default until a provider is chosen.
|
|
|
|
|
gatus_alerting: {}
|
|
|
|
|
gatus_default_alerts: []
|
|
|
|
|
|
|
|
|
|
# ── Security ─────────────────────────────────────────────────────────────────
|
|
|
|
|
# gatus_basic_auth: {username: admin, password-bcrypt-base64: "..."}
|
|
|
|
|
gatus_basic_auth: {}
|
|
|
|
|
|
|
|
|
|
gatus_maintenance: {}
|
monitoring: recover host checks for the whole estate, reported to Gatus
Five checks, 27 endpoints, replacing what Uptime Kuma used to watch:
is it up every 5min, all hosts
is disk full daily, all hosts
is CPU hot every 5min, nodito
is ZFS broken daily, nodito
is UPS online every 5min, nodito
Two roles, kept separate so neither knows about the other - they meet at a URL
and a token, the same way caddy_site and each service meet at a vhost:
roles/gatus_endpoint runs on the observability host, writes ONE file into
/opt/gatus/config/endpoints/. Gatus merges every *.yaml
there and appends lists, so callers compose without
coordinating.
roles/healthcheck runs on the monitored host: a check script, a systemd
service, a timer, and an optional push. Ships a library
of check bodies under templates/checks/.
Everything PUSHES. Gatus never reaches out, which matters because nodito and its
VMs are behind NAT, and because four of the five checks are internal state with
no pollable surface at all. Liveness pushes too, deliberately: a heartbeat proves
the host is running AND can reach the internet, where an ICMP probe from one
vantage point only proves it answers pings from there. And since Gatus alerts
when a heartbeat window expires, a check that stops running raises the alarm by
itself - a dead timer looks exactly like a dead host, which is the correct
reading.
One bearer token per host, generated straight into the vault and never printed.
A token only writes results for its own host's endpoints, so a compromised host
can lie about itself, which it could do anyway.
Three things learned from the source that shaped this:
* Gatus polls its own config every 30s and reloads
(main.listenToConfigurationFileChanges), so gatus_endpoint needs no restart
handler - writing the file IS the deploy.
* ...but on a reload it panics if the new config fails to parse, unless
skip-invalid-config-update is set. Endpoint files are contributed by other
playbooks, so one malformed file would take the monitor down at the worst
possible moment. Now set.
* The push URL uses a key Gatus computes, not the name you write:
sanitize(group) + "_" + sanitize(name), lowercased with / _ . , space # + &
replaced by "-" (config/key/key.go). So knots_box_local is
knots-box-local in the URL. The playbook derives it rather than hand-writing.
storage: maximum-number-of-results 900, up from upstream's 100. Gatus bounds the
database by COUNT and trims inline on insert, so there is no retention job and
no way to fill a disk - but history depth is then a function of check frequency,
and 100 results at a 5-minute interval is 8 hours. 900 is ~3 days of liveness
and ~2.5 years of the daily disk check. The uptime table is separate and its
30-day retention is hard-coded upstream.
A bug worth recording: the first deploy shipped five scripts that all died with
"syntax error: unexpected end of file". Jinja strips an included template's
trailing newline and trim_blocks then eats the newline after {% endif %}, so the
closing brace of check() landed on the same line as the body's last statement -
`return 0}`. Every check was broken and the deploy still reported failed=0,
because the role's "run once" task has failed_when: false and reports the result
as a debug message nobody read. The blank line that fixes it is now load-bearing
and commented as such.
Verified by triggering every unit by hand rather than waiting on timers: all
checks exit 0 on all hosts, and Gatus shows 26 UP / 1 DOWN. The one DOWN is
liveness_watchtower, which is honest - that host currently refuses SSH (TCP
connects, no banner exchange) and is excluded from this deploy. It is also the
box still running Uptime Kuma.
Known waste, not yet fixed: healthcheck installs its dependencies per CHECK
rather than per HOST, so apt runs 29 times estate-wide for a curl that is
already present, and daemon_reload runs 4x per host.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 08:54:10 +02:00
|
|
|
|
|
|
|
|
# Keep serving the previous config if a contributed endpoint file is malformed,
|
|
|
|
|
# instead of panicking. See the note in config.yaml.j2.
|
|
|
|
|
gatus_skip_invalid_config_update: true
|
gatus: deploy on prd-monitoring, behind Caddy basic auth
First step of replacing Uptime Kuma and ntfy. Gatus runs on the new
observability VPS, fronted by Caddy at status.contrapeso.xyz.
Deployed as the upstream container image, not built from source. The role did
build from source first - their Dockerfile is a bare `CGO_ENABLED=0 go build`,
the Vue dashboard is compiled in via `//go:embed static` in web/static.go, and
CGO can stay off because the sqlite driver is pure-Go modernc.org/sqlite - but
that produces a binary upstream never ran, and it meant compiling the AWS SDK
and gRPC on the smallest box in the estate. That load was heavy enough that
unrelated Ansible tasks timed out while it ran. The cost of the container is a
daemon on the machine whose job is to notice when everything else breaks; that
trade is made deliberately and is written down in the role README.
Pinned by DIGEST, not tag. A tag is mutable - v5.36.0 can be repushed - so
pinning it alone is a weaker promise than it looks:
gatus_image: "ghcr.io/twin/gatus@sha256:c5f210d0..."
`docker compose pull` now either fetches exactly the reviewed image or fails.
gatus_version is kept beside it only so a human can read the release; the two
move together.
The image is FROM scratch, so it has no /etc/passwd and its default user is
root. The container runs as 10001:10001 with the host data dir owned to match,
plus read_only, cap_drop ALL, and no-new-privileges. NET_RAW is added back only
when gatus_allow_icmp, so the capability for icmp:// checks is a visible grant
rather than something inherited from running as root.
Config is a DIRECTORY, not a file. Gatus merges every *.yaml under
GATUS_CONFIG_PATH - maps deep-merge, lists append - so the role owns
00-base.yaml (web, storage, ui, alerting, security) and each service will drop
its own file into endpoints/, the same shape as caddy_site. A primitive defined
twice is ambiguous and upstream refuses it, so anything that is not a list lives
in the base file and nowhere else.
Two bugs the deploy caught:
* Gatus panics on a config with no endpoints ("configuration should contain at
least one endpoint or suite"), so "install now, add endpoints later" is not
a valid state. The role ships endpoints/00-self.yaml checking its own
/health. Less circular than it looks: it proves the directory merged, the
listener serves, and storage accepted a write.
* web.address was carried over from the systemd design as 127.0.0.1. Inside a
container that is the CONTAINER's loopback, which docker-proxy cannot reach
- gatus came up healthy, self-check passing, while every connection to the
published port was refused. It now always binds 0.0.0.0 inside the
container; the isolation comes from publishing to 127.0.0.1 on the host.
Auth is done at the edge, NOT with Gatus's own security.basic. Reading
api/api.go, that middleware protects exactly four routes - the statuses
endpoints. Everything else is registered on the unprotected router, including
/api/v1/config, every badge, and /api/v1/endpoints/:key/uptimes/:duration and
.../response-times/:duration/history, which return real data to anyone who can
guess a key ("<group>_<name>"). Verified against the live instance: all seven
routes returned 200 unauthenticated, and /uptimes/24h returned "1.000000".
So the vhost uses caddy_site_body with a path carve-out rather than
caddy_site_basic_auth, which has no way to exempt a path. The external-endpoint
push API must NOT sit behind basic auth: it authenticates with
`Authorization: Bearer <token>`, and basic auth wants the same header. It is not
unauthenticated - the handler 401s on a missing prefix, an empty token, or a
token that does not match that endpoint's own.
Verified end to end. All seven previously-open routes now 401. The push path
distinguishes cleanly: POST with no auth gets Gatus's own "invalid Authorization
header" with NO WWW-Authenticate; POST with a bogus Bearer gets 404 (key looked
up, no external endpoints yet); GET on the same path gets Caddy's 401 with
WWW-Authenticate: Basic, so the exemption is scoped to POST alone. The
self-check still passes because it polls localhost inside the container and
never traverses Caddy.
The host itself was rebuilt from scratch: 01 (ok=9 changed=8), 02 (ok=12
changed=6), 910_docker (--limit, since that playbook still wrongly claims all of
`managed` needs Docker), caddy (ok=13 changed=8), gatus (ok=18 changed=2).
Not done here: gatus_alerting is still {} - valid, and every condition is
evaluated and recorded, there is just nowhere to shout until a provider is
chosen to replace ntfy.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 22:29:03 +02:00
|
|
|
gatus_log_level: INFO
|
|
|
|
|
|
|
|
|
|
# ── Self-check ───────────────────────────────────────────────────────────────
|
|
|
|
|
# Gatus panics on a config with no endpoints, so the role always ships one.
|
|
|
|
|
# Turning this off is only safe once another file in endpoints/ provides one.
|
|
|
|
|
gatus_self_check: true
|
|
|
|
|
|
|
|
|
|
# ── ICMP ─────────────────────────────────────────────────────────────────────
|
|
|
|
|
# Gatus supports icmp:// endpoints. Raw ICMP needs CAP_NET_RAW, which the
|
|
|
|
|
# container would get free only if it ran as root; it does not. Set false if
|
|
|
|
|
# you never use icmp:// checks and want the capability dropped entirely.
|
|
|
|
|
gatus_allow_icmp: true
|