caddy: add the caddy_site role
Replaces the four-task Caddy vhost block currently copy-pasted into 10
playbooks. Nothing calls it yet; this commit only adds the role.
Verified by rendering all 10 sites through the template and diffing against
what the current playbooks produce: 9 of 10 byte-identical. The tenth is
datum-gateway, where the resolvers comment is standardised, rewriting one
comment line Caddy ignores.
Then dry-run against the live hosts (--check, nothing written):
- vipy: forgejo, vaultwarden, lnbits, personal-blog, ntfy-emergency-app
all report ok/unchanged against the real files
- watchtower: ntfy renders identical via caddy_site_body, blank line and
{host}{uri} placeholders intact
- spacey: headscale renders identical when given the config that is
actually running
- memos, mempool, datum-gateway report changed - the comment, as expected
All 14 site files on all 3 hosts confirmed unchanged afterwards.
Two things the build turned up:
- Ansible does not template dict *keys*, so caddy_site_basic_auth is a list
of {user, hash}. As a dict, a Jinja username passes through literally.
The assert refuses a mapping.
- `caddy validate` does accept a single site fragment - rc=0 on a good one,
rc=1 with a line number on a broken one. This was the plan's one untested
claim. A failed validate leaves the live file untouched.
The reload is now a handler, so it fires once at end of play rather than
immediately; anything needing the new config live mid-play must
flush_handlers first.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-11 23:10:43 +02:00
|
|
|
# `caddy_site`
|
|
|
|
|
|
|
|
|
|
Writes one Caddy site file into `{{ caddy_sites_dir }}`, makes sure the main
|
|
|
|
|
Caddyfile imports that directory, validates the result, and reloads Caddy once.
|
|
|
|
|
|
|
|
|
|
Replaces the four-task block that was copy-pasted into 10 playbooks.
|
|
|
|
|
|
|
|
|
|
Runs on any host in the `[caddy]` group — `edge` (vipy), `monitoring`
|
|
|
|
|
(watchtower) and `vpn_control` (spacey).
|
|
|
|
|
|
|
|
|
|
## Usage
|
|
|
|
|
|
|
|
|
|
```yaml
|
|
|
|
|
- ansible.builtin.include_role:
|
|
|
|
|
name: caddy_site
|
|
|
|
|
vars:
|
|
|
|
|
caddy_site_name: forgejo # -> forgejo.conf
|
|
|
|
|
caddy_site_domain: "{{ forgejo_domain }}"
|
|
|
|
|
caddy_site_upstream: "localhost:{{ forgejo_port }}"
|
|
|
|
|
```
|
|
|
|
|
|
|
|
|
|
Use `include_role`, not a `roles:` block, so the call stays in task order next
|
|
|
|
|
to the tasks it depends on. Variables passed this way are scoped to the include
|
|
|
|
|
and do not leak into later calls — so **every call must pass everything it
|
|
|
|
|
needs**; nothing carries over.
|
|
|
|
|
|
|
|
|
|
## Shapes
|
|
|
|
|
|
|
|
|
|
Pick exactly one of `caddy_site_upstream`, `caddy_site_root`, `caddy_site_body`.
|
|
|
|
|
|
|
|
|
|
| Want | Set |
|
|
|
|
|
|---|---|
|
|
|
|
|
| `reverse_proxy host:port` | `caddy_site_upstream` |
|
|
|
|
|
| static `root *` + `file_server` | `caddy_site_root` |
|
|
|
|
|
| anything else | `caddy_site_body` (raw, indented 4 for you) |
|
|
|
|
|
|
|
|
|
|
`caddy_site_upstream` accepts two modifiers, which add a block to the
|
|
|
|
|
`reverse_proxy`:
|
|
|
|
|
|
|
|
|
|
- `caddy_site_headers_up: {"X-Forwarded-Host": "..."}`
|
|
|
|
|
- `caddy_site_resolvers: "100.100.100.100"` — Tailscale MagicDNS
|
|
|
|
|
|
|
|
|
|
and `caddy_site_basic_auth` wraps the site in a `basic_auth` block.
|
|
|
|
|
|
|
|
|
|
## `caddy_site_basic_auth` is a LIST, not a dict
|
|
|
|
|
|
|
|
|
|
```yaml
|
|
|
|
|
caddy_site_basic_auth:
|
|
|
|
|
- user: "{{ datum_dashboard_username }}"
|
|
|
|
|
hash: "{{ datum_dashboard_password_hash }}"
|
|
|
|
|
```
|
|
|
|
|
|
|
|
|
|
**Ansible does not template dictionary keys.** With `{ "{{ user }}": "hash" }`
|
|
|
|
|
the value is rendered and the key is not, so the literal string
|
|
|
|
|
`{{ datum_dashboard_username }}` lands in the config file. Found while building
|
|
|
|
|
this role; the `assert` refuses a mapping so it cannot happen again.
|
|
|
|
|
|
|
|
|
|
## Secrets and `--diff`
|
|
|
|
|
|
|
|
|
|
Rendered site files can carry credentials — `datum-gateway.conf` holds a bcrypt
|
|
|
|
|
hash — and `--diff` prints rendered content. The template task therefore sets
|
|
|
|
|
`diff: "{{ caddy_site_reveal | bool }}"`, default `false`, so `--diff` runs are
|
|
|
|
|
safe everywhere. Pass `-e caddy_site_reveal=true` to see what moved on a site
|
|
|
|
|
you know is not secret.
|
|
|
|
|
|
|
|
|
|
## Validation
|
|
|
|
|
|
|
|
|
|
`validate: "caddy validate --adapter caddyfile --config %s"` runs against the
|
|
|
|
|
rendered temp file before it is moved into place. Verified on vipy that a single
|
|
|
|
|
site fragment validates cleanly (rc=0, `Valid configuration`) and that a
|
|
|
|
|
malformed one is rejected (rc=1, with the syntax error and line number). A
|
|
|
|
|
failed validate leaves the live file untouched, so a broken config can no longer
|
|
|
|
|
reach a running Caddy.
|
|
|
|
|
|
|
|
|
|
What it cannot catch is a conflict with the global `/etc/caddy/Caddyfile`.
|
|
|
|
|
|
|
|
|
|
## The reload is a handler
|
|
|
|
|
|
|
|
|
|
`Reload caddy` fires **once, at the end of the play**, however many sites
|
|
|
|
|
notified it. The code this replaced ran `command: systemctl reload caddy`
|
|
|
|
|
immediately, mid-play. If a later task in the same play needs the new config to
|
|
|
|
|
be live, flush first:
|
|
|
|
|
|
|
|
|
|
```yaml
|
|
|
|
|
- ansible.builtin.meta: flush_handlers
|
|
|
|
|
```
|
|
|
|
|
|
|
|
|
|
## Known intentional difference
|
|
|
|
|
|
|
|
|
|
The `resolvers` block is commented `# Use Tailscale MagicDNS to resolve the
|
|
|
|
|
upstream hostname` in every case. `datum-gateway` previously said `# Resolve via
|
|
|
|
|
Tailscale MagicDNS`. Migrating it therefore rewrites one comment line, which
|
|
|
|
|
Caddy ignores. Every other site renders byte-identical to what its playbook
|
|
|
|
|
produced.
|
caddy: close out Plan 4
All five close-out greps return nothing: no sites-enabled handling outside
roles/, no `systemctl reload caddy`, no caddy_sites_dir self-reference, no
inline proxy unit writes. 37 playbooks syntax clean. The 14 Caddy site files
and 6 proxy units on the hosts are byte-identical to the Stage 0 baseline.
Seven play names still said "on vipy" while the play targeted a group. Renamed
to "on the edge host" - the last place a play claimed a hostname after Plan 2.
Documented the four vhosts in /etc/caddy/sites-enabled that no playbook writes
(uptime-kuma, arbretstaging, bitcoininfra, scriberr) in the caddy_site README.
None deleted.
uptime-kuma.conf was going to be deleted as dead config. It is not dead: the
louislam/uptime-kuma container is STILL RUNNING on watchtower - created
2026-02-07, restart=unless-stopped, healthy - and uptime.contrapeso.xyz returns
302, not the 502 a dead backend would give. The "decommissioning" retired the
Ansible code and the vault credentials, not the service. PLAN_3 claimed "the
tokens died with the server"; that is corrected there.
The Caddyfile.* backups are kept: one per host, Nov-Dec 2025, not churning, and
the only record of each Caddyfile before the import line was added.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-11 23:53:09 +02:00
|
|
|
|
|
|
|
|
## Sites on the hosts that this role does NOT manage
|
|
|
|
|
|
|
|
|
|
Four vhosts exist in `/etc/caddy/sites-enabled/` that no playbook writes. They
|
|
|
|
|
were made by hand. The role only ever writes the one file it is told to, so it
|
|
|
|
|
leaves them alone — but nothing in the repo records them, and that is why they
|
|
|
|
|
are listed here. Checked 2026-09-11:
|
|
|
|
|
|
|
|
|
|
| File | Host | Serves | State |
|
|
|
|
|
|---|---|---|---|
|
|
|
|
|
| `uptime-kuma.conf` | watchtower | `localhost:3001` | **HTTP 302 — still live**, see below |
|
|
|
|
|
| `arbretstaging.conf` | vipy | `arbret-staging-box:80` via MagicDNS | HTTP 200 |
|
|
|
|
|
| `bitcoininfra.conf` | vipy | static `file_server` from `/var/www/bitcoin-services-home` | HTTP 200 |
|
|
|
|
|
| `scriberr.conf` | vipy | `scriberr-box:8080` via MagicDNS | HTTP 502 — upstream down |
|
|
|
|
|
|
|
|
|
|
**`uptime-kuma.conf` must not be deleted as dead config.** Uptime Kuma was
|
|
|
|
|
"decommissioned" in the repo — its playbooks archived and its credentials pulled
|
|
|
|
|
from the vault — but the container is **still running** on watchtower
|
|
|
|
|
(`louislam/uptime-kuma:latest`, created 2026-02-07, `restart=unless-stopped`)
|
|
|
|
|
and is still reachable at its public subdomain. Only the Ansible code was
|
|
|
|
|
retired; the service was not. See `archive/uptime_kuma/`.
|
|
|
|
|
|
|
|
|
|
`scriberr` returning 502 is the one that looks like genuine rot: it proxies to a
|
|
|
|
|
`scriberr-box` that is not answering, and `scriberr-box` is not in the inventory.
|