personal_infra/ansible/roles/socket_proxy/README.md
counterweight c4094b692f
bitcoin-knots, fulcrum, datum-gateway: add and use the socket_proxy role
Three near-identical hosts: edge plays become one role plus three short
calls. 183 lines removed, 34 added, plus a 111-line role.

Verified before touching any playbook: all six live units on vipy reproduced
byte-identically. Then --limit edge --check per playbook - bitcoin-knots and
fulcrum changed=0; datum-gateway changed=2, both attributable to the already
known caddy_site comment line and the Reload caddy handler it triggers.
The 6 units and 14 Caddy files on the hosts are byte-identical afterwards.

PLAN_4 claimed these three plays had "no behavioural drift at all". That was
wrong - it came from a diff truncated by head -60. The live bitcoin-p2p-proxy
units carry four settings this playbook never wrote:

  .socket   Documentation=, FreeBind=true
  .service  Documentation=, TimeoutStopSec=5,
            StandardOutput=journal, StandardError=journal

FreeBind is the one that matters: it lets the socket bind to an address that
is not up yet, so without it the socket can fail to start on boot. Running
the bitcoin-knots playbook would have stripped it. Same class of hazard as
headscale. The role expresses all four; bitcoin-p2p is the only caller that
passes any.

Also: UFW treats the rule comment as part of the rule. datum-stratum's live
comment is "DATUM Gateway Stratum public access" but the role's derived
default produced "DATUM Stratum public access", which rewrote the rule.
Caught in the dry-run; datum now passes the comment explicitly.

Two deliberate differences from the original, both documented in the README:
ignore_errors: yes on the upstream check became failed_when: false, and the
handler restarts the .socket, which drops connections open through it - it
fires only when a unit file actually changes.

The inert Uptime Kuma TCP monitor blocks stay in the playbooks rather than
being pulled into a new role (12/12/18 guarded tasks).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-11 23:50:15 +02:00

2.7 KiB

socket_proxy

Exposes a service running on a private Tailscale host through a public TCP port on an edge machine, using systemd-socket-proxyd. Writes a .socket and a .service unit, enables the socket, opens the UFW port, and checks the upstream is reachable.

Usage

- ansible.builtin.include_role:
    name: socket_proxy
  vars:
    socket_proxy_name: fulcrum-ssl          # -> fulcrum-ssl-proxy.{socket,service}
    socket_proxy_description: "Fulcrum SSL" # -> "Fulcrum SSL Proxy Socket"
    socket_proxy_listen_port: "{{ fulcrum_ssl_port }}"
    socket_proxy_upstream_host: "{{ fulcrum_tailscale_hostname }}"

socket_proxy_upstream_port defaults to socket_proxy_listen_port, which is what all three current callers want.

Optional unit settings

These exist because the live bitcoin-p2p-proxy units on vipy carried settings the playbook never wrote. Somebody added them by hand, so running deploy_bitcoin_knots_playbook.yml would have silently removed them:

Variable Emits Why it matters
socket_proxy_free_bind FreeBind=true in [Socket] Lets the socket bind to an address that is not up yet. Without it the socket can fail to start on boot.
socket_proxy_documentation Documentation= in both units Cosmetic.
socket_proxy_timeout_stop_sec TimeoutStopSec= Bounds how long a stop can hang.
socket_proxy_log_to_journal StandardOutput=journal + StandardError=journal Cosmetic on modern systemd, which defaults to the journal anyway.

Only bitcoin-p2p passes any of them.

socket_proxy_ufw_comment

Defaults to "<description> public access", which reproduces the live rule comment for bitcoin-p2p and fulcrum-ssl. datum-stratum must pass it explicitly — its live comment is DATUM Gateway Stratum public access while the derived default would be DATUM Stratum public access, and UFW treats the comment as part of the rule, so the mismatch rewrites the rule on every run.

The upstream check never fails the play

wait_for on the upstream carries failed_when: false. The proxy is correctly configured whether or not the backend happens to be up, and this is the one task that depends on another machine. The original plays used ignore_errors: yes, which prints a red "ignoring" line; failed_when: false is the quieter equivalent.

Restarts

The handler restarts the .socket, not the .service — that is what picks up a changed unit; the service is started by the socket on the next connection.

Restarting a socket drops connections that are currently open through it. For bitcoin-p2p that means peers reconnect; for datum-stratum it means a mining client has to reconnect and may lose in-flight shares. The handler only fires when a unit file actually changes.