Skip to content

feat(hub): make a remote caucus hub work, and shut the doors it left open - #101

Merged
obeone merged 36 commits into
mainfrom
feat/remote-hub
Sep 18, 2026
Merged

obeone merged 36 commits into
mainfrom
feat/remote-hub

Conversation

@obeone

@obeone obeone commented Sep 18, 2026

Copy link
Copy Markdown
Owner

What changed

A hub bound anywhere but loopback was an open room. POST /register was unauthenticated and /mcp had no auth at all, so anything that could reach the port joined the caucus and read everything said in it. caucus-hub accepted --host 0.0.0.0 silently while caucus-setup-service refused it, the /mcp DNS-rebinding guard only ever learned the bind address (so a remote client's Host header was rejected with no way to allow it), and several pieces a remote setup needs already existed but were documented nowhere: CAUCUS_ALLOW_REMOTE_HUB, --allowed-origin.

This branch makes a multi-machine caucus actually work, and says plainly what it does and does not protect.

  • A shared agent key (--agent-key / CAUCUS_AGENT_KEY) on /register and on the whole /mcp mount, independent of the existing operator and observer console tokens.
  • --allowed-host and --public-url, so a remote client is not rejected by the host guard and a remote agent is not handed a 127.0.0.1 URL it cannot reach.
  • A non-loopback bind now refuses to start without both credentials, plus --public-url on a wildcard bind. --allow-insecure-bind is the escape hatch. The refusal message is written for a human who just got refused: it names both flags, both env vars, and the way back to loopback.
  • /peers, /channels, /forms and /ping are gated too. They were handing out the peer roster, every private channel's name, topic and members, the text of every pending operator form, and each peer's own set_status string.
  • watch_command on a remote hub hands out a single-use 120 second ticket, redeemed at POST /watch-ticket/redeem, instead of printing the peer token into the agent's transcript. That token is the room bearer for /receive, /send, /ack, /channels/*, /ask and /floor, and the agent key gates none of those.
  • docs/remote-hub.md: a walkthrough for both connection paths, a flags and env reference, a Caddy TLS example (the hub has no TLS of its own), troubleshooting, and a section on what the key does not buy.

Loopback deployments are unchanged. With no agent key configured, nothing in here fires.

Breaking change

Anyone binding caucus-hub to a non-loopback address today must now set --agent-key and --operator-token (and --public-url on 0.0.0.0 or ::), or pass --allow-insecure-bind.

What the key does not buy

This is in the doc for the same reason it is here: the key is a perimeter, not authorization. It turns "anyone who can reach the port" into "anyone holding one shared secret", and inside the room it grants nothing extra and restricts nothing. Any keyholder can self-join any private channel with no invite check, take the broadcast floor and block every other sender's say(), and take over a peer name whenever that peer's listener is not currently polling /receive. There is no per-agent identity, and no revocation short of restarting the hub with a new key.

Two more limits worth having in the open. The hub still has no native TLS and expects a reverse proxy in front of anything crossing a network you do not already trust. And the /register rate-limit bucket is keyed on request.client.host with no X-Forwarded-For handling, so behind a proxy every agent shares one bucket: a registration burst can trip it for the whole fleet, not just the noisy one.

How I verified it

1072 tests pass, mypy src/ is clean under strict, and ruff check src/ sits at its pre-existing 9 errors, all unrelated to this branch and being fixed separately. No version bump, the version comes from git tags in this repo.

One thing I want the reviewer to see rather than find. tests/test_remote_hub.py::test_a_reaped_peer_is_revived_through_a_redeemed_ticket failed once on a full-suite run; the 151 runs after that were green, and I never got it to fail again. Reading the test, it ran two clocks at once: the ticket was minted against the wall clock while the reap was driven off a synthetic instant, so the ordering it depends on only held while the suite was fast. 4d75bfc drives both off one reference instant, placed at the midpoint of client_ttl and WATCH_TICKET_TTL, so the relationship is arithmetic on the constants instead of a race with the wall clock. The original traceback is lost, so that diagnosis is inspection, not a captured failure. If someone sees it go red again, I would like to know.

Checklist

  • User-visible changes are recorded under ## [Unreleased] in CHANGELOG.md.
  • If PROTOCOL_TEXT changed in hub.py, PROTOCOL_VERSION was bumped too. (PROTOCOL_TEXT is untouched on this branch.)

/register was unauthenticated and /mcp had no auth at all, so anything
that could reach the port could join the caucus and read the room. Add
--agent-key (env CAUCUS_AGENT_KEY): when set, both doors demand
Authorization: Bearer <key>.

/register checks it before the per-host throttle, so a bogus key never
spends another caller's budget. /mcp is gated once at the HTTP layer
rather than per tool, inside the CORS layer so a preflight OPTIONS (which
carries no Authorization by definition) is still answered. The
DNS-rebinding Host/Origin allowlist is untouched.

The key is independent of --operator-token/--observer-token, which keep
guarding only the console. With no key set, nothing changes. A
non-loopback bind without one now warns at startup.
HubConnector gains an agent_key argument defaulting to the environment
variable, so caucus-claude-agent inherits it with no wiring; the bridge
reads the same variable at module level. Both present it on /register
only: that is the one call made before the process holds a peer token,
and every later call already spends that token as its own bearer.

caucus-watch needs nothing: it never registers, it is handed a peer
token.
The plugin connects straight to /mcp, so it needs the key too. Claude
Code expands ${VAR:-default} in an http server's headers, and an unset
variable with an empty default becomes an empty string, so the same
entry serves a keyed remote hub and an open local one.
Open behaviour unchanged with no key; correct key accepted on /register
and /mcp; wrong, missing and malformed keys refused with 401; the
refusal precedes the register throttle; the CORS preflight still passes;
and the in-process HubConnector tool path still works behind the gate.
validate_public_url is the server-side counterpart to validate_hub_url: it
checks the origin the hub advertises to agents is http or https, names a
host, and carries nothing past the origin, since the hub appends its own
paths to that value.
The watcher token file lives on the hub's filesystem, which is not the
agent's when the hub serves other machines, so its path names nothing the
agent can run. build_mcp_server now takes a remote flag; when set,
watch_command returns the CAUCUS_TOKEN env form caucus-watch already
accepts and writes no file. The loopback path is untouched.
--allowed-host (env CAUCUS_ALLOWED_HOSTS) adds entries to the /mcp
DNS-rebinding guard, which previously only ever learned the bind address:
a hub on 0.0.0.0 reached as hub.lan:8765 was refused with no way to allow
it. A bare host is allowed on the hub's own port, host:port is kept as
given, and the guard itself is unchanged.

--public-url (env CAUCUS_PUBLIC_URL) is the base URL other machines reach
the hub at. It replaces the 127.0.0.1 a wildcard bind used to advertise,
so watch_command hands out a runnable command, and its netloc joins the
Host allowlist.

A non-loopback bind now refuses to start unless both an operator token and
an agent key are set; --allow-insecure-bind is the escape hatch. The
refusal names both doors, both flags, both env vars, the two flags needed
to make the hub reachable, and the way back to loopback.
check_bind asked only for the operator token, which never guarded
/register or /mcp: the installer would happily write a unit whose agent
door was open to the whole network. It now demands both credentials,
matching the gate caucus-hub applies at startup, and --agent-key carries
the key into the launchd plist and the systemd env file alongside the
dashboard tokens. README's one-line description of that gate follows.
allowed-host parsing (bare, host:port, bracketed IPv6, env, merge),
public-url accept/reject cases, watch_command in both deployments, the
_mount_mcp_http wiring, and the startup refusal driven through main()
with uvicorn stubbed. test_setup_service follows check_bind's new
signature and covers the agent key's plumbing.
Document running a hub on one machine with agents joining from others:
the /mcp and caucus-bridge connection paths end to end, the flags/env
reference, a Caddy TLS example, and the failure modes an operator
actually hits. Also corrects ARCHITECTURE.md's outdated claim that /ui
carries no authentication now that AuthConfig exists, and notes the
agent-key middleware guarding /register and /mcp.
Three modules spelled "loopback" as the exact-string frozenset
{127.0.0.1, localhost, ::1} while urlguard did a real ipaddress check,
so --host 127.0.0.2 was refused as non-loopback by the hub's bind gate
and by the service installer while urlguard called it loopback, and
--host [::1] likewise. Promote urlguard's check to is_loopback_host()
and make it the one definition every call site reads.

Add needs_remote_optin() alongside it: validate_hub_url's question
asked without consulting this process's environment, so the hub can
tell whether a URL will be refused on somebody else's machine.
/peers, /channels and /forms answer before a caller holds a peer token,
so the only credential they can carry is the shared agent key. The
bridge and the connector sent it on /register alone, and the in-process
connector behind the MCP tools read it from the environment, which
--agent-key on the command line never sets.

Send it on all three read calls and take it from the live AuthConfig
in-process, so the hub can start gating them without cutting off its
own clients. No behaviour change against a hub with no key configured.

Rename the bridge's _register_headers to _agent_headers, which is now
what it is.
With --public-url http://hub.lan:8765, watch_command handed a remote
agent `caucus-watch --hub http://hub.lan:8765`. caucus-watch runs
validate_hub_url on that value and exits 2 before its first poll, since
plain http to a non-loopback host needs CAUCUS_ALLOW_REMOTE_HUB. The
agent backgrounds the process and believes a watcher is listening,
while nothing is.

Prefix the command with the opt-in when the advertised URL is that
shape. Decided with needs_remote_optin rather than validate_hub_url, so
the answer does not depend on the hub process's own environment: the
command runs in the agent's.
Three fixes at the HTTP layer, all of them about a hub bound wider than
loopback with an agent key set:

/peers, /channels and /forms were unauthenticated, so anyone who could
reach the port still read the peer roster, every peer's self-reported
status, every private channel's name, topic and membership, and the
text of every pending operator form. Gate all three on the agent key,
accepting an operator or observer token too so the console keeps
working. The role check is guarded on AuthConfig.enabled, because
role_for grades every caller as operator when no operator token is set.
/ping stays open as the liveness probe, with its docstring now saying
what it does disclose.

secrets.compare_digest raises TypeError on a str holding anything
outside ASCII, and ASGI decodes headers as latin-1, so
`Authorization: Bearer \xe9` was an unhandled 500 rather than a 401 --
and on /register the gate runs before the rate-limit bucket, so nothing
throttled a caller repeating it. Compare UTF-8 bytes in all three
credential comparisons.

An OPTIONS on /mcp skipped the key gate and reached the SDK, which
built a session transport and a task group before answering 405. A
genuine CORS preflight is answered by the CORS layer wrapping this
gate, so an OPTIONS arriving here is not one: gate it like any other
method.
--host 0.0.0.0 is rewritten to 127.0.0.1 for the address the hub
advertises, but the mount is still marked remote, so a remote agent got
a watcher command pointing back at its own machine. Fix it at the
source: fold --public-url into the existing non-loopback bind gate as a
third required setting when the bind host is a wildcard, keeping the
per-credential MISSING / already set shape and the --allow-insecure-bind
escape hatch. A concrete bind such as --host 192.168.1.10 already
advertises itself, so nothing is demanded there.

Also in this pass:

- Use is_loopback_host everywhere the bind gate, the /mcp default and
  the launcher precondition used to spell loopback as three exact
  strings. --host 127.0.0.2 and --host ::1 now behave as the loopback
  addresses they are, so they skip the gate.
- Bracket a bare IPv6 --allowed-host: `::1` and `2001:db8::1` contain a
  colon, so they were passed through verbatim and matched no Host header
  the guard will ever see, while the operator believed otherwise.
- De-duplicate the allowlist: the bind address is no longer appended a
  second time when the operator also named it with --allowed-host.
- Normalise blank credentials to None. `--agent-key ""` made agent_ok
  reject everyone and `--operator-token ""` turned auth on with a token
  no first frame can match; the clients already normalise this way.
- Warn once at startup when the advertised URL is plain http to a
  non-loopback host: peer tokens and message content cross the network
  in clear, and the watcher command has to carry the opt-in to run.
check_bind demanded an agent key for a non-loopback install, then wrote
a unit passing only --host/--port/--no-browser: an operator following
the refusal installed a hub with /mcp off (the non-loopback default)
advertising an address nothing off-box can dial.

Plumb --public-url, --allowed-host and --mcp-http through the same
route --agent-key takes -- validated, then into the plist's
EnvironmentVariables and the systemd environment file, both now built
from one service_environment() so neither can drift -- and show them in
the plan. Extend check_bind to demand --public-url on a wildcard bind,
matching the gate caucus-hub applies at startup, so the refusal comes
while installing rather than from a unit that loads and then exits.

Loopback is decided by the shared is_loopback_host, so the installer
and the hub no longer disagree about 127.0.0.2.
Fold the public-url/allowed-host, agent-key and non-loopback bind
refusal refinements into their existing entries, and add the new
Fixed and Security items from this round of review.
A remote agent cannot be handed a token file (the path names nothing on its
machine), and handing it the peer token puts the room bearer for /receive,
/send, /ack, /channels/*, /ask and /floor into the agent's transcript, its
shell history and the watcher's environ. The token file exists precisely to
keep it out of argv and out of the launching transcript.

Add the store the remote path needs instead: issue_watch_ticket mints an
opaque claim check with a 120s TTL, redeem_watch_ticket spends it exactly
once and returns the peer token. Expired entries are pruned on issue, on
redeem and in the idle-reaper sweep, so the store is bounded by the tickets
issued inside one TTL window. Redemption compares with secrets.compare_digest
on UTF-8 bytes, like the hub's other credentials.

A ticket resurrects nothing: it yields a token, and a token the hub has since
forgotten is refused by client_for exactly as it is today.
Two holes in the same wall, both only reachable once a hub leaves loopback.

POST /watch-ticket/redeem exchanges a single-use watch ticket for the peer
token it stands for. It is gated on the shared agent key the same way
/register is, so a keyed hub does not hand tokens to whoever can reach the
port. An unknown, already spent or expired ticket answers 404, not 401: the
caller's credential was fine, and conflating the two would have the watcher
report a dead session when all it needs is a fresh watch_command(). Neither
the ticket nor the token is ever logged.

/ping returned far more than liveness: whether a named peer exists, how long
since it last touched the hub, whether a listener is attached, and the peer's
own set_status prose. On a keyed non-loopback hub that was an unauthenticated
disclosure, so it now goes through _require_agent_read like /peers, /channels
and /forms. The connector and the stdio bridge present the key on their ping
calls accordingly; with no key configured the endpoint stays open, as before.

PROTOCOL_TEXT is unchanged (it never named the credential form), so
PROTOCOL_VERSION is not bumped.
--ticket and CAUCUS_TICKET join the credential chain. When a ticket is
supplied and no direct token is, the watcher spends it once against
POST /watch-ticket/redeem at startup, presenting CAUCUS_AGENT_KEY the way
hub_connector already reads it, then polls exactly as before.

Precedence, flags before environment as the module docstring has always
promised: --token > --token-file > --ticket > CAUCUS_TOKEN > CAUCUS_TICKET.
The historical order between the three token forms is untouched.

A refused or unredeemable ticket exits 1 after printing the remedy to stdout,
not a traceback: the agent is woken by this process exiting and reads what it
left there, so the line says the ticket is single-use and short-lived and that
watch_command() issues another.
The remote branch of watch_command emitted CAUCUS_TOKEN=<peer token> inline,
which is the room bearer for /receive, /send, /ack, /channels/*, /ask and
/floor, and the agent key gates none of those. It landed in the agent's
transcript, its shell history and the watcher's environ, undoing exactly what
the loopback token file was built to prevent.

Emit caucus-watch --hub <url> --ticket <ticket> instead, keeping the
CAUCUS_ALLOW_REMOTE_HUB=1 prefix logic. The token now appears nowhere in the
tool result. The loopback branch and its 0600 token file are untouched.
…ed /ping

Ticket store and endpoint: redeems once and the replay 404s, an expired ticket
fails with the clock driven rather than slept on, an unknown ticket fails, the
idle reaper sweeps the store, redemption demands the agent key when one is
configured (and the refusal fires before the lookup, so the ticket survives
it), and a ticket for a dead peer hands back a token client_for still refuses.

watch_command: the remote result carries the ticket and the peer token appears
nowhere in the payload, not merely outside the command string; the loopback
0600 token file is unchanged.

Watcher: the five-step credential precedence, and redemption against a live
hub including the keyed and unreachable cases. The four _resolve_token cases
were rewritten onto _resolve_credential, which replaces it.

/ping: open on an unkeyed hub, 401 without a credential on a keyed one, served
with the agent key, and the connector carries that key on the call.
…reat model

Brings docs/remote-hub.md and docs/ARCHITECTURE.md in line with what landed
on the branch since the guide was written: watch_command() on a remote hub
now hands out a single-use ticket instead of the peer token, the bind gate
also requires --public-url on a wildcard host, and /peers, /channels, /forms
and /ping are gated on the agent key. Adds a threat-model section spelling
out what the agent key does not buy.
watch_command() minted a fresh single-use ticket on every call without
retiring the one it replaced, so N refresh calls left N live bearer
credentials outstanding for the same peer token, each already sitting in
the agent's transcript for the full 120s TTL. Track the current ticket
on the session's membership and revoke it (new HubState.revoke_watch_ticket)
before minting the next one, on leave(), and in the dead-session sweep, so
at most one ticket per member is ever redeemable.
issue_watch_ticket()'s docstring claimed a reaped peer can never be
resurrected through a stale ticket, but client_for() revives from the
reaped graveyard inside reaped_grace regardless of how the token arrived,
so a ticket redeemed in that window resurrects the peer same as a direct
token would. Correct the docstring, and cover the grace-window path with
a test (the existing dead-peer test only exercises a hard unregister()).

watch.py's docstring also still said --token was the only credential form
landed in argv, but --ticket is passed in argv too. State that plainly,
what it exposes to another local uid via ps, and why the exposure stays
bounded (single-use, short TTL) without moving it to stdin.
The startup guard only ever looked at the bind address, so a hub on
127.0.0.1 exposed through a Cloudflare Tunnel, ngrok or a local reverse
proxy started without a word: no agent key, so agent_ok() passed every
caller, and no operator token, so every browser was graded as operator.
/register, /peers and /watch-ticket/redeem were open to whoever reached
the tunnel.

_mount_mcp_http already read a declared public URL as "this hub is
remote" when deciding what watch_command hands an agent. The credentials
gate now reads it the same way: a --public-url naming a non-loopback host
demands both --operator-token and --agent-key, with --allow-insecure-bind
as the same escape hatch the bind guard offers. A loopback public URL is
just a nicer address for this machine and arms nothing.

_insecure_public_url_message carries the refusal; the doors inventory and
the openssl snippet both messages share move into _doors_block so the two
cannot drift apart on a flag rename.
check_bind returned early on any loopback host, so the installer happily
wrote a unit for --host 127.0.0.1 --public-url https://hub.example.net
with neither credential. The hub now refuses that configuration at
startup, which would turn a clean install into a unit that loads and
immediately exits.

Same predicate as the hub's, same reasoning: a public URL naming a
non-loopback host is the operator saying agents on other machines dial
this hub, whatever the socket is bound to.
Both halves of the new gate: the hub CLI (refusal, which credential is
named as missing, both credentials starting, --allow-insecure-bind, a
loopback public URL arming nothing, the CAUCUS_PUBLIC_URL path, and a
public URL with no parseable host still falling through to
validate_public_url) and check_bind in the installer.

test_an_invalid_public_url_refuses_at_startup now passes both credentials:
its URL names a non-loopback host, so without them the new gate answers
first and the test no longer exercises validate_public_url.
The TLS section told operators that binding to loopback behind a proxy
"sidesteps its own non-loopback-bind refusal entirely" and that the two
credentials were worth setting anyway. They are now required, for exactly
the reason that paragraph described. The flags table row for
--allow-insecure-bind names the public-url case too.
test_a_reaped_peer_is_revived_through_a_redeemed_ticket minted its watch
ticket on the wall clock and then reaped on a synthetic one derived from a
last_seen snapshot, at last_seen + client_ttl + 1.

reap_stale measures two different deadlines against the single now it is
handed: the idle window it reaps against, and the watch-ticket expiry it
prunes against. Mixing the two clocks made the test hold only while the
synthetic instant stayed within WATCH_TICKET_TTL of the real mint, and while
the live last_seen had not moved past the literal one second of slack in the
snapshot arithmetic.

Mint the ticket against the same reference instant the reap is derived from,
and place the reap at the midpoint of client_ttl and WATCH_TICKET_TTL, so the
relationship the test rests on is arithmetic on the two constants instead of
an agreement between a synthetic clock and the wall clock. Assertions and
production code are unchanged.
Takes main's ruff 0.16 pin, the explicit S310 rule, and both relocks
wholesale; this branch changed nothing under pyproject.toml or uv.lock.

Two conflicts, both resolved keeping each side's intent:

- src/caucus/hub_connector.py: the branch's new _agent_headers() helper
  sits next to main's PYI034 suppression on __aenter__.
- CHANGELOG.md: one [Unreleased] section holding both sets of bullets.
@obeone
obeone merged commit dfaedfe into main Sep 18, 2026
8 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant