Repository navigation
feat(hub): make a remote caucus hub work, and shut the doors it left open - #101
Merged
Merged
Conversation
/register was unauthenticated and /mcp had no auth at all, so anything that could reach the port could join the caucus and read the room. Add --agent-key (env CAUCUS_AGENT_KEY): when set, both doors demand Authorization: Bearer <key>. /register checks it before the per-host throttle, so a bogus key never spends another caller's budget. /mcp is gated once at the HTTP layer rather than per tool, inside the CORS layer so a preflight OPTIONS (which carries no Authorization by definition) is still answered. The DNS-rebinding Host/Origin allowlist is untouched. The key is independent of --operator-token/--observer-token, which keep guarding only the console. With no key set, nothing changes. A non-loopback bind without one now warns at startup.
HubConnector gains an agent_key argument defaulting to the environment variable, so caucus-claude-agent inherits it with no wiring; the bridge reads the same variable at module level. Both present it on /register only: that is the one call made before the process holds a peer token, and every later call already spends that token as its own bearer. caucus-watch needs nothing: it never registers, it is handed a peer token.
The plugin connects straight to /mcp, so it needs the key too. Claude
Code expands ${VAR:-default} in an http server's headers, and an unset
variable with an empty default becomes an empty string, so the same
entry serves a keyed remote hub and an open local one.
Open behaviour unchanged with no key; correct key accepted on /register and /mcp; wrong, missing and malformed keys refused with 401; the refusal precedes the register throttle; the CORS preflight still passes; and the in-process HubConnector tool path still works behind the gate.
validate_public_url is the server-side counterpart to validate_hub_url: it checks the origin the hub advertises to agents is http or https, names a host, and carries nothing past the origin, since the hub appends its own paths to that value.
The watcher token file lives on the hub's filesystem, which is not the agent's when the hub serves other machines, so its path names nothing the agent can run. build_mcp_server now takes a remote flag; when set, watch_command returns the CAUCUS_TOKEN env form caucus-watch already accepts and writes no file. The loopback path is untouched.
--allowed-host (env CAUCUS_ALLOWED_HOSTS) adds entries to the /mcp DNS-rebinding guard, which previously only ever learned the bind address: a hub on 0.0.0.0 reached as hub.lan:8765 was refused with no way to allow it. A bare host is allowed on the hub's own port, host:port is kept as given, and the guard itself is unchanged. --public-url (env CAUCUS_PUBLIC_URL) is the base URL other machines reach the hub at. It replaces the 127.0.0.1 a wildcard bind used to advertise, so watch_command hands out a runnable command, and its netloc joins the Host allowlist. A non-loopback bind now refuses to start unless both an operator token and an agent key are set; --allow-insecure-bind is the escape hatch. The refusal names both doors, both flags, both env vars, the two flags needed to make the hub reachable, and the way back to loopback.
check_bind asked only for the operator token, which never guarded /register or /mcp: the installer would happily write a unit whose agent door was open to the whole network. It now demands both credentials, matching the gate caucus-hub applies at startup, and --agent-key carries the key into the launchd plist and the systemd env file alongside the dashboard tokens. README's one-line description of that gate follows.
allowed-host parsing (bare, host:port, bracketed IPv6, env, merge), public-url accept/reject cases, watch_command in both deployments, the _mount_mcp_http wiring, and the startup refusal driven through main() with uvicorn stubbed. test_setup_service follows check_bind's new signature and covers the agent key's plumbing.
Document running a hub on one machine with agents joining from others: the /mcp and caucus-bridge connection paths end to end, the flags/env reference, a Caddy TLS example, and the failure modes an operator actually hits. Also corrects ARCHITECTURE.md's outdated claim that /ui carries no authentication now that AuthConfig exists, and notes the agent-key middleware guarding /register and /mcp.
Three modules spelled "loopback" as the exact-string frozenset
{127.0.0.1, localhost, ::1} while urlguard did a real ipaddress check,
so --host 127.0.0.2 was refused as non-loopback by the hub's bind gate
and by the service installer while urlguard called it loopback, and
--host [::1] likewise. Promote urlguard's check to is_loopback_host()
and make it the one definition every call site reads.
Add needs_remote_optin() alongside it: validate_hub_url's question
asked without consulting this process's environment, so the hub can
tell whether a URL will be refused on somebody else's machine.
/peers, /channels and /forms answer before a caller holds a peer token, so the only credential they can carry is the shared agent key. The bridge and the connector sent it on /register alone, and the in-process connector behind the MCP tools read it from the environment, which --agent-key on the command line never sets. Send it on all three read calls and take it from the live AuthConfig in-process, so the hub can start gating them without cutting off its own clients. No behaviour change against a hub with no key configured. Rename the bridge's _register_headers to _agent_headers, which is now what it is.
With --public-url http://hub.lan:8765, watch_command handed a remote agent `caucus-watch --hub http://hub.lan:8765`. caucus-watch runs validate_hub_url on that value and exits 2 before its first poll, since plain http to a non-loopback host needs CAUCUS_ALLOW_REMOTE_HUB. The agent backgrounds the process and believes a watcher is listening, while nothing is. Prefix the command with the opt-in when the advertised URL is that shape. Decided with needs_remote_optin rather than validate_hub_url, so the answer does not depend on the hub process's own environment: the command runs in the agent's.
Three fixes at the HTTP layer, all of them about a hub bound wider than loopback with an agent key set: /peers, /channels and /forms were unauthenticated, so anyone who could reach the port still read the peer roster, every peer's self-reported status, every private channel's name, topic and membership, and the text of every pending operator form. Gate all three on the agent key, accepting an operator or observer token too so the console keeps working. The role check is guarded on AuthConfig.enabled, because role_for grades every caller as operator when no operator token is set. /ping stays open as the liveness probe, with its docstring now saying what it does disclose. secrets.compare_digest raises TypeError on a str holding anything outside ASCII, and ASGI decodes headers as latin-1, so `Authorization: Bearer \xe9` was an unhandled 500 rather than a 401 -- and on /register the gate runs before the rate-limit bucket, so nothing throttled a caller repeating it. Compare UTF-8 bytes in all three credential comparisons. An OPTIONS on /mcp skipped the key gate and reached the SDK, which built a session transport and a task group before answering 405. A genuine CORS preflight is answered by the CORS layer wrapping this gate, so an OPTIONS arriving here is not one: gate it like any other method.
--host 0.0.0.0 is rewritten to 127.0.0.1 for the address the hub advertises, but the mount is still marked remote, so a remote agent got a watcher command pointing back at its own machine. Fix it at the source: fold --public-url into the existing non-loopback bind gate as a third required setting when the bind host is a wildcard, keeping the per-credential MISSING / already set shape and the --allow-insecure-bind escape hatch. A concrete bind such as --host 192.168.1.10 already advertises itself, so nothing is demanded there. Also in this pass: - Use is_loopback_host everywhere the bind gate, the /mcp default and the launcher precondition used to spell loopback as three exact strings. --host 127.0.0.2 and --host ::1 now behave as the loopback addresses they are, so they skip the gate. - Bracket a bare IPv6 --allowed-host: `::1` and `2001:db8::1` contain a colon, so they were passed through verbatim and matched no Host header the guard will ever see, while the operator believed otherwise. - De-duplicate the allowlist: the bind address is no longer appended a second time when the operator also named it with --allowed-host. - Normalise blank credentials to None. `--agent-key ""` made agent_ok reject everyone and `--operator-token ""` turned auth on with a token no first frame can match; the clients already normalise this way. - Warn once at startup when the advertised URL is plain http to a non-loopback host: peer tokens and message content cross the network in clear, and the watcher command has to carry the opt-in to run.
check_bind demanded an agent key for a non-loopback install, then wrote a unit passing only --host/--port/--no-browser: an operator following the refusal installed a hub with /mcp off (the non-loopback default) advertising an address nothing off-box can dial. Plumb --public-url, --allowed-host and --mcp-http through the same route --agent-key takes -- validated, then into the plist's EnvironmentVariables and the systemd environment file, both now built from one service_environment() so neither can drift -- and show them in the plan. Extend check_bind to demand --public-url on a wildcard bind, matching the gate caucus-hub applies at startup, so the refusal comes while installing rather than from a unit that loads and then exits. Loopback is decided by the shared is_loopback_host, so the installer and the hub no longer disagree about 127.0.0.2.
Fold the public-url/allowed-host, agent-key and non-loopback bind refusal refinements into their existing entries, and add the new Fixed and Security items from this round of review.
A remote agent cannot be handed a token file (the path names nothing on its machine), and handing it the peer token puts the room bearer for /receive, /send, /ack, /channels/*, /ask and /floor into the agent's transcript, its shell history and the watcher's environ. The token file exists precisely to keep it out of argv and out of the launching transcript. Add the store the remote path needs instead: issue_watch_ticket mints an opaque claim check with a 120s TTL, redeem_watch_ticket spends it exactly once and returns the peer token. Expired entries are pruned on issue, on redeem and in the idle-reaper sweep, so the store is bounded by the tickets issued inside one TTL window. Redemption compares with secrets.compare_digest on UTF-8 bytes, like the hub's other credentials. A ticket resurrects nothing: it yields a token, and a token the hub has since forgotten is refused by client_for exactly as it is today.
Two holes in the same wall, both only reachable once a hub leaves loopback. POST /watch-ticket/redeem exchanges a single-use watch ticket for the peer token it stands for. It is gated on the shared agent key the same way /register is, so a keyed hub does not hand tokens to whoever can reach the port. An unknown, already spent or expired ticket answers 404, not 401: the caller's credential was fine, and conflating the two would have the watcher report a dead session when all it needs is a fresh watch_command(). Neither the ticket nor the token is ever logged. /ping returned far more than liveness: whether a named peer exists, how long since it last touched the hub, whether a listener is attached, and the peer's own set_status prose. On a keyed non-loopback hub that was an unauthenticated disclosure, so it now goes through _require_agent_read like /peers, /channels and /forms. The connector and the stdio bridge present the key on their ping calls accordingly; with no key configured the endpoint stays open, as before. PROTOCOL_TEXT is unchanged (it never named the credential form), so PROTOCOL_VERSION is not bumped.
--ticket and CAUCUS_TICKET join the credential chain. When a ticket is supplied and no direct token is, the watcher spends it once against POST /watch-ticket/redeem at startup, presenting CAUCUS_AGENT_KEY the way hub_connector already reads it, then polls exactly as before. Precedence, flags before environment as the module docstring has always promised: --token > --token-file > --ticket > CAUCUS_TOKEN > CAUCUS_TICKET. The historical order between the three token forms is untouched. A refused or unredeemable ticket exits 1 after printing the remedy to stdout, not a traceback: the agent is woken by this process exiting and reads what it left there, so the line says the ticket is single-use and short-lived and that watch_command() issues another.
The remote branch of watch_command emitted CAUCUS_TOKEN=<peer token> inline, which is the room bearer for /receive, /send, /ack, /channels/*, /ask and /floor, and the agent key gates none of those. It landed in the agent's transcript, its shell history and the watcher's environ, undoing exactly what the loopback token file was built to prevent. Emit caucus-watch --hub <url> --ticket <ticket> instead, keeping the CAUCUS_ALLOW_REMOTE_HUB=1 prefix logic. The token now appears nowhere in the tool result. The loopback branch and its 0600 token file are untouched.
…ed /ping Ticket store and endpoint: redeems once and the replay 404s, an expired ticket fails with the clock driven rather than slept on, an unknown ticket fails, the idle reaper sweeps the store, redemption demands the agent key when one is configured (and the refusal fires before the lookup, so the ticket survives it), and a ticket for a dead peer hands back a token client_for still refuses. watch_command: the remote result carries the ticket and the peer token appears nowhere in the payload, not merely outside the command string; the loopback 0600 token file is unchanged. Watcher: the five-step credential precedence, and redemption against a live hub including the keyed and unreachable cases. The four _resolve_token cases were rewritten onto _resolve_credential, which replaces it. /ping: open on an unkeyed hub, 401 without a credential on a keyed one, served with the agent key, and the connector carries that key on the call.
…reat model Brings docs/remote-hub.md and docs/ARCHITECTURE.md in line with what landed on the branch since the guide was written: watch_command() on a remote hub now hands out a single-use ticket instead of the peer token, the bind gate also requires --public-url on a wildcard host, and /peers, /channels, /forms and /ping are gated on the agent key. Adds a threat-model section spelling out what the agent key does not buy.
watch_command() minted a fresh single-use ticket on every call without retiring the one it replaced, so N refresh calls left N live bearer credentials outstanding for the same peer token, each already sitting in the agent's transcript for the full 120s TTL. Track the current ticket on the session's membership and revoke it (new HubState.revoke_watch_ticket) before minting the next one, on leave(), and in the dead-session sweep, so at most one ticket per member is ever redeemable.
issue_watch_ticket()'s docstring claimed a reaped peer can never be resurrected through a stale ticket, but client_for() revives from the reaped graveyard inside reaped_grace regardless of how the token arrived, so a ticket redeemed in that window resurrects the peer same as a direct token would. Correct the docstring, and cover the grace-window path with a test (the existing dead-peer test only exercises a hard unregister()). watch.py's docstring also still said --token was the only credential form landed in argv, but --ticket is passed in argv too. State that plainly, what it exposes to another local uid via ps, and why the exposure stays bounded (single-use, short TTL) without moving it to stdin.
The startup guard only ever looked at the bind address, so a hub on 127.0.0.1 exposed through a Cloudflare Tunnel, ngrok or a local reverse proxy started without a word: no agent key, so agent_ok() passed every caller, and no operator token, so every browser was graded as operator. /register, /peers and /watch-ticket/redeem were open to whoever reached the tunnel. _mount_mcp_http already read a declared public URL as "this hub is remote" when deciding what watch_command hands an agent. The credentials gate now reads it the same way: a --public-url naming a non-loopback host demands both --operator-token and --agent-key, with --allow-insecure-bind as the same escape hatch the bind guard offers. A loopback public URL is just a nicer address for this machine and arms nothing. _insecure_public_url_message carries the refusal; the doors inventory and the openssl snippet both messages share move into _doors_block so the two cannot drift apart on a flag rename.
check_bind returned early on any loopback host, so the installer happily wrote a unit for --host 127.0.0.1 --public-url https://hub.example.net with neither credential. The hub now refuses that configuration at startup, which would turn a clean install into a unit that loads and immediately exits. Same predicate as the hub's, same reasoning: a public URL naming a non-loopback host is the operator saying agents on other machines dial this hub, whatever the socket is bound to.
Both halves of the new gate: the hub CLI (refusal, which credential is named as missing, both credentials starting, --allow-insecure-bind, a loopback public URL arming nothing, the CAUCUS_PUBLIC_URL path, and a public URL with no parseable host still falling through to validate_public_url) and check_bind in the installer. test_an_invalid_public_url_refuses_at_startup now passes both credentials: its URL names a non-loopback host, so without them the new gate answers first and the test no longer exercises validate_public_url.
The TLS section told operators that binding to loopback behind a proxy "sidesteps its own non-loopback-bind refusal entirely" and that the two credentials were worth setting anyway. They are now required, for exactly the reason that paragraph described. The flags table row for --allow-insecure-bind names the public-url case too.
test_a_reaped_peer_is_revived_through_a_redeemed_ticket minted its watch ticket on the wall clock and then reaped on a synthetic one derived from a last_seen snapshot, at last_seen + client_ttl + 1. reap_stale measures two different deadlines against the single now it is handed: the idle window it reaps against, and the watch-ticket expiry it prunes against. Mixing the two clocks made the test hold only while the synthetic instant stayed within WATCH_TICKET_TTL of the real mint, and while the live last_seen had not moved past the literal one second of slack in the snapshot arithmetic. Mint the ticket against the same reference instant the reap is derived from, and place the reap at the midpoint of client_ttl and WATCH_TICKET_TTL, so the relationship the test rests on is arithmetic on the two constants instead of an agreement between a synthetic clock and the wall clock. Assertions and production code are unchanged.
…' into feat/remote-hub
Takes main's ruff 0.16 pin, the explicit S310 rule, and both relocks wholesale; this branch changed nothing under pyproject.toml or uv.lock. Two conflicts, both resolved keeping each side's intent: - src/caucus/hub_connector.py: the branch's new _agent_headers() helper sits next to main's PYI034 suppression on __aenter__. - CHANGELOG.md: one [Unreleased] section holding both sets of bullets.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What changed
A hub bound anywhere but loopback was an open room.
POST /registerwas unauthenticated and/mcphad no auth at all, so anything that could reach the port joined the caucus and read everything said in it.caucus-hubaccepted--host 0.0.0.0silently whilecaucus-setup-servicerefused it, the/mcpDNS-rebinding guard only ever learned the bind address (so a remote client'sHostheader was rejected with no way to allow it), and several pieces a remote setup needs already existed but were documented nowhere:CAUCUS_ALLOW_REMOTE_HUB,--allowed-origin.This branch makes a multi-machine caucus actually work, and says plainly what it does and does not protect.
--agent-key/CAUCUS_AGENT_KEY) on/registerand on the whole/mcpmount, independent of the existing operator and observer console tokens.--allowed-hostand--public-url, so a remote client is not rejected by the host guard and a remote agent is not handed a127.0.0.1URL it cannot reach.--public-urlon a wildcard bind.--allow-insecure-bindis the escape hatch. The refusal message is written for a human who just got refused: it names both flags, both env vars, and the way back to loopback./peers,/channels,/formsand/pingare gated too. They were handing out the peer roster, every private channel's name, topic and members, the text of every pending operator form, and each peer's ownset_statusstring.watch_commandon a remote hub hands out a single-use 120 second ticket, redeemed atPOST /watch-ticket/redeem, instead of printing the peer token into the agent's transcript. That token is the room bearer for/receive,/send,/ack,/channels/*,/askand/floor, and the agent key gates none of those.docs/remote-hub.md: a walkthrough for both connection paths, a flags and env reference, a Caddy TLS example (the hub has no TLS of its own), troubleshooting, and a section on what the key does not buy.Loopback deployments are unchanged. With no agent key configured, nothing in here fires.
Breaking change
Anyone binding
caucus-hubto a non-loopback address today must now set--agent-keyand--operator-token(and--public-urlon0.0.0.0or::), or pass--allow-insecure-bind.What the key does not buy
This is in the doc for the same reason it is here: the key is a perimeter, not authorization. It turns "anyone who can reach the port" into "anyone holding one shared secret", and inside the room it grants nothing extra and restricts nothing. Any keyholder can self-join any private channel with no invite check, take the broadcast floor and block every other sender's
say(), and take over a peer name whenever that peer's listener is not currently polling/receive. There is no per-agent identity, and no revocation short of restarting the hub with a new key.Two more limits worth having in the open. The hub still has no native TLS and expects a reverse proxy in front of anything crossing a network you do not already trust. And the
/registerrate-limit bucket is keyed onrequest.client.hostwith noX-Forwarded-Forhandling, so behind a proxy every agent shares one bucket: a registration burst can trip it for the whole fleet, not just the noisy one.How I verified it
1072 tests pass,
mypy src/is clean under strict, andruff check src/sits at its pre-existing 9 errors, all unrelated to this branch and being fixed separately. No version bump, the version comes from git tags in this repo.One thing I want the reviewer to see rather than find.
tests/test_remote_hub.py::test_a_reaped_peer_is_revived_through_a_redeemed_ticketfailed once on a full-suite run; the 151 runs after that were green, and I never got it to fail again. Reading the test, it ran two clocks at once: the ticket was minted against the wall clock while the reap was driven off a synthetic instant, so the ordering it depends on only held while the suite was fast. 4d75bfc drives both off one reference instant, placed at the midpoint ofclient_ttlandWATCH_TICKET_TTL, so the relationship is arithmetic on the constants instead of a race with the wall clock. The original traceback is lost, so that diagnosis is inspection, not a captured failure. If someone sees it go red again, I would like to know.Checklist
## [Unreleased]inCHANGELOG.md.PROTOCOL_TEXTchanged inhub.py,PROTOCOL_VERSIONwas bumped too. (PROTOCOL_TEXTis untouched on this branch.)