Skip to content

Arm the credentials guard on a non-loopback --public-url - #100

Merged
obeone merged 5 commits into
feat/remote-hubfrom
fix/public-url-credentials-guard
Sep 18, 2026
Merged

obeone merged 5 commits into
feat/remote-hubfrom
fix/public-url-credentials-guard

Conversation

@obeone

@obeone obeone commented Sep 18, 2026

Copy link
Copy Markdown
Owner

A hub bound to 127.0.0.1 and exposed through a Cloudflare Tunnel, ngrok or a
reverse proxy on the same machine starts with no credentials and says nothing
about it. The startup guard on this branch only looks at the bind address, so
caucus-hub --host 127.0.0.1 --public-url https://hub.example.net walks
straight past it. With no --agent-key, agent_ok() returns True for
everyone, which leaves /register, /peers and /watch-ticket/redeem open to
whoever reaches the tunnel. With no --operator-token, role_for grades every
browser that loads /ui as operator.

The gap predates this branch, but it is the same gate, and docs/remote-hub.md
was recommending the shape: its TLS section said a loopback bind "sidesteps its
own non-loopback-bind refusal entirely", then suggested setting both credentials
anyway.

What changed

_mount_mcp_http already read a declared public URL as remote when deciding
what watch_command hands an agent
(remote=public_url is not None or not is_loopback_host(host)). The credentials
gate now agrees with it: a --public-url whose host is not loopback demands both
--operator-token and --agent-key, whatever the socket is bound to.

The call worth arguing about is whether declaring a public URL should arm the
guard on its own. It only arms when the URL names a non-loopback host, so
--public-url http://localhost:8765 stays a free cosmetic choice and anyone
using it that way sees no change. A URL naming a routable host is the operator
saying agents on other machines dial this hub, which is exactly the exposure the
guard exists for. --allow-insecure-bind covers this case too, same flag, same
meaning.

check_bind in setup_service.py gets the same predicate. Without it the
installer would happily write a unit the hub then refuses to start.

The refusal, in the style of _insecure_bind_message:

refusing to start: --public-url https://hub.example.net says agents on other
machines dial this hub, and by default neither of its two doors is
locked. The bind is loopback, so the socket is not the exposure --
whatever sits in front of it is, and anything reaching that front
reaches both of these:

  agent door     /register and /mcp - join the room, read everything
                 said in it
                 --agent-key KEY  (env CAUCUS_AGENT_KEY): already set
  operator door  /ui - pause, stop, kick, read the whole transcript
                 --operator-token TOKEN  (env CAUCUS_OPERATOR_TOKEN): MISSING

Both are required once the hub is advertised off-box. Generate and set them:

  export CAUCUS_AGENT_KEY="$(openssl rand -hex 24)"
  export CAUCUS_OPERATOR_TOKEN="$(openssl rand -hex 24)"
  caucus-hub --host 127.0.0.1 --public-url https://hub.example.net

Not what you meant? A loopback public URL (http://localhost:8765) is
just a nicer address for this machine and needs none of this.
Meant it, behind something that already authenticates? --allow-insecure-bind
starts anyway, with both doors open.

The doors inventory and the openssl lines that both messages share moved into
_doors_block, so a flag rename cannot fix one message and miss the other.

How to test

# refused
caucus-hub --host 127.0.0.1 --public-url https://hub.example.net

# starts
caucus-hub --host 127.0.0.1 --public-url https://hub.example.net \
  --operator-token "$(openssl rand -hex 24)" --agent-key "$(openssl rand -hex 24)"

# starts, unchanged
caucus-hub --host 127.0.0.1 --public-url http://localhost:8765

11 new tests, 7 on the hub CLI in tests/test_remote_hub.py and 4 on
check_bind in tests/test_setup_service.py. They cover the refusal, which
credential the message reports as missing, both credentials starting,
--allow-insecure-bind, a loopback public URL arming nothing, the
CAUCUS_PUBLIC_URL path, and a public URL with no parseable host still falling
through to validate_public_url.

pytest: 1079 passed, 0 failed. mypy src/: clean. ruff check src/: 9
findings, byte-identical to the ones already on this branch (checked against a
pristine git archive of feat/remote-hub), none of them in the changed lines.

One existing test moved: test_an_invalid_public_url_refuses_at_startup now
passes both credentials, because its URL names a non-loopback host and the new
gate answers before validate_public_url gets to reject the path. That ordering
is deliberate. Missing credentials on an advertised hub outrank the shape of the
address being advertised.

Risk

Breaking for one configuration: a hub started with a non-loopback --public-url
and fewer than both credentials now exits 2 instead of starting. That is the
point, but it will bite anyone running the reverse proxy example from
docs/remote-hub.md as it was written, and a service installed by
caucus-setup-service in that shape stops loading. --allow-insecure-bind is
the one flag back.

The startup guard only ever looked at the bind address, so a hub on
127.0.0.1 exposed through a Cloudflare Tunnel, ngrok or a local reverse
proxy started without a word: no agent key, so agent_ok() passed every
caller, and no operator token, so every browser was graded as operator.
/register, /peers and /watch-ticket/redeem were open to whoever reached
the tunnel.

_mount_mcp_http already read a declared public URL as "this hub is
remote" when deciding what watch_command hands an agent. The credentials
gate now reads it the same way: a --public-url naming a non-loopback host
demands both --operator-token and --agent-key, with --allow-insecure-bind
as the same escape hatch the bind guard offers. A loopback public URL is
just a nicer address for this machine and arms nothing.

_insecure_public_url_message carries the refusal; the doors inventory and
the openssl snippet both messages share move into _doors_block so the two
cannot drift apart on a flag rename.
check_bind returned early on any loopback host, so the installer happily
wrote a unit for --host 127.0.0.1 --public-url https://hub.example.net
with neither credential. The hub now refuses that configuration at
startup, which would turn a clean install into a unit that loads and
immediately exits.

Same predicate as the hub's, same reasoning: a public URL naming a
non-loopback host is the operator saying agents on other machines dial
this hub, whatever the socket is bound to.
Both halves of the new gate: the hub CLI (refusal, which credential is
named as missing, both credentials starting, --allow-insecure-bind, a
loopback public URL arming nothing, the CAUCUS_PUBLIC_URL path, and a
public URL with no parseable host still falling through to
validate_public_url) and check_bind in the installer.

test_an_invalid_public_url_refuses_at_startup now passes both credentials:
its URL names a non-loopback host, so without them the new gate answers
first and the test no longer exercises validate_public_url.
The TLS section told operators that binding to loopback behind a proxy
"sidesteps its own non-loopback-bind refusal entirely" and that the two
credentials were worth setting anyway. They are now required, for exactly
the reason that paragraph described. The flags table row for
--allow-insecure-bind names the public-url case too.
@obeone
obeone force-pushed the fix/public-url-credentials-guard branch from 70e73b3 to 26ae438 Compare September 18, 2026 02:21
@obeone
obeone merged commit 225724c into feat/remote-hub Sep 18, 2026
5 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant