Skip to content

[guardian-proxy] Forward only committee handoffs the chain stores - #1365

Merged
0xsiddharthks merged 8 commits into
mainfrom
siddharth/proxy-onchain-handoffs
Oct 10, 2026
Merged

0xsiddharthks merged 8 commits into
mainfrom
siddharth/proxy-onchain-handoffs

Conversation

@0xsiddharthks

@0xsiddharthks 0xsiddharthks commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor

Summary

A committee member can push a handoff to the guardian while its reconfig can still abort, which leaves the guardian stuck on a committee the chain never activated and stops withdrawals (IOP-792). The proxy now forwards a handoff only once the chain stores it, which end_reconfig does.

Changes

  • New node::handoffs::HandoffGate admits a transition only when the chain stores its handoff.
  • Forwarding runs the gate before UpdateCommittee and UpdateCommitteeChain, and refuses when Sui can't be read.
  • ChainMemberSource becomes ChainSource and also reads handoffs.
  • New guardian_proxy_handoff_refused_total{reason} metric.
  • The key-rotation e2e test runs the gate against localnet.

Nodes need no change. They already push only stored handoffs and retry a refused push, so a lagging proxy fullnode only delays the update.

The gate matches a handoff's two epochs, not its certificate bytes. An epoch only ever forms one committee, and the enclave still verifies the certificate.

@github-actions

github-actions Bot commented Oct 8, 2026 •

Copy link
Copy Markdown
PR Preview Action v1.8.1
Preview removed because the pull request was closed.
2026-10-10 07:46 UTC

@0xsiddharthks
0xsiddharthks marked this pull request as ready for review October 8, 2026 17:30
@0xsiddharthks
0xsiddharthks requested a review from bmwill as a code owner October 8, 2026 17:30
@0xsiddharthks
0xsiddharthks force-pushed the siddharth/proxy-onchain-handoffs branch 2 times, most recently from 5336d74 to e4a0aea Compare October 8, 2026 18:18
Comment thread crates/hashi-guardian-proxy/src/node/handoffs.rs
Comment thread crates/hashi-guardian-proxy/src/node/handoffs.rs Outdated
Comment thread crates/hashi-guardian-proxy/src/node/handoffs.rs Outdated
Comment thread design/docs/guardian.mdx
Comment thread crates/e2e-tests/src/lib.rs Outdated

@zhouwfang zhouwfang left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The gate only checks each hop's epochs, so a member can stuff a chain with thousands of fake hops that reuse a stored handoff's epochs. The enclave skips them as no-ops, but it still parses every signature first, which gets expensive. Could the gate also check each hop's certificate and committee against the chain?

@zhouwfang
zhouwfang dismissed their stale review October 8, 2026 20:59

out-of-scope hardening

@0xsiddharthks
0xsiddharthks marked this pull request as draft October 8, 2026 22:11
@0xsiddharthks
0xsiddharthks force-pushed the siddharth/proxy-onchain-handoffs branch 2 times, most recently from 962e469 to dfe780b Compare October 9, 2026 18:17
@0xsiddharthks
0xsiddharthks marked this pull request as ready for review October 9, 2026 18:47
@0xsiddharthks

Copy link
Copy Markdown
Contributor Author

The gate only checks each hop's epochs, so a member can stuff a chain with thousands of fake hops that reuse a stored handoff's epochs.

502b164 makes a chain consecutive again, so it can use each stored handoff's epochs once, not thousands of times. I'd dropped that rule with the cache in 212eb46. Matching each hop's certificate and committee against the chain isn't in this PR.

if known.is_some() {
return Ok(known);
}
let read = tokio::time::timeout(LOOKUP_TIMEOUT, self.source.next_epoch(from_epoch))

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Misses aren't cached, so every request with an unstored epoch costs two Sui reads on the same endpoint the allowlist uses. Could we bound misses the way RosterCache does?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

A shared miss budget would let one member starve another's update here. A roster re-read serves every signer, but a lookup is per epoch, so a member spamming unstored epochs would keep taking the budget an honest push needs. I'll bound it per member in a follow-up, since that needs the caller's key in the gate.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Makes sense, a per-member bound in a follow-up works for me.

let mut reached = None;
for transition in transitions {
let (from_epoch, to_epoch) = epochs(transition).ok_or(Refusal::Malformed)?;
// A handoff reaches a later epoch, so consecutive ones are distinct

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This limits hops per request, but a member can still resend the chain as often as they like, and the enclave parses every hop each time. Could we drop hops at or below the guardian's epoch before forwarding?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That needs the guardian's epoch in the proxy, kept in step with the enclave: GetGuardianInfo is cached for 30s and an update doesn't refresh it. The per-member limit from the other thread caps this too, since a member could then send one chain per interval.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It doesn't need to be in step. The guardian's epoch only goes up, so a stale value just skips fewer no-op hops. Also, resent chains are all cache hits, so the follow-up would need to limit every update per member, not just misses.

Comment thread crates/hashi-guardian-proxy/src/node/handoffs.rs Outdated

@zhouwfang zhouwfang left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approving with the per-member limit as a follow-up.

The two Docker builds failed on Docker Hub's pull rate limit, not on the change. It has cleared now, so could you re-run the failed jobs?

A committee member can push a handoff certificate to the guardian while
its reconfig is still pending. If that reconfig then aborts, the guardian
is stuck on a committee the chain never activated and every withdrawal
stops.

The proxy now forwards UpdateCommittee and UpdateCommitteeChain only for
handoffs the chain stores, which end_reconfig does: the CommitteeHandoff
out of a transition's signing epoch must name its target epoch. A failed
Sui read refuses the request.
Stored handoffs never change, so the gate remembers the ones it has read.
Replaying the chain's history then costs no Sui reads, and a request at
most one lookup that finds nothing. That bound replaces the consecutive
rule, which grew with every epoch.
A handoff the chain does not store yet is refused as Unavailable under
not_on_chain: its reconfig is pending, or the proxy's fullnode lags. One
the chain stored a different handoff over is refused as FailedPrecondition
under superseded, which no stale read can produce.
The key-rotation e2e test refused a handoff before any reconfig had
started. It now checks the refusal once the reconfig is pending, which is
the window the gate exists for.
The handoff key's type address is the Hashi type's only because both are
in the original package. The design doc now says the member and handoff
checks trust the proxy's Sui fullnode and assume nodes cannot reach the
enclave directly.
A chain could repeat one stored handoff any number of times, and the
enclave parses every transition's signature before it skips the ones
that change nothing. Each handoff must again leave the epoch the one
before it reached, so a chain holds a stored handoff at most once.
A separate connection keeps request-driven lookups off the allowlist
refresh's, but both still share the endpoint, so "can't hold up" claimed
too much.
admit and the module summaries read as if the whole handoff were compared
with the chain's. They now say a handoff is matched by its two epochs, and
the gate's module doc says its certificate and committee are not compared.
@0xsiddharthks
0xsiddharthks force-pushed the siddharth/proxy-onchain-handoffs branch from 532750d to f644e14 Compare October 10, 2026 02:00
@0xsiddharthks
0xsiddharthks merged commit 71aca1a into main Oct 10, 2026
13 checks passed
@0xsiddharthks
0xsiddharthks deleted the siddharth/proxy-onchain-handoffs branch October 10, 2026 07:46
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants