Conversation
|
Have you tried the following guarded change? In
The unified formulas are there to pin the slope when a denominator vanishes. At
|
AssertIsInSubgroup delegated to assertPointInSubgroup, which returned immediately when CofactorClearing was nil. That nil meant "prime-order curve, nothing to check" for secp256k1/BN254/P-256/P-384/STARK, but it also held for BW6-761 G1, which has a nontrivial cofactor and simply routes around the binding internally via PreferClassicGLV. Exporting the method made that case reachable from outside, where it silently accepted on-curve, off-subgroup points: AssertIsOnCurve + AssertIsInSubgroup solved for (2, y) on y² = x³ - 1, a point whose native IsInSubGroup() is false. Add CurveParams.PrimeOrder to carry the fact that licensed the no-op, and panic when it is absent and no clearing constant is available. The zero value fails closed, so a newly added cofactor curve is rejected until its membership check is supplied. Set PrimeOrder on the five h = 1 curves; BW6-761 G1 and G2 keep it false with a note that the omission is deliberate. No behaviour change on any internal path: every assertPointInSubgroup call site is already gated on !PreferClassicGLV. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Bugbot needs on-demand usage enabledBugbot uses usage-based billing for this team and requires on-demand usage to be enabled. A team admin can enable on-demand usage in the Cursor dashboard. |
The comment claimed every prime dividing h occurs squared with rank-two
torsion, which the factorization directly above it contradicts: h =
3·11²·10177²·859267²·52437899² has a single factor of 3.
The rank-two claim belongs to n = (x-1)/3, not to h. ePrint 2021/1359
§3.2 Cor. 1 proves for every BLS curve that the full n-torsion is
rational — E[n] ⊂ E(Fp), n² points of order n and none of order n² — so
the n-part is Z_n × Z_n. Since h = 3n² with 3 ∤ n, the lone factor 3 is
exactly the 3 in 3n = x-1 and contributes a cyclic Z_3, about which no
rank-two claim is made. The cofactor torsion is Z_n × Z_{3n}, of order
3n² = h and exponent lcm(n, 3n) = 3n = x-1, which agrees with the
Wahby-Boneh decomposition already cited.
The constant |x-1| is unchanged; only its justification was wrong.
TestBLS12381CofactorClearingConstant now pins h = 3n², c = 3n, 3 ∤ n and
the resulting exponent, so the comment is checkable.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
TestScalarMulBaseCombConstraints logged constraint counts and asserted nothing about the window. It also logged compile errors and continued, so a width could drop out of the comparison silently. It would have passed with combWindow picking a width that loses to another measured one. Replace it with TestScalarMulBaseCombOptimalWindow, which compiles the comb across a sweep, takes the measured argmin and asserts that combOptimalWindow agrees with it. Compile failures now fail the test, and the argmin must be interior to the sweep so it is not an artifact of where the sweep was cut. R1CS confirms w = 9, 9, 10. Add TestCombWindowDispatch to cover the R1CS/PLONK routing in combWindow itself, which no test touched. Doing the same for PLONK found a real bug: combPlonkWindow was 5 and documented as the measured SCS optimum, but w = 6 is cheaper on every supported curve: secp256k1 92472 -> 89890 (-2.8%) BN254 G1 91162 -> 89652 (-1.7%) BLS12-381 132226 -> 129875 (-1.8%) Set it to 6 and record the w=4..8 table next to the constant. The old test could not have caught this because it compiled PLONK only at w = 8, a width combPlonkWindow never selects. The correctness sweeps now include w = 6; internal/stats compares strictly and still passes. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
sw_bls12381's randomCurvePoint was a verbatim copy of sw_emulated's
randomBLS12381CurvePoint. Two copies of a sampler invite the two tests to
drift apart in what they assume about it.
The two call sites do not want the same thing, so sharing one helper
across the package boundary (which would need a new test-only package)
is the wrong shape. sw_emulated's check is statistical — it asserts [c]P
lands in G1 over many random P — and keeps sampling. sw_bls12381 just
needs some off-subgroup points, so it now constructs them deterministically
with curvePointAtX for x in {4, 5, 6}. Each helper documents why it is
random or deterministic so they are not re-merged later.
Two defects fixed in passing:
- the random loop guarded with `if p.IsInSubGroup() { continue }`,
which skipped the assertion rather than making it; with deterministic
points the off-subgroup property is asserted instead.
- the sampler's doc claimed a uniformly random point of E(Fp) while
always taking the canonical Sqrt root, i.e. one point out of each
{P, -P} pair. Harmless for a subgroup test since G1 is closed under
negation, but not what it said; it now picks the sign at random.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The bound nbits = (BitLen+3)/4 + 1 was attributed to a "LLL Hermite
bound, gamma_4 ~ 1.25". gamma_4 is sqrt(2), and a Hermite bound is about
lambda_1, the shortest vector, which LLL is not guaranteed to find — so
the justification did not support the constant even though the constant
is right.
1.25 turns out to be the LLL approximation factor, not a Hermite
constant. gnark-crypto reduces with delta = 99/100, giving
||b_1|| <= (1/(delta - 1/4))^((n-1)/4) * (det L)^(1/n)
= 1.3514^(3/4) * r^(1/4) = 1.2534 * r^(1/4)
on L = {(x,y,z,t) : x + \lambda y = k(z + \lambda t) mod r}, which has
det r as the kernel of a surjection Z^4 -> Z/r. Since 1.2534 < 2, every
coordinate fits in ceil(BitLen/4) + 1 bits.
That still leaves the gap Ivo identified: the hint returns the
minimum-infinity-norm row with a nonzero denominator, not b_1. b_1 always
qualifies. A vector with (z,t) = (0,0) satisfies x + \lambda y = 0 mod r
and so lies in the 2D GLV sublattice, whose minimum is ~sqrt(r) ~ 2^127 —
three orders of magnitude above ||b_1|| ~ 2^64. The denominator condition
therefore cannot skip b_1, and the selected row is bounded by it.
+1 is minimal rather than conservative: 1.2534*r^(1/4) ~ 2^64.33 for a
256-bit r, so ceil(BitLen/4) alone is not provable. The hint does try an
early-termination path bounded by r^(1/4), but it falls through to the
general reduction for roughly 1 in 6000 random scalars, so the bound must
cover the fallback. This is why the 64-bit bound used by hand-built
circuits that search for a Minkowski-optimal vector does not transfer.
Add TestSubScalarBound: deterministic scalars (extremes, powers of two at
the 64/65-bit boundary, lambda, and a fixed LCG spread) asserted within
nbits, the reconstruction relation checked to vanish mod r, the
denominator checked nonzero, and the 2D sublattice minimum checked to
dwarf the bound. Fix the stale "+2" on the range-check line and the same
claim on the BN254 and BLS12-381 G2 copies.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
Thanks @ivokub, all were worth fixing. Pushed as four commits. 1. BW6-761 G1 no-op. You're right, and I reproduced it — 2. Subscalar bound. You're right about γ₄, and right that an existence bound doesn't bound the selected row. The 1.25 turned out not to be a Hermite constant at all: gnark-crypto reduces at δ = 99/100, and the LLL guarantee For the row-selection gap: b₁ is always a candidate. A vector with +1 is minimal, not conservative: 1.2534·r^(1/4) ≈ 2^64.33 for 256-bit r. Worth noting the hint's early-termination path is bounded by r^(1/4), which would give 64 bits, but it falls through to the general reduction for ~1 in 6000 random scalars, so the bound has to cover the fallback. Added 3. 3-primary component. Correct, the sentence was wrong. The rank-2 claim belongs to 4. Comb window test. Agreed it wasn't a test. It now takes the measured argmin and asserts Doing the same for PLONK found a real one: 5. Duplicated helper. Took the second option. The two call sites want different things — sw_emulated's check is statistical ([c]P ∈ G1 over many P) so it keeps sampling; sw_bls12381 just needs off-subgroup points, so it builds them deterministically now. Each helper says why it's random or deterministic. Two things fell out: the old loop's |
Per @YaoJGalteland's suggestion. The cofactor-clearing ladder ran on unified formulas throughout, which is far more than it needs. assertedRatio pins a slope with lambda*den - num == 0. That is unsatisfiable whenever den == 0 with num != 0, and leaves lambda free only when den == num == 0. So the incomplete group law already fails closed everywhere except two spots: - tangent with a = 0: 2y == 0 and 3x^2 == 0, i.e. the point (0,0). Not on the curve, but AssertIsOnCurve accepts it as the infinity encoding, so a malicious preimage hint can supply it. A free slope there leaves the ladder output unconstrained, and intermediates are never re-checked against the curve equation. - chord: q.x == t.x and q.y == t.y, i.e. adding a point to itself. Both are closed by one emulated non-zero check each: reject (0,0) up front, assert y != 0 at every doubling, assert the x-coordinates differ at every addition (which also excludes q = -t, whose sum is the unrepresentable infinity). A guard can only reject, so this cannot cost soundness; completeness holds because an honest preimage has order r, a 255-bit prime, while every intermediate NAF scalar stays in (0, 2^64). Measured on the BN254 scalar field, matching Yao's estimates exactly: subgroup check 80021 -> 62163 R1CS, 301425 -> 229182 PLONK on-curve plus subgroup 80500 -> 62692 R1CS, 303239 -> 230996 PLONK AssertIsOnG1 80696 -> 62911 R1CS, 304006 -> 231763 PLONK Add TestMulByConstantRejectsInfinity. It compiles mulByConstant with no constraint on its output, so a guard is the only thing that can reject, which makes it a real discriminator: with the guards the (0,0) witness is unsatisfiable, with all three removed the honest tangent hint returns lambda = 0 and the circuit solves. Verified by reverting. The test documents that it covers the guard rather than an end-to-end forgery, since mounting one also needs the tangent hint replaced. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Thanks @YaoJGalteland! implemented, and your numbers were exact. One refinement the implementation made clear: Added |
Description
Three constraint-count optimisations for emulated-curve scalar multiplication, plus a subgroup-membership soundness fix found during review.
1. Optimal comb window for fixed-base scalar multiplication
combWindow()previously returned a compile-time constantw=8for all R1CS circuits regardless of the base field. The constraint count for one comb step is dominated bycombSelect, whose minimum cost is:Wider windows reduce the number of chain additions but increase per-window selector cost. The trade-off shifts with
nbLimbsbecause the2^rb · 2·nbLimbsterm scales linearly with the limb count, making wider windows relatively cheaper for larger base fields.combWindow()now callscombOptimalWindow(nbLimbs, n), which minimises the modelwhere
100·nbLimbsapproximates the per-step chain-addition cost (fitted to measured data: ≈440 constraints/step for nbLimbs=4, ≈614 for nbLimbs=6). Results:combPlonkWindowwas also wrong. It was5, documented as the measured SCS optimum, butw=6is cheaper on every supported curve. It is now6. SCS counts,w=4…8:This is a single constant shared by all curves, so the relevant claim is that one width is simultaneously optimal for all of them — which
TestScalarMulBaseCombPlonkWindownow asserts, rather than assuming.2. Tight subscalar bound for 4D Eisenstein (fake-GLV) decomposition
The 4D scalar decomposition via
rationalReconstructExt/rationalReconstructExtG2(gnark-crypto/algebra/lattice) decomposes a full-size scalar into four subscalars. The bound tightens from(BitLen+3)/4 + 2to+ 1, removing one iteration of the 4-way loop and saving one complete-addition step per scalar multiplication.The hint reduces the rank-4 lattice
(
Lis the kernel of a surjectionℤ⁴ → ℤ/r, hence of indexr) and returns a row with a nonzero denominator(z,t). gnark-crypto reduces atδ = 99/100, so the first reduced vector satisfies the LLL guaranteeand
1.2534 < 2, so every coordinate fits in⌈BitLen/4⌉ + 1bits.Note this is the LLL approximation factor, not the Hermite constant.
γ₄is√2, and a Hermite bound would only boundλ₁(L) ≤ 2^(1/4)·r^(1/4), the length of the shortest vector — which LLL is not guaranteed to find.Why the returned row is bounded, not just some short vector. The hint selects the minimum-infinity-norm row among those with
(z,t) ≠ (0,0), so boundingb₁only helps ifb₁is itself a candidate. It always is: a lattice vector with(z,t) = (0,0)satisfiesx + λy ≡ 0 mod r, i.e. lies in the 2D GLV sublattice, whose minimum is≈ √r. Measured by Gauss reduction:2^128for secp256k1 and BLS12-381,2^127for BN254 — far above‖b₁‖ ≈ 2^64. The denominator condition therefore can never skipb₁.+1is minimal, not conservative.1.2534·r^(1/4) ≈ 2^64.33for a 256-bitr, so⌈BitLen/4⌉bits alone is not provable. The hint does try an early-termination path bounded by exactlyr^(1/4)— which would fit⌈BitLen/4⌉— but it falls through to the general reduction for roughly 1 in 6000 random scalars, so the bound has to cover the fallback. This is why the 64-bit bound used by hand-built circuits that search for a Minkowski-optimal vector does not transfer to a deterministic single-shot LLL.Applied to:
scalarMulGLVAndFakeGLVinsw_emulated/point.go(all j=0 emulated curves)sw_bls12381/g2.go)sw_bn254/g2.go)Not applied to BW6-761 G2: the cofactor clearing constant is 97 bits (log₂ ≈ 96.4), which requires
nbits ≥ 97, i.e. the+2formula. Not applied to 2D classic GLV (scalarMulGLV,jointScalarMulGLVUnsafe):ecc.SplitScalar(Babai rounding / HalfGCD) is empirically verified to produce subscalars up toBitLen>>1 + 2bits for BW6-761.3. Cheaper BLS12-381 G1 subgroup check: clear by
|x-1|, not by the full cofactorCofactorClearingfor BLS12-381 G1 was the full cofactorh = (x-1)²/3 = 3·11²·10177²·859267²·52437899², a 126-bit constant, built by taking every prime power that exactly dividesh. That rule is sufficient but not necessary, and here it over-charges by a factor of two.Write
n = (x-1)/3 = 11·10177·859267·52437899, so thath = 3n²andx-1 = 3n, with3 ∤ n.El Housni-Guillevic §3.2, Cor. 1 prove for every BLS curve that the full
n-torsion is rational —E[n] ⊂ E(𝔽ₚ), i.e. there aren²points of ordernand none of ordern². So then-part of the cofactor torsion isℤ_n × ℤ_n: rank 2, with exponentnrather thann². The remaining lone factor3ofhis exactly the3in3n = x-1; it is not squared inhand contributes a cyclicℤ_3, so no rank-2 claim is made about it.The cofactor torsion is therefore
ℤ_n × ℤ_{3n}, of order3n² = hand exponentlcm(n, 3n) = 3n = x-1. This agrees withE(𝔽ₚ) ≅ ℤ_{(x-1)/3} × ℤ_{(x-1)·r}(Wahby-Boneh, §5). The exponent of the torsion subgroup is thus(x-1), noth.That is exactly what the preimage binding needs:
[x-1]E(𝔽ₚ) = G₁exactly, so a point carrying cofactor torsion has no on-curve preimage under[x-1].gcd(x-1, r) = 1, so[x-1]is a bijection onG₁and an honest point keeps a preimage.CofactorClearingis now|x-1| = 0xd201000000010001(64 bits, Hamming weight 7) — the same constantgnark-crypto'sG1.ClearCofactorhas always used. TheCurveParams.CofactorClearingdoc now states the actual requirement (divisibility by the torsion exponent, plusgcd(c, r) = 1) rather than the prime-power rule.AssertIsOnG1now uses the same binding. It previously ran Bowe's endomorphism testP = -[x²]ϕ(P)(eprint 2019/814) viascalarMulBySeedSquare, a 128-bit ladder. It now hints a preimageSand asserts[x-1]S == P, a 64-bit ladder. To avoid a second copy of the binding,assertPointInSubgroupis exposed asCurve.AssertIsInSubgroupandsw_bls12381delegates to it.scalarMulBySeedSquareis retained —pairing.gostill uses it.128 bits is the floor for the endomorphism approach:
G₁has CM byℤ[ω]only, so the lattice is rank 2 with determinantrand its shortest vector has norm ≈√r≈ 2¹²⁷·⁵. Bowe already sits at that bound, and Scott 2021/1130 §6 and Dai-Lin-Zhao-Zhou 2022/348 (Table 4, short vector(z², 1)) both restate the same test. The preimage route beats it only because a circuit has hints and native code does not.The 64-bit ladder uses the incomplete group law, with a guard on every step (thanks @YaoJGalteland for the suggestion).
assertedRatiopins a slope withλ·den − num ≡ 0. That is already unsatisfiable wheneverden ≡ 0withnum ≢ 0, and leavesλfree only whenden ≡ num ≡ 0. So the incomplete law fails closed everywhere except two spots, each closed by one emulated non-zero check:a = 0:den = 2yandnum = 3x²vanish together only at(0,0). That point is not on the curve, butAssertIsOnCurveaccepts it as the infinity encoding, so a malicious preimage hint can supply it; a free slope there leaves the ladder output unconstrained, and intermediates are never re-checked against the curve equation. Rejected up front, and every doubling additionally assertsy ≠ 0.den = q.x − t.xandnum = q.y − t.yvanish together iffq = t, i.e. adding a point to itself. Every addition asserts the two x-coordinates differ, which also excludesq = −t(whose sum is the unrepresentable infinity).A guard can only reject, so this cannot cost soundness. Completeness: an honest preimage has order
r, a 255-bit prime, while every intermediate width-4 NAF scalar stays in(0, 2^64), so no step doubles infinity or adds a point to its negative or to itself.This is the subgroup-binding ladder only (
AssertIsInSubgroup, andAssertIsOnG1which calls it). BLS12-381 G1ScalarMulstill takes the classic-GLV path.4.
AssertIsInSubgroupno longer accepts off-subgroup points on cofactor curvesExporting
assertPointInSubgroupturned an internal shortcut into a public method that could silently assert nothing. It returns immediately whenCofactorClearing == nil, and thatnilmeant two different things: "prime-order curve, nothing to check" (secp256k1, BN254 G1, P-256, P-384, STARK curve) and "cofactor curve that routes around the binding internally viaPreferClassicGLV" (BW6-761 G1). Only the first justifies a no-op. A circuit callingAssertIsOnCurvefollowed byAssertIsInSubgroupon the on-curve, off-subgroup BW6-761 point(2, y)ony² = x³ − 1— whose nativeIsInSubGroup()is false — solved.CurveParamsgains aPrimeOrder boolcarrying the fact that actually licenses the no-op, andAssertIsInSubgroupnow panics at circuit-definition time when neither that nor a clearing constant is present. The zero value fails closed, so a newly added cofactor curve is rejected until it supplies a membership check rather than silently accepting everything.PrimeOrder: trueis set on the five genuinely cofactor-1 curves; BW6-761 G1 and G2 keep it false with a comment that the omission is deliberate.No behaviour change on any internal path: every
assertPointInSubgroupcall site is already gated on!PreferClassicGLV.Type of change
AssertIsInSubgroupaccepting off-subgroup points on BW6-761 G1, andcombPlonkWindowbeing suboptimalCurveParamsgains an exported field (PrimeOrder). Anyone constructingCurveParamsliterals for a custom prime-order curve must set it, otherwiseAssertIsInSubgrouppanics instead of silently passing.How has this been tested?
go test ./std/algebra/emulated/sw_emulated/...go test ./std/algebra/emulated/sw_bls12381/...go test ./std/algebra/emulated/sw_bn254/...go test ./std/algebra/emulated/sw_bw6761/...(BW6-761 G2 unchanged)go test ./internal/stats/— the committed snapshot compares strictly and is unaffectedgolangci-lintcleanNew tests:
TestScalarMulBaseCombOptimalWindow— compiles the comb acrossw=7…12, takes the measured argmin and assertscombOptimalWindowagrees. Compile failures fail the test instead of being logged and skipped, and the argmin must be strictly interior to the sweep so it is not an artifact of where the sweep was cut. (ReplacesTestScalarMulBaseCombConstraints, which only logged counts and would have passed with a losing window selected.)TestScalarMulBaseCombPlonkWindow— same for SCS acrossw=4…8, asserted for all three curves at once.TestCombWindowDispatch— covers the R1CS/PLONK branch incombWindowitself.TestSubScalarBound— deterministic scalars (extremes, powers of two at the 64/65-bit boundary,λ, and a fixed LCG spread) asserted withinnbits; the reconstruction relation checked to vanish modr; the denominator checked nonzero; the 2D sublattice minimum checked to dwarf the bound.TestBLS12381CofactorClearingConstant— extended to pinh = 3n²,c = 3n,3 ∤ n, the factorisation ofn, andlcm(n, 3n) = c.TestAssertIsInSubgroupBW6761FailsExplicitly— regression for §4, asserting the circuit fails and fails for the stated reason rather than as an incidental unsatisfied constraint.TestAssertIsInSubgroupPrimeOrder,TestAssertIsInSubgroupBLS12381,TestCofactorCurvesDeclareAMembershipCheck— the other side of the guard, and the fail-closed invariant over all seven supported curves.TestMulByConstantRejectsInfinity— compilesmulByConstantwith nothing constraining its output, so a guard is the only thing that can reject; verified to fail with the guards removed. It covers the guard rather than an end-to-end forgery, since mounting one also requires replacing the tangent hint.How has this been benchmarked?
Constraint counts from
frontend.Compileon the BN254 scalar field (GetNbConstraints), Apple M5, 32 GB RAM.Fixed-base scalar multiplication (
ScalarMulBase)R1CS:
PLONK (
combPlonkWindow5 → 6):Variable-base scalar multiplication (4D Eisenstein GLV, subscalar bound −1 bit)
ScalarMulScalarMulScalarMulBLS12-381 G1 subgroup check
"Unified" is the 64-bit
|x−1|ladder on unified formulas; "guarded incomplete" is the same ladder on the incomplete group law with the assertions above (62 doublings and 7 additions, each with an emulated non-zero check, plus the up-front(0,0)reject). All columns are compiles, not estimates.|x−1|AssertIsOnG1(R1CS)AssertIsOnG1(PLONK)Against the pre-PR baselines, the subgroup check is 178,084 → 62,163 R1CS (−115,921, −65.1%) and
AssertIsOnG1is 127,328 → 62,911 R1CS (−64,417, −50.6%) and 479,218 → 231,763 PLONK (−247,455, −51.6%).For reference, the same ladder with no guards at all (compiled, not used, and unsound) is 51,756 / 195,880 for the subgroup check, so the three guards cost 10,407 R1CS and 33,302 PLONK.
Checklist:
golangci-lintdoes not output errors locallyNote
High Risk
Changes affect cryptographic constraint systems (subgroup binding, cofactor clearing, GLV bounds) and exported
CurveParams; incorrect math would break soundness, though the PR adds extensive regression tests.Overview
This PR reduces constraint cost for emulated-curve scalar multiplication and tightens subgroup-membership soundness in
sw_emulatedand related curve packages.Performance: Fixed-base
ScalarMulBasenow picks an R1CS-optimal comb window viacombOptimalWindow(e.g. w=9 for 4-limb curves, w=10 for BLS12-381 G1) instead of a fixed w=8; PLONK usescombPlonkWindow6 (was 5). Fake-GLV subscalar bit bounds drop from(BitLen+3)/4 + 2to+ 1on G1/G2 paths (with updatedg2GenNbitsconstants). BLS12-381 G1CofactorClearingis |x−1| (64 bits) instead of the full cofactor, andAssertIsOnG1uses the shared[x−1]preimage check viaCurve.AssertIsInSubgrouprather than a 128-bit endomorphism ladder; the clearing ladder uses guarded incomplete group law inmulByConstant.Soundness / API:
CurveParamsgainsPrimeOrder;AssertIsInSubgroupis exported and panics on cofactor curves without a clearing constant (e.g. BW6-761 G1), fixing silent no-ops. Benchmark stats ininternal/stats/latest_stats.csvare refreshed; tests cover comb windows, subscalar bounds, cofactor arithmetic, and membership guards.Reviewed by Cursor Bugbot for commit fd1b5bc. Bugbot is set up for automated code reviews on this repo. Configure here.