Skip to content

perf: emulated-curve scalar multiplication - #1826

Open
yelhousni wants to merge 14 commits into
masterfrom
perf/sm
Open

yelhousni wants to merge 14 commits into
masterfrom
perf/sm

Conversation

@yelhousni

@yelhousni yelhousni commented Sep 18, 2026 •

Copy link
Copy Markdown
Contributor

Description

Three constraint-count optimisations for emulated-curve scalar multiplication, plus a subgroup-membership soundness fix found during review.

1. Optimal comb window for fixed-base scalar multiplication

combWindow() previously returned a compile-time constant w=8 for all R1CS circuits regardless of the base field. The constraint count for one comb step is dominated by combSelect, whose minimum cost is:

selectCost(w) = min_rb [ oneHotCost(rb) + oneHotCost(w-rb) + 2^rb · 2·nbLimbs ]

Wider windows reduce the number of chain additions but increase per-window selector cost. The trade-off shifts with nbLimbs because the 2^rb · 2·nbLimbs term scales linearly with the limb count, making wider windows relatively cheaper for larger base fields.

combWindow() now calls combOptimalWindow(nbLimbs, n), which minimises the model

cost(w) = (nw−1)·selectCost(w) + selectCost(tw) + (nw−1)·100·nbLimbs

where 100·nbLimbs approximates the per-step chain-addition cost (fitted to measured data: ≈440 constraints/step for nbLimbs=4, ≈614 for nbLimbs=6). Results:

curve nbLimbs old w new w
secp256k1 G1, BN254 G1 4 8 9
BLS12-381 G1, BW6-761 G1 6 8 10

combPlonkWindow was also wrong. It was 5, documented as the measured SCS optimum, but w=6 is cheaper on every supported curve. It is now 6. SCS counts, w=4…8:

curve w=4 w=5 w=6 w=7 w=8
secp256k1 101,776 92,472 89,890 97,852 123,576
BN254 G1 101,538 91,162 89,652 97,614 120,402
BLS12-381 G1 148,187 132,226 129,875 141,905 176,699

This is a single constant shared by all curves, so the relevant claim is that one width is simultaneously optimal for all of them — which TestScalarMulBaseCombPlonkWindow now asserts, rather than assuming.

2. Tight subscalar bound for 4D Eisenstein (fake-GLV) decomposition

The 4D scalar decomposition via rationalReconstructExt / rationalReconstructExtG2 (gnark-crypto/algebra/lattice) decomposes a full-size scalar into four subscalars. The bound tightens from (BitLen+3)/4 + 2 to + 1, removing one iteration of the 4-way loop and saving one complete-addition step per scalar multiplication.

The hint reduces the rank-4 lattice

L = { (x,y,z,t) ∈ ℤ⁴ : x + λy ≡ k(z + λt) mod r },   det L = r

(L is the kernel of a surjection ℤ⁴ → ℤ/r, hence of index r) and returns a row with a nonzero denominator (z,t). gnark-crypto reduces at δ = 99/100, so the first reduced vector satisfies the LLL guarantee

‖b₁‖ ≤ (1/(δ − 1/4))^((n−1)/4) · (det L)^(1/n) = 1.3514^(3/4) · r^(1/4) = 1.2534 · r^(1/4)

and 1.2534 < 2, so every coordinate fits in ⌈BitLen/4⌉ + 1 bits.

Note this is the LLL approximation factor, not the Hermite constant. γ₄ is √2, and a Hermite bound would only bound λ₁(L) ≤ 2^(1/4)·r^(1/4), the length of the shortest vector — which LLL is not guaranteed to find.

Why the returned row is bounded, not just some short vector. The hint selects the minimum-infinity-norm row among those with (z,t) ≠ (0,0), so bounding b₁ only helps if b₁ is itself a candidate. It always is: a lattice vector with (z,t) = (0,0) satisfies x + λy ≡ 0 mod r, i.e. lies in the 2D GLV sublattice, whose minimum is ≈ √r. Measured by Gauss reduction: 2^128 for secp256k1 and BLS12-381, 2^127 for BN254 — far above ‖b₁‖ ≈ 2^64. The denominator condition therefore can never skip b₁.

+1 is minimal, not conservative. 1.2534·r^(1/4) ≈ 2^64.33 for a 256-bit r, so ⌈BitLen/4⌉ bits alone is not provable. The hint does try an early-termination path bounded by exactly r^(1/4) — which would fit ⌈BitLen/4⌉ — but it falls through to the general reduction for roughly 1 in 6000 random scalars, so the bound has to cover the fallback. This is why the 64-bit bound used by hand-built circuits that search for a Minkowski-optimal vector does not transfer to a deterministic single-shot LLL.

Applied to:

  • scalarMulGLVAndFakeGLV in sw_emulated/point.go (all j=0 emulated curves)
  • BLS12-381 G2 (sw_bls12381/g2.go)
  • BN254 G2 (sw_bn254/g2.go)

Not applied to BW6-761 G2: the cofactor clearing constant is 97 bits (log₂ ≈ 96.4), which requires nbits ≥ 97, i.e. the +2 formula. Not applied to 2D classic GLV (scalarMulGLV, jointScalarMulGLVUnsafe): ecc.SplitScalar (Babai rounding / HalfGCD) is empirically verified to produce subscalars up to BitLen>>1 + 2 bits for BW6-761.

3. Cheaper BLS12-381 G1 subgroup check: clear by |x-1|, not by the full cofactor

CofactorClearing for BLS12-381 G1 was the full cofactor h = (x-1)²/3 = 3·11²·10177²·859267²·52437899², a 126-bit constant, built by taking every prime power that exactly divides h. That rule is sufficient but not necessary, and here it over-charges by a factor of two.

Write n = (x-1)/3 = 11·10177·859267·52437899, so that h = 3n² and x-1 = 3n, with 3 ∤ n.

El Housni-Guillevic §3.2, Cor. 1 prove for every BLS curve that the full n-torsion is rational — E[n] ⊂ E(𝔽ₚ), i.e. there are n² points of order n and none of order n². So the n-part of the cofactor torsion is ℤ_n × ℤ_n: rank 2, with exponent n rather than n². The remaining lone factor 3 of h is exactly the 3 in 3n = x-1; it is not squared in h and contributes a cyclic ℤ_3, so no rank-2 claim is made about it.

The cofactor torsion is therefore ℤ_n × ℤ_{3n}, of order 3n² = h and exponent lcm(n, 3n) = 3n = x-1. This agrees with E(𝔽ₚ) ≅ ℤ_{(x-1)/3} × ℤ_{(x-1)·r} (Wahby-Boneh, §5). The exponent of the torsion subgroup is thus (x-1), not h.

That is exactly what the preimage binding needs:

  • Soundness — [x-1]E(𝔽ₚ) = G₁ exactly, so a point carrying cofactor torsion has no on-curve preimage under [x-1].
  • Completeness — gcd(x-1, r) = 1, so [x-1] is a bijection on G₁ and an honest point keeps a preimage.

CofactorClearing is now |x-1| = 0xd201000000010001 (64 bits, Hamming weight 7) — the same constant gnark-crypto's G1.ClearCofactor has always used. The CurveParams.CofactorClearing doc now states the actual requirement (divisibility by the torsion exponent, plus gcd(c, r) = 1) rather than the prime-power rule.

AssertIsOnG1 now uses the same binding. It previously ran Bowe's endomorphism test P = -[x²]ϕ(P) (eprint 2019/814) via scalarMulBySeedSquare, a 128-bit ladder. It now hints a preimage S and asserts [x-1]S == P, a 64-bit ladder. To avoid a second copy of the binding, assertPointInSubgroup is exposed as Curve.AssertIsInSubgroup and sw_bls12381 delegates to it. scalarMulBySeedSquare is retained — pairing.go still uses it.

128 bits is the floor for the endomorphism approach: G₁ has CM by ℤ[ω] only, so the lattice is rank 2 with determinant r and its shortest vector has norm ≈ √r ≈ 2¹²⁷·⁵. Bowe already sits at that bound, and Scott 2021/1130 §6 and Dai-Lin-Zhao-Zhou 2022/348 (Table 4, short vector (z², 1)) both restate the same test. The preimage route beats it only because a circuit has hints and native code does not.

The 64-bit ladder uses the incomplete group law, with a guard on every step (thanks @YaoJGalteland for the suggestion). assertedRatio pins a slope with λ·den − num ≡ 0. That is already unsatisfiable whenever den ≡ 0 with num ≢ 0, and leaves λ free only when den ≡ num ≡ 0. So the incomplete law fails closed everywhere except two spots, each closed by one emulated non-zero check:

  • tangent, a = 0: den = 2y and num = 3x² vanish together only at (0,0). That point is not on the curve, but AssertIsOnCurve accepts it as the infinity encoding, so a malicious preimage hint can supply it; a free slope there leaves the ladder output unconstrained, and intermediates are never re-checked against the curve equation. Rejected up front, and every doubling additionally asserts y ≠ 0.
  • chord: den = q.x − t.x and num = q.y − t.y vanish together iff q = t, i.e. adding a point to itself. Every addition asserts the two x-coordinates differ, which also excludes q = −t (whose sum is the unrepresentable infinity).

A guard can only reject, so this cannot cost soundness. Completeness: an honest preimage has order r, a 255-bit prime, while every intermediate width-4 NAF scalar stays in (0, 2^64), so no step doubles infinity or adds a point to its negative or to itself.

This is the subgroup-binding ladder only (AssertIsInSubgroup, and AssertIsOnG1 which calls it). BLS12-381 G1 ScalarMul still takes the classic-GLV path.

4. AssertIsInSubgroup no longer accepts off-subgroup points on cofactor curves

Exporting assertPointInSubgroup turned an internal shortcut into a public method that could silently assert nothing. It returns immediately when CofactorClearing == nil, and that nil meant two different things: "prime-order curve, nothing to check" (secp256k1, BN254 G1, P-256, P-384, STARK curve) and "cofactor curve that routes around the binding internally via PreferClassicGLV" (BW6-761 G1). Only the first justifies a no-op. A circuit calling AssertIsOnCurve followed by AssertIsInSubgroup on the on-curve, off-subgroup BW6-761 point (2, y) on y² = x³ − 1 — whose native IsInSubGroup() is false — solved.

CurveParams gains a PrimeOrder bool carrying the fact that actually licenses the no-op, and AssertIsInSubgroup now panics at circuit-definition time when neither that nor a clearing constant is present. The zero value fails closed, so a newly added cofactor curve is rejected until it supplies a membership check rather than silently accepting everything. PrimeOrder: true is set on the five genuinely cofactor-1 curves; BW6-761 G1 and G2 keep it false with a comment that the omission is deliberate.

No behaviour change on any internal path: every assertPointInSubgroup call site is already gated on !PreferClassicGLV.

Type of change

  • New feature (non-breaking change which adds functionality)
  • Bug fix (non-breaking change which fixes an issue) — AssertIsInSubgroup accepting off-subgroup points on BW6-761 G1, and combPlonkWindow being suboptimal

CurveParams gains an exported field (PrimeOrder). Anyone constructing CurveParams literals for a custom prime-order curve must set it, otherwise AssertIsInSubgroup panics instead of silently passing.

How has this been tested?

  • go test ./std/algebra/emulated/sw_emulated/...
  • go test ./std/algebra/emulated/sw_bls12381/...
  • go test ./std/algebra/emulated/sw_bn254/...
  • go test ./std/algebra/emulated/sw_bw6761/... (BW6-761 G2 unchanged)
  • go test ./internal/stats/ — the committed snapshot compares strictly and is unaffected
  • golangci-lint clean

New tests:

  • TestScalarMulBaseCombOptimalWindow — compiles the comb across w=7…12, takes the measured argmin and asserts combOptimalWindow agrees. Compile failures fail the test instead of being logged and skipped, and the argmin must be strictly interior to the sweep so it is not an artifact of where the sweep was cut. (Replaces TestScalarMulBaseCombConstraints, which only logged counts and would have passed with a losing window selected.)
  • TestScalarMulBaseCombPlonkWindow — same for SCS across w=4…8, asserted for all three curves at once.
  • TestCombWindowDispatch — covers the R1CS/PLONK branch in combWindow itself.
  • TestSubScalarBound — deterministic scalars (extremes, powers of two at the 64/65-bit boundary, λ, and a fixed LCG spread) asserted within nbits; the reconstruction relation checked to vanish mod r; the denominator checked nonzero; the 2D sublattice minimum checked to dwarf the bound.
  • TestBLS12381CofactorClearingConstant — extended to pin h = 3n², c = 3n, 3 ∤ n, the factorisation of n, and lcm(n, 3n) = c.
  • TestAssertIsInSubgroupBW6761FailsExplicitly — regression for §4, asserting the circuit fails and fails for the stated reason rather than as an incidental unsatisfied constraint.
  • TestAssertIsInSubgroupPrimeOrder, TestAssertIsInSubgroupBLS12381, TestCofactorCurvesDeclareAMembershipCheck — the other side of the guard, and the fail-closed invariant over all seven supported curves.
  • TestMulByConstantRejectsInfinity — compiles mulByConstant with nothing constraining its output, so a guard is the only thing that can reject; verified to fail with the guards removed. It covers the guard rather than an end-to-end forgery, since mounting one also requires replacing the tangent hint.

How has this been benchmarked?

Constraint counts from frontend.Compile on the BN254 scalar field (GetNbConstraints), Apple M5, 32 GB RAM.

Fixed-base scalar multiplication (ScalarMulBase)

R1CS:

circuit before (w=8) after (w=optimal) Δ
secp256k1 G1 16,634 16,248 (w=9) −386 (−2.3%)
BN254 G1 16,568 16,218 (w=9) −350 (−2.1%)
BLS12-381 G1 22,615 21,895 (w=10) −720 (−3.2%)

PLONK (combPlonkWindow 5 → 6):

circuit before (w=5) after (w=6) Δ
secp256k1 G1 92,472 89,890 −2,582 (−2.8%)
BN254 G1 91,162 89,652 −1,510 (−1.7%)
BLS12-381 G1 132,226 129,875 −2,351 (−1.8%)

Variable-base scalar multiplication (4D Eisenstein GLV, subscalar bound −1 bit)

circuit before after Δ
secp256k1 G1 ScalarMul 167,434 165,358 −2,076 (−1.2%)
BLS12-381 G2 ScalarMul 604,203 599,199 −5,004 (−0.8%)
BN254 G2 ScalarMul 452,817 449,427 −3,390 (−0.7%)

BLS12-381 G1 subgroup check

"Unified" is the 64-bit |x−1| ladder on unified formulas; "guarded incomplete" is the same ladder on the incomplete group law with the assertions above (62 doublings and 7 additions, each with an emulated non-zero check, plus the up-front (0,0) reject). All columns are compiles, not estimates.

circuit before unified |x−1| guarded incomplete Δ vs unified
subgroup check (R1CS) 178,084 80,021 62,163 −17,858 (−22.3%)
subgroup check (PLONK) 301,425 229,182 −72,243 (−24.0%)
on-curve plus subgroup (R1CS) 80,500 62,692 −17,808 (−22.1%)
on-curve plus subgroup (PLONK) 303,239 230,996 −72,243 (−23.8%)
AssertIsOnG1 (R1CS) 127,328 80,696 62,911 −17,785 (−22.0%)
AssertIsOnG1 (PLONK) 479,218 304,006 231,763 −72,243 (−23.8%)

Against the pre-PR baselines, the subgroup check is 178,084 → 62,163 R1CS (−115,921, −65.1%) and AssertIsOnG1 is 127,328 → 62,911 R1CS (−64,417, −50.6%) and 479,218 → 231,763 PLONK (−247,455, −51.6%).

For reference, the same ladder with no guards at all (compiled, not used, and unsound) is 51,756 / 195,880 for the subgroup check, so the three guards cost 10,407 R1CS and 33,302 PLONK.

Checklist:

  • I have performed a self-review of my code
  • I have commented my code, particularly in hard-to-understand areas
  • I have made corresponding changes to the documentation
  • I have added tests that prove my fix is effective or that my feature works
  • I did not modify files generated from templates
  • golangci-lint does not output errors locally
  • New and existing unit tests pass locally with my changes
  • Any dependent changes have been merged and published in downstream modules

Note

High Risk
Changes affect cryptographic constraint systems (subgroup binding, cofactor clearing, GLV bounds) and exported CurveParams; incorrect math would break soundness, though the PR adds extensive regression tests.

Overview
This PR reduces constraint cost for emulated-curve scalar multiplication and tightens subgroup-membership soundness in sw_emulated and related curve packages.

Performance: Fixed-base ScalarMulBase now picks an R1CS-optimal comb window via combOptimalWindow (e.g. w=9 for 4-limb curves, w=10 for BLS12-381 G1) instead of a fixed w=8; PLONK uses combPlonkWindow 6 (was 5). Fake-GLV subscalar bit bounds drop from (BitLen+3)/4 + 2 to + 1 on G1/G2 paths (with updated g2GenNbits constants). BLS12-381 G1 CofactorClearing is |x−1| (64 bits) instead of the full cofactor, and AssertIsOnG1 uses the shared [x−1] preimage check via Curve.AssertIsInSubgroup rather than a 128-bit endomorphism ladder; the clearing ladder uses guarded incomplete group law in mulByConstant.

Soundness / API: CurveParams gains PrimeOrder; AssertIsInSubgroup is exported and panics on cofactor curves without a clearing constant (e.g. BW6-761 G1), fixing silent no-ops. Benchmark stats in internal/stats/latest_stats.csv are refreshed; tests cover comb windows, subscalar bounds, cofactor arithmetic, and membership guards.

Reviewed by Cursor Bugbot for commit fd1b5bc. Bugbot is set up for automated code reviews on this repo. Configure here.

@yelhousni yelhousni changed the title Perf/sm perf: emulated-curve scalar multiplication Sep 18, 2026
@yelhousni yelhousni self-assigned this Sep 18, 2026
@yelhousni yelhousni added type: perf dep: linea Issues affecting Linea downstream labels Sep 18, 2026
@yelhousni
yelhousni requested a review from a team as a code owner September 24, 2026 16:03
@yelhousni
yelhousni requested a review from Tabaie September 24, 2026 16:31
Comment thread std/algebra/emulated/sw_emulated/point.go
Comment thread std/algebra/emulated/sw_emulated/point.go Outdated
Comment thread std/algebra/emulated/sw_emulated/params.go Outdated
Comment thread std/algebra/emulated/sw_emulated/fixedbase_test.go Outdated
Comment thread std/algebra/emulated/sw_bls12381/g1_test.go
@YaoJGalteland

Copy link
Copy Markdown

Have you tried the following guarded change?

In mulByConstant, drop the unified formulas and use the incomplete tangent and the incomplete chord, with a guard on every step:

  • reject (0,0) at the start of the ladder
  • assert y ≠ 0 before every doubling
  • assert the two x coordinates differ before every addition

The unified formulas are there to pin the slope when a denominator vanishes. At (0,0) with a = 0 the incomplete tangent λ·2y = 3x² is 0 = 0, so λ is free. The guards close that case. They are emulated non-zero checks, so they cost something, but much less than the unified formula on all 62 doublings and 7 additions.

frontend.Compile on the BN254 scalar field, GetNbConstraints (vibe-coded the perf as follows):

Circuit This PR (unified 64-bit) Proposed (guarded incomplete) Δ
Subgroup check, R1CS 80,021 62,163 −17,858 (−22.3%)
Subgroup check, PLONK 301,425 229,182 −72,243 (−24.0%)
On-curve plus subgroup, R1CS 80,500 62,692 −17,808 (−22.1%)
On-curve plus subgroup, PLONK 303,239 230,996 −72,243 (−23.8%)
AssertIsOnG1, R1CS 80,696 62,911 −17,785 (−22.0%)
AssertIsOnG1, PLONK 304,006 231,763 −72,243 (−23.8%)

AssertIsInSubgroup delegated to assertPointInSubgroup, which returned
immediately when CofactorClearing was nil. That nil meant "prime-order
curve, nothing to check" for secp256k1/BN254/P-256/P-384/STARK, but it
also held for BW6-761 G1, which has a nontrivial cofactor and simply
routes around the binding internally via PreferClassicGLV. Exporting the
method made that case reachable from outside, where it silently accepted
on-curve, off-subgroup points: AssertIsOnCurve + AssertIsInSubgroup
solved for (2, y) on y² = x³ - 1, a point whose native IsInSubGroup() is
false.

Add CurveParams.PrimeOrder to carry the fact that licensed the no-op,
and panic when it is absent and no clearing constant is available. The
zero value fails closed, so a newly added cofactor curve is rejected
until its membership check is supplied. Set PrimeOrder on the five
h = 1 curves; BW6-761 G1 and G2 keep it false with a note that the
omission is deliberate.

No behaviour change on any internal path: every assertPointInSubgroup
call site is already gated on !PreferClassicGLV.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@cursor

cursor Bot commented Oct 1, 2026

Copy link
Copy Markdown

Bugbot needs on-demand usage enabled

Bugbot uses usage-based billing for this team and requires on-demand usage to be enabled.

A team admin can enable on-demand usage in the Cursor dashboard.

yelhousni and others added 4 commits October 1, 2026 14:16
The comment claimed every prime dividing h occurs squared with rank-two
torsion, which the factorization directly above it contradicts: h =
3·11²·10177²·859267²·52437899² has a single factor of 3.

The rank-two claim belongs to n = (x-1)/3, not to h. ePrint 2021/1359
§3.2 Cor. 1 proves for every BLS curve that the full n-torsion is
rational — E[n] ⊂ E(Fp), n² points of order n and none of order n² — so
the n-part is Z_n × Z_n. Since h = 3n² with 3 ∤ n, the lone factor 3 is
exactly the 3 in 3n = x-1 and contributes a cyclic Z_3, about which no
rank-two claim is made. The cofactor torsion is Z_n × Z_{3n}, of order
3n² = h and exponent lcm(n, 3n) = 3n = x-1, which agrees with the
Wahby-Boneh decomposition already cited.

The constant |x-1| is unchanged; only its justification was wrong.
TestBLS12381CofactorClearingConstant now pins h = 3n², c = 3n, 3 ∤ n and
the resulting exponent, so the comment is checkable.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
TestScalarMulBaseCombConstraints logged constraint counts and asserted
nothing about the window. It also logged compile errors and continued, so
a width could drop out of the comparison silently. It would have passed
with combWindow picking a width that loses to another measured one.

Replace it with TestScalarMulBaseCombOptimalWindow, which compiles the
comb across a sweep, takes the measured argmin and asserts that
combOptimalWindow agrees with it. Compile failures now fail the test, and
the argmin must be interior to the sweep so it is not an artifact of
where the sweep was cut. R1CS confirms w = 9, 9, 10. Add
TestCombWindowDispatch to cover the R1CS/PLONK routing in combWindow
itself, which no test touched.

Doing the same for PLONK found a real bug: combPlonkWindow was 5 and
documented as the measured SCS optimum, but w = 6 is cheaper on every
supported curve:

	secp256k1   92472 -> 89890  (-2.8%)
	BN254 G1    91162 -> 89652  (-1.7%)
	BLS12-381  132226 -> 129875 (-1.8%)

Set it to 6 and record the w=4..8 table next to the constant. The old
test could not have caught this because it compiled PLONK only at w = 8,
a width combPlonkWindow never selects. The correctness sweeps now include
w = 6; internal/stats compares strictly and still passes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
sw_bls12381's randomCurvePoint was a verbatim copy of sw_emulated's
randomBLS12381CurvePoint. Two copies of a sampler invite the two tests to
drift apart in what they assume about it.

The two call sites do not want the same thing, so sharing one helper
across the package boundary (which would need a new test-only package)
is the wrong shape. sw_emulated's check is statistical — it asserts [c]P
lands in G1 over many random P — and keeps sampling. sw_bls12381 just
needs some off-subgroup points, so it now constructs them deterministically
with curvePointAtX for x in {4, 5, 6}. Each helper documents why it is
random or deterministic so they are not re-merged later.

Two defects fixed in passing:

  - the random loop guarded with `if p.IsInSubGroup() { continue }`,
    which skipped the assertion rather than making it; with deterministic
    points the off-subgroup property is asserted instead.
  - the sampler's doc claimed a uniformly random point of E(Fp) while
    always taking the canonical Sqrt root, i.e. one point out of each
    {P, -P} pair. Harmless for a subgroup test since G1 is closed under
    negation, but not what it said; it now picks the sign at random.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The bound nbits = (BitLen+3)/4 + 1 was attributed to a "LLL Hermite
bound, gamma_4 ~ 1.25". gamma_4 is sqrt(2), and a Hermite bound is about
lambda_1, the shortest vector, which LLL is not guaranteed to find — so
the justification did not support the constant even though the constant
is right.

1.25 turns out to be the LLL approximation factor, not a Hermite
constant. gnark-crypto reduces with delta = 99/100, giving

	||b_1|| <= (1/(delta - 1/4))^((n-1)/4) * (det L)^(1/n)
	        = 1.3514^(3/4) * r^(1/4) = 1.2534 * r^(1/4)

on L = {(x,y,z,t) : x + \lambda y = k(z + \lambda t) mod r}, which has
det r as the kernel of a surjection Z^4 -> Z/r. Since 1.2534 < 2, every
coordinate fits in ceil(BitLen/4) + 1 bits.

That still leaves the gap Ivo identified: the hint returns the
minimum-infinity-norm row with a nonzero denominator, not b_1. b_1 always
qualifies. A vector with (z,t) = (0,0) satisfies x + \lambda y = 0 mod r
and so lies in the 2D GLV sublattice, whose minimum is ~sqrt(r) ~ 2^127 —
three orders of magnitude above ||b_1|| ~ 2^64. The denominator condition
therefore cannot skip b_1, and the selected row is bounded by it.

+1 is minimal rather than conservative: 1.2534*r^(1/4) ~ 2^64.33 for a
256-bit r, so ceil(BitLen/4) alone is not provable. The hint does try an
early-termination path bounded by r^(1/4), but it falls through to the
general reduction for roughly 1 in 6000 random scalars, so the bound must
cover the fallback. This is why the 64-bit bound used by hand-built
circuits that search for a Minkowski-optimal vector does not transfer.

Add TestSubScalarBound: deterministic scalars (extremes, powers of two at
the 64/65-bit boundary, lambda, and a fixed LCG spread) asserted within
nbits, the reconstruction relation checked to vanish mod r, the
denominator checked nonzero, and the 2D sublattice minimum checked to
dwarf the bound. Fix the stale "+2" on the range-check line and the same
claim on the BN254 and BLS12-381 G2 copies.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@yelhousni

Copy link
Copy Markdown
Contributor Author

Thanks @ivokub, all were worth fixing. Pushed as four commits.

1. BW6-761 G1 no-op. You're right, and I reproduced it — (2, y) on y² = x³ - 1 is on-curve, off-subgroup, and used to satisfy the check. CofactorClearing == nil was doing double duty for "prime-order" and "cofactor curve that routes around the binding internally"; exporting the method made the second case reachable. Added CurveParams.PrimeOrder to carry the fact that licensed the no-op, and AssertIsInSubgroup now panics when neither that nor a clearing constant is present. The zero value fails closed, so a new cofactor curve is rejected until it supplies a check. BW6-761 G2 had the same latent shape. Regression test asserts it fails and fails for the stated reason.

2. Subscalar bound. You're right about γ₄, and right that an existence bound doesn't bound the selected row. The 1.25 turned out not to be a Hermite constant at all: gnark-crypto reduces at δ = 99/100, and the LLL guarantee ‖b₁‖ ≤ (1/(δ−¼))^((n−1)/4)·det^(1/4) gives exactly 1.2534·r^(1/4) on the rank-4 lattice of determinant r. Right number, wrong name.

For the row-selection gap: b₁ is always a candidate. A vector with (z,t) = (0,0) lies in the 2D GLV sublattice, whose minimum is ≈√r — measured 2^128 (secp256k1, BLS12-381) and 2^127 (BN254), against ‖b₁‖ ≈ 2^64. The denominator condition can't skip b₁, so the selected row is bounded by it.

+1 is minimal, not conservative: 1.2534·r^(1/4) ≈ 2^64.33 for 256-bit r. Worth noting the hint's early-termination path is bounded by r^(1/4), which would give 64 bits, but it falls through to the general reduction for ~1 in 6000 random scalars, so the bound has to cover the fallback. AddedTestSubScalarBound with deterministic vectors at the 64/65 boundary, and fixed the stale +2 at the range check and the same claim on both G2 copies.

3. 3-primary component. Correct, the sentence was wrong. The rank-2 claim belongs to n = (x-1)/3, not to h: by eprint 2021/1359 §3.2 Cor. 1 the full n-torsion is rational, so the n-part is ℤ_n × ℤ_n, while h = 3n² with 3 ∤ n and the lone 3 contributes a cyclic ℤ_3. Torsion is ℤ_n × ℤ_{3n}, exponent 3n = x-1. So the single factor of 3 is predicted by the shape rather than an exception to it. Constant unchanged; PR description updated too.

4. Comb window test. Agreed it wasn't a test. It now takes the measured argmin and asserts combOptimalWindow agrees, fails on compile errors, and requires the argmin to be interior to the sweep. R1CS confirms 9, 9, 10 as you measured.

Doing the same for PLONK found a real one: combPlonkWindow was 5 and documented as the measured SCS optimum, but w=6 is cheaper on every curve — secp256k1 92472→89890 (−2.8%), BN254 91162→89652, BLS12-381 132226→129875. Changed it to 6 with the table recorded. The old test couldn't have caught it: it compiled PLONK only at w=8, a width combPlonkWindow never selects.

5. Duplicated helper. Took the second option. The two call sites want different things — sw_emulated's check is statistical ([c]P ∈ G1 over many P) so it keeps sampling; sw_bls12381 just needs off-subgroup points, so it builds them deterministically now. Each helper says why it's random or deterministic. Two things fell out: the old loop's if p.IsInSubGroup() { continue } skipped the assertion rather than making it, and the sampler claimed uniformity while only ever returning the canonical Sqrt root.

yelhousni and others added 2 commits October 1, 2026 17:12
Per @YaoJGalteland's suggestion. The cofactor-clearing ladder ran on
unified formulas throughout, which is far more than it needs.

assertedRatio pins a slope with lambda*den - num == 0. That is
unsatisfiable whenever den == 0 with num != 0, and leaves lambda free
only when den == num == 0. So the incomplete group law already fails
closed everywhere except two spots:

  - tangent with a = 0: 2y == 0 and 3x^2 == 0, i.e. the point (0,0).
    Not on the curve, but AssertIsOnCurve accepts it as the infinity
    encoding, so a malicious preimage hint can supply it. A free slope
    there leaves the ladder output unconstrained, and intermediates are
    never re-checked against the curve equation.
  - chord: q.x == t.x and q.y == t.y, i.e. adding a point to itself.

Both are closed by one emulated non-zero check each: reject (0,0) up
front, assert y != 0 at every doubling, assert the x-coordinates differ
at every addition (which also excludes q = -t, whose sum is the
unrepresentable infinity). A guard can only reject, so this cannot cost
soundness; completeness holds because an honest preimage has order r, a
255-bit prime, while every intermediate NAF scalar stays in (0, 2^64).

Measured on the BN254 scalar field, matching Yao's estimates exactly:

	subgroup check          80021 -> 62163 R1CS, 301425 -> 229182 PLONK
	on-curve plus subgroup  80500 -> 62692 R1CS, 303239 -> 230996 PLONK
	AssertIsOnG1            80696 -> 62911 R1CS, 304006 -> 231763 PLONK

Add TestMulByConstantRejectsInfinity. It compiles mulByConstant with no
constraint on its output, so a guard is the only thing that can reject,
which makes it a real discriminator: with the guards the (0,0) witness
is unsatisfiable, with all three removed the honest tangent hint returns
lambda = 0 and the circuit solves. Verified by reverting. The test
documents that it covers the guard rather than an end-to-end forgery,
since mounting one also needs the tangent hint replaced.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@yelhousni

Copy link
Copy Markdown
Contributor Author

Have you tried the following guarded change?

In mulByConstant, drop the unified formulas and use the incomplete tangent and the incomplete chord, with a guard on every step:

  • reject (0,0) at the start of the ladder
  • assert y ≠ 0 before every doubling
  • assert the two x coordinates differ before every addition

The unified formulas are there to pin the slope when a denominator vanishes. At (0,0) with a = 0 the incomplete tangent λ·2y = 3x² is 0 = 0, so λ is free. The guards close that case. They are emulated non-zero checks, so they cost something, but much less than the unified formula on all 62 doublings and 7 additions.

frontend.Compile on the BN254 scalar field, GetNbConstraints (vibe-coded the perf as follows):

Circuit This PR (unified 64-bit) Proposed (guarded incomplete) Δ
Subgroup check, R1CS 80,021 62,163 −17,858 (−22.3%)
Subgroup check, PLONK 301,425 229,182 −72,243 (−24.0%)
On-curve plus subgroup, R1CS 80,500 62,692 −17,808 (−22.1%)
On-curve plus subgroup, PLONK 303,239 230,996 −72,243 (−23.8%)
AssertIsOnG1, R1CS 80,696 62,911 −17,785 (−22.0%)
AssertIsOnG1, PLONK 304,006 231,763 −72,243 (−23.8%)

Thanks @YaoJGalteland! implemented, and your numbers were exact.

One refinement the implementation made clear: assertedRatio pins λ with λ·den − num ≡ 0, which is already unsatisfiable when den ≡ 0 and num ≢ 0. So the free-slope case is only where both vanish — the two you named, (0,0) for the tangent at a = 0 and adding a point to itself for the chord. Three non-zero checks are enough.

Added TestMulByConstantRejectsInfinity, which constrains nothing but the guard so it actually discriminates; verified it fails with the guards removed.

@yelhousni
yelhousni requested a review from ivokub October 1, 2026 21:23
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

dep: linea Issues affecting Linea downstream type: perf

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants