Skip to content

hydra: interleaved merge split; merge exactness and bound tests - #198

Merged
zzylol merged 3 commits into
mainfrom
hydra/merge-split
Oct 11, 2026
Merged

zzylol merged 3 commits into
mainfrom
hydra/merge-split

Conversation

@zzylol

@zzylol zzylol commented Oct 10, 2026 •

Copy link
Copy Markdown
Collaborator

What changed

  • approxbench --merge-split contiguous|interleaved (default contiguous, today's behavior for every row). Interleaved deals the stream round-robin by record, so each shard is a sample of the whole stream; shard sizes equal the contiguous split's (record j goes to the next shard with room).
  • Records carry merge_split next to merge_shards (BenchSection, MergeMetrics, merged_record.schema.json regenerated). Absent on older records = contiguous, so existing results parse; --flat refuses two different splits in one row, as it does two shard counts. MergeSplit lives in aqpbm-core and clap parses it directly.
  • The interleaved copy is built only when a merge is requested.
  • Plumbing: Requirement.merge_split → scored_row, which hands every merge body (merge, merge-step, merge-accuracy) the stream reordered by wrappers::interleave, so the existing contiguous partition cut yields the round-robin shards. No wrapper signature changes; every row with a merge gets it.
  • Documented in README (flag list next to the Hydra merge example; states the round-robin exactly: it runs over the contiguous shard sizes, so with a short last shard shard 0 is not exactly 0, k, 2k, ...) and docs/saturation_study.md.

Why

Spec PR 4 for the Hydra §6.3 study (hydra doc #193 §2.2–2.3): merges across ingest workers split by record, not by time range.

Tests (wrappers::tests, wrappers::hydra_shared::tests, 3.7 s debug)

  • interleave_deals_round_robin_at_the_contiguous_sizes.
  • merge_split_on_two_squares_must_agree (flatten: absent = contiguous; interleaved + absent is refused).
  • hydra-cms / hydra-cs / hydra-hll: a 3×16 grid folded from 4 and 7 shards, under both splits, answers every probed subpopulation (each first label and each pair; every occurring (group, value) for CMS/CS) exactly as the single-stream grid.
  • §2.2 bounds on a fixed 20k-record, 2-label stream, W=16 (so groups collide), cells 3×512, β=8: every scored group's error is inside its bound. CMS: f ≤ f̂ ≤ f + ε_c·N_q + β(F_v + ε_c·M)/W, ε_c = e/512. CS: [f − ε_s(L2_q + L2_C), f + ε_s·L2_q + βF_v/W + ε_s·L2_C], ε_s = 1/√512, L2_C = √(β·F2/W) with F2 = Σ_v F_v². HLL: [(1−ε_h)D_q, (1+ε_h)(D_q + β·ΣD/W)], ε_h = 3·1.04/√2^14.

Mutation checks (each fails the named test): interleave returning the stream unchanged; each hydra merge skipping a shard; β = 0 (all three bound tests); inner ε = 0 (CMS, CS); CMS answer 1 under (never-under assert). Inner ε_h = 0 does not fail HLL: at these distinct counts HLL is near-exact. CS margin: passes from β ≈ 4.8 (fails at 4.7), so β = 8 leaves under 2x; cell_cols 1024 would need β ≈ 14 (ε shrinks, the colliders' error does not), so cells stay 3×512. Removing the absent→contiguous default fails the flatten test.

Evidence

cargo fmt --check, cargo clippy --workspace --all-targets -D warnings, cargo test --workspace clean. Release CLI, --merge-shards 8 --metrics accuracy: hydra-cms (hydra_columns.yaml, 3×128, 3×512) and hydra-hll merge_accuracy equal query_accuracy under both splits (are_all 2.5929 / 1.3396); kll-percall --library oxide (deterministic; repeated runs identical) --config k=50 --dataset zipf --size 100000 --zipf-s 1.1 --cardinality 100000 --merge-shards 8 --operations merge --metrics accuracy: merge mean_rank_err 0.012426 contiguous vs 0.012599 interleaved, so the split reaches lossy merges too. Same contiguous command on 653d6a8 vs this head: the record is identical except the new "merge_split": "contiguous" (timings excluded).

🤖 Generated with Claude Code

https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

zzylol and others added 2 commits October 10, 2026 04:09
`--merge-split contiguous|interleaved` (default contiguous, unchanged) picks
how merge runs split the stream into shards. Interleaved reorders the stream
so the existing contiguous `partition` deals it round-robin by record, at the
contiguous split's shard sizes, so every row's folds take it unchanged.

Tests: hydra-cms, hydra-cs and hydra-hll grids folded from 4 and 7 shards,
under both splits, answer every probed subpopulation exactly as the single
stream; and every scored group on a fixed 20k-record stream lands inside its
hydra doc §2.2 bound (CMS also never under).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis
- `merge_split` (contiguous|interleaved) is a record field next to
  `merge_shards` (BenchSection and MergeMetrics; schema regenerated).
  Merge records stamp it; a record without it split contiguously, so
  existing results parse, and flatten refuses two different splits in
  one row as it does two shard counts.
- `MergeSplit` moves to aqpbm-core (it is a record type) and is parsed by
  clap directly instead of matching strings.
- The interleaved copy of the stream is built only when a merge runs.
- README states the round-robin exactly (contiguous shard sizes, so a
  short last shard shifts the pattern); MergeMetrics docs name the split.
- hydra-cs bound test: note its margin (passes from beta ~4.8).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis
@zzylol
zzylol merged commit be74be4 into main Oct 11, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant