Repository navigation
hydra: interleaved merge split; merge exactness and bound tests - #198
Merged
Merged
Conversation
`--merge-split contiguous|interleaved` (default contiguous, unchanged) picks how merge runs split the stream into shards. Interleaved reorders the stream so the existing contiguous `partition` deals it round-robin by record, at the contiguous split's shard sizes, so every row's folds take it unchanged. Tests: hydra-cms, hydra-cs and hydra-hll grids folded from 4 and 7 shards, under both splits, answer every probed subpopulation exactly as the single stream; and every scored group on a fixed 20k-record stream lands inside its hydra doc §2.2 bound (CMS also never under). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis
- `merge_split` (contiguous|interleaved) is a record field next to `merge_shards` (BenchSection and MergeMetrics; schema regenerated). Merge records stamp it; a record without it split contiguously, so existing results parse, and flatten refuses two different splits in one row as it does two shard counts. - `MergeSplit` moves to aqpbm-core (it is a record type) and is parsed by clap directly instead of matching strings. - The interleaved copy of the stream is built only when a merge runs. - README states the round-robin exactly (contiguous shard sizes, so a short last shard shifts the pattern); MergeMetrics docs name the split. - hydra-cs bound test: note its margin (passes from beta ~4.8). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis
This was referenced Oct 10, 2026
…side merge_split)
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What changed
approxbench --merge-split contiguous|interleaved(defaultcontiguous, today's behavior for every row). Interleaved deals the stream round-robin by record, so each shard is a sample of the whole stream; shard sizes equal the contiguous split's (recordjgoes to the next shard with room).merge_splitnext tomerge_shards(BenchSection,MergeMetrics,merged_record.schema.jsonregenerated). Absent on older records = contiguous, so existing results parse;--flatrefuses two different splits in one row, as it does two shard counts.MergeSplitlives in aqpbm-core and clap parses it directly.Requirement.merge_split→scored_row, which hands every merge body (merge, merge-step, merge-accuracy) the stream reordered bywrappers::interleave, so the existing contiguouspartitioncut yields the round-robin shards. No wrapper signature changes; every row with a merge gets it.0, k, 2k, ...) anddocs/saturation_study.md.Why
Spec PR 4 for the Hydra §6.3 study (hydra doc #193 §2.2–2.3): merges across ingest workers split by record, not by time range.
Tests (
wrappers::tests,wrappers::hydra_shared::tests, 3.7 s debug)interleave_deals_round_robin_at_the_contiguous_sizes.merge_split_on_two_squares_must_agree(flatten: absent = contiguous; interleaved + absent is refused).f ≤ f̂ ≤ f + ε_c·N_q + β(F_v + ε_c·M)/W,ε_c = e/512. CS:[f − ε_s(L2_q + L2_C), f + ε_s·L2_q + βF_v/W + ε_s·L2_C],ε_s = 1/√512,L2_C = √(β·F2/W)withF2 = Σ_v F_v². HLL:[(1−ε_h)D_q, (1+ε_h)(D_q + β·ΣD/W)],ε_h = 3·1.04/√2^14.Mutation checks (each fails the named test):
interleavereturning the stream unchanged; each hydra merge skipping a shard; β = 0 (all three bound tests); inner ε = 0 (CMS, CS); CMS answer 1 under (never-under assert). Inner ε_h = 0 does not fail HLL: at these distinct counts HLL is near-exact. CS margin: passes from β ≈ 4.8 (fails at 4.7), so β = 8 leaves under 2x; cell_cols 1024 would need β ≈ 14 (ε shrinks, the colliders' error does not), so cells stay 3×512. Removing the absent→contiguous default fails the flatten test.Evidence
cargo fmt --check,cargo clippy --workspace --all-targets -D warnings,cargo test --workspaceclean. Release CLI,--merge-shards 8 --metrics accuracy: hydra-cms (hydra_columns.yaml, 3×128, 3×512) and hydra-hllmerge_accuracyequalquery_accuracyunder both splits (are_all 2.5929 / 1.3396); kll-percall--library oxide(deterministic; repeated runs identical)--config k=50 --dataset zipf --size 100000 --zipf-s 1.1 --cardinality 100000 --merge-shards 8 --operations merge --metrics accuracy: merge mean_rank_err 0.012426 contiguous vs 0.012599 interleaved, so the split reaches lossy merges too. Same contiguous command on 653d6a8 vs this head: the record is identical except the new"merge_split": "contiguous"(timings excluded).🤖 Generated with Claude Code
https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis