Skip to content

kll: report asap_sketchlib's allocation as the lib rows' footprint - #192

Merged
zzylol merged 2 commits into
mainfrom
191-kll-lib-footprint
Oct 10, 2026
Merged

zzylol merged 2 commits into
mainfrom
191-kll-lib-footprint

Conversation

@zzylol

@zzylol zzylol commented Oct 9, 2026 •

Copy link
Copy Markdown
Collaborator

Closes #191.

The kll-percall lib rows now report what asap_sketchlib::KLL allocates at construction, kll_lib_bytes::<T>(k):

  • the retained-item slots, compute_max_capacity(k, m) (a copy of the library's private function);
  • the level index, (MAX_LEVELS + 1) usizes;
  • the merge buffer's capacity, k items.

The size is still independent of the item count, because the library allocates once. At k = 50, 200, 800 with i64 it is 5536, 10128 and 29136 B, against 1600, 6400 and 25600 B before.

What changed

  • wrappers/kll/mod.rs: the capacity replica and its constants moved here from hydra_kll/mod.rs (renamed kll_lib_slots, KLL_LIB_*). kll_lib_bytes was added, and kll_lib_k reproduces the library's floor and cap on k.
  • wrappers/kll/sketchlib.rs: both lib rows (per-call and cdf) use kll_lib_bytes.
  • hydra_kll: imports the moved replica. Its footprint is unchanged: slots plus level index per cell, without the merge buffer. Hydra is out of scope here.
  • The oxide rows keep the nominal 4·k, because sketch_oxide's KLL grows its level vectors as it fills, so there is no fixed allocation to report. The doc comment now says so.

Tests

  • New lib_footprint_is_the_library_allocation. At k = 269 it gives 12296 B, the figure aqpbm-planeval derives independently for the same library. It also pins kll_lib_slots(50) = 580 and the floor below LIB_K_MIN.
  • cargo fmt --check, clippy --workspace --all-targets -D warnings (without rqe-optimizer, which needs libclang locally), cargo test -p sketch-bench (96 passed) and -p aqpbm-cli pass.

What this PR doesn't count

  • The cdf row's prepared CDF table (about 16 B per retained item), and the oxide cdf row's, built for queries.
  • Each hydra-kll cell's merge buffer. hydra-kll is out of scope; a comment at memory_hydra_kll notes the difference.

Follow-up (not in this PR): committed data with the old KLL memory (1600, 6400, 25600 B), to regenerate together

🤖 Generated with Claude Code

https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

The kll-percall lib rows reported a nominal 4·k·sizeof(T), which is not
what asap_sketchlib::KLL allocates: init_internal boxes
compute_max_capacity(k, m) slots once, plus a level index and a merge
buffer of capacity k (5536 B at k = 50, not 1600 B). Report that, from
the capacity replica the hydra-kll wrapper already had, now shared in
the kll module. The oxide rows keep the nominal size: sketch_oxide's
KLL grows as it fills.

Closes #191.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis
zzylol added a commit that referenced this pull request Oct 9, 2026
- candidates: eligibility allows subset groupings for families that
  merge across groups; generation shows the roll-up RAQE sets; accuracy
  is read at G_r with KLL's fan-out merge count, and past the measured
  merge counts the guarantee applies, floored by the last measurement;
  pruning lists every compared cost; the MILP lives only in the cost
  doc; one roll-up section.
- cost model: top-k output and query CPU scale with k; k_D sums parts;
  mergeable_across_groups listed; KLL footprint points at #191/#192;
  Hydra moved out.
- v1 and TODO: phases include compaction and the query job; TODOs match
  cost by use.
- rqe_optimizer_hydra.md (new): Hydra in sketch-bench, accuracy bounds
  per combination (CMS, CountSketch, HLL, UnivMon, KLL) checked against
  the paper's Theorem 2 and the library (correlated inner seeds, fixed
  column hash, few rows), cost model, candidates and MILP edges,
  sketch-bench requirements and open decisions.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis
The lib footprint is the sketch's own allocation; say that the cdf
row's prepared CDF table is not counted. KLL_LIB_MIN_LEVEL is the m
init_kll passes, 8, with a const assertion that LIB_K_MIN agrees,
instead of being defined through it. hydra-kll's footprint notes it
omits each cell's merge buffer.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis
@zzylol
zzylol merged commit f9c4b68 into main Oct 10, 2026
2 checks passed
zzylol added a commit that referenced this pull request Oct 10, 2026
- candidates: eligibility allows subset groupings for families that
  merge across groups; generation shows the roll-up RAQE sets; accuracy
  is read at G_r with KLL's fan-out merge count, and past the measured
  merge counts the guarantee applies, floored by the last measurement;
  pruning lists every compared cost; the MILP lives only in the cost
  doc; one roll-up section.
- cost model: top-k output and query CPU scale with k; k_D sums parts;
  mergeable_across_groups listed; KLL footprint points at #191/#192;
  Hydra moved out.
- v1 and TODO: phases include compaction and the query job; TODOs match
  cost by use.
- rqe_optimizer_hydra.md (new): Hydra in sketch-bench, accuracy bounds
  per combination (CMS, CountSketch, HLL, UnivMon, KLL) checked against
  the paper's Theorem 2 and the library (correlated inner seeds, fixed
  column hash, few rows), cost model, candidates and MILP edges,
  sketch-bench requirements and open decisions.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis
zzylol added a commit that referenced this pull request Oct 10, 2026
Tests that fail without their rule:
- a DDSketch roll-up sizes its accumulator at the coarse grouping;
- validate_facts rejects a RAQE grouping with more groups than series;
- pruning drops a heap whose larger answered k only adds output memory.

Docs: the KLL footprint paragraph now that #192 is in; univmon-cardinality
reads merge curves like KLL (and has no guarantee past them); the TODO's
done list covers roll-ups and the pruning costs; "temporal pre-merging";
enumerate.rs names the MILP variables z and u as the docs do.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis
zzylol added a commit that referenced this pull request Oct 10, 2026
…QEs (#190)

* rqe-optimizer: let mergeable finer-group deployments serve coarser RAQEs

A deployment grouped by G_d now serves a RAQE grouped by any subset G_r when
its family merges across groups (exact sum/min/max, HLL, UnivMon cardinality,
KLL, DD). The read is priced at G_r: card(G_d)·L/x − card(G_r) merges, with
card(G_r) accumulators, probes and output. Accuracy reads G_r's shape and
items, and KLL's merge curve at the average fan-out times L/x. Fine candidates
take windows from coarser RAQEs. validate_facts rejects a label set with more
groups than a superset.

Closes #187

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs: split RQE optimizer design by concern

* docs: clarify RQE optimizer cost model status

* docs: one cost table for every phase, with roll-ups

Merge the cost-by-use CPU and memory tables into one table per phase
(ingest, compaction, query job, storage) in terms of I_d and I_r, so
roll-ups (G_r ⊂ G_d) are priced as the code prices them:
ℓ = card(G_r)·c_qry + (I_d·n − I_r)·c_mrg, accumulators at G_r.

The snapshot-AUC section now lists only how it differs (k = 1, query
memory always held, per-RAQE serial latency) instead of repeating the
table; the shape/size-law side-by-side gives each family's I_d, I_r and
tracker; the proposed Hydra cost model is §3's table with I_d = I_r = 1,
adding compaction and using the doc's symbols.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* docs: KLL's memory is sketch-bench's nominal footprint, not measured at 1e6

The cost table's KLL m is kll_footprint = 4·k·sizeof(T), independent of
the item count (1600 B at k = 50 for 1e3, 1e6 and 1e8 items), so it is not
"the size measured at 1e6 values" and does not grow with n. Say so, and
how it compares with asap_sketchlib's preallocated capacity.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* rqe-optimizer: prune on every cost cost-by-use bills; fix stale comments

Dominance pruning compared ingest, query latency, merge memory and
storage. Cost by use also bills compaction (per window in every chain,
per slide in CPU and byte-seconds) and query output memory, which
differs between top-k configs answering different k. A direct RAQE
never merges at query time, so a candidate with cheaper ingest but
costlier merges could prune one that is cheaper by use. Compare those
too; the new test fails without it.

Also: workload_scenarios no longer says groupings never share; the
exact-increase exclusion states its real (conservative) reason;
merged_instance_count notes roll-ups; a circular doc comment; stale §N
citations now name the docs.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* docs: consolidate the optimizer docs; add the Hydra admission design

- candidates: eligibility allows subset groupings for families that
  merge across groups; generation shows the roll-up RAQE sets; accuracy
  is read at G_r with KLL's fan-out merge count, and past the measured
  merge counts the guarantee applies, floored by the last measurement;
  pruning lists every compared cost; the MILP lives only in the cost
  doc; one roll-up section.
- cost model: top-k output and query CPU scale with k; k_D sums parts;
  mergeable_across_groups listed; KLL footprint points at #191/#192;
  Hydra moved out.
- v1 and TODO: phases include compaction and the query job; TODOs match
  cost by use.
- rqe_optimizer_hydra.md (new): Hydra in sketch-bench, accuracy bounds
  per combination (CMS, CountSketch, HLL, UnivMon, KLL) checked against
  the paper's Theorem 2 and the library (correlated inner seeds, fixed
  column hash, few rows), cost model, candidates and MILP edges,
  sketch-bench requirements and open decisions.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* rqe-optimizer: univmon-cardinality merges lossily; increase's real reason

UnivMon's counters add exactly on merge, but each level's heavy-hitter
heap is rebuilt from the union of the two heaps, so a key heavy only in
the union is lost: a merged (or rolled-up) answer now reads merge
curves, as KLL's does.

exact-increase's accumulator merge joins two pieces of one counter in
time, so merging two groups' counters is wrong; say so instead of
"unconfirmed".

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* rqe-optimizer: check group counts only between label sets a plan uses

validate_facts compared every pair of label sets in a metric's
cardinality map, so an inconsistent entry no RAQE groups by rejected the
whole workload. The roll-up merge count only reads RAQE groupings and the
full label set, so check X ⊂ Y among those.

Docs: why rate/increase, top-k and Hydra don't roll up.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* docs(hydra): hydra-kll is measured in the same Hydra rerun

Every Hydra measurement predates the admission requirements, so they
are all rerun; hydra-kll goes in that run with the other four variants.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* docs: move the Hydra admission design out of this PR

Keep #190 to roll-ups. docs/rqe_optimizer_hydra.md and its links move
to a separate PR stacked on this one; the docs here only say Hydra is
future work and doesn't roll up.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* rqe-optimizer: tests for roll-up DD memory, facts and pruning; doc fixes

Tests that fail without their rule:
- a DDSketch roll-up sizes its accumulator at the coarse grouping;
- validate_facts rejects a RAQE grouping with more groups than series;
- pruning drops a heap whose larger answered k only adds output memory.

Docs: the KLL footprint paragraph now that #192 is in; univmon-cardinality
reads merge curves like KLL (and has no guarantee past them); the TODO's
done list covers roll-ups and the pruning costs; "temporal pre-merging";
enumerate.rs names the MILP variables z and u as the docs do.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* rqe-optimizer: name every lossy non-heap sketch that needs merge curves

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: zz_y <zeyingz@umd.edu>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

kll-percall (lib) reports a nominal 4·k footprint, not asap_sketchlib's allocation

1 participant