Repository navigation
rqe-optimizer: let mergeable finer-group deployments serve coarser RAQEs - #190
Conversation
4bc0aa3 to
c20748e
Compare
c20748e to
f5297d3
Compare
| | Ingest, per active `D` | `lambda × a_D × c_ins` | `card(G) × m × a_D` (open windows) | | ||
| | Merge, per RAQE | `card(G) × (n_i − 1) × c_mrg / T_i` | `card(G) × m` (one accumulator per group); 0 when `n_i = 1` | | ||
| | Query, per RAQE | `card(G) × c_qry / T_i` | `card(G) ×` output bytes | | ||
| | Storage, per active `D` | 0 | `card(G) × m × ((max_i S_i − x) / y + 1)` (closed windows) | |
There was a problem hiding this comment.
@zzylol These are part of the old cost model. We should check these against the new cost model and merge tables?
There was a problem hiding this comment.
Yes we should update and merge
There was a problem hiding this comment.
Merged in 332ed6a. Cost by use §3 is now the one cost table: ingest, compaction, query job (merge, then estimate) and storage, each with CPU per event, mean vCPUs, memory, and how long the memory is held. It is written in terms of I_d (instances per window, card(G_d) or 1 for a shared sketch) and I_r (states a query ends with, card(G_r) or 1), summed over parts (sketch + key tracker).
The old snapshot table is gone. The snapshot-AUC section (milp::minimize, still supported) now just lists how it accounts §3's table differently: k_D = 1 so no compaction; query memory held all the time; per-RAQE serial latency against latency_sla_ms instead of batch latency. The DD m definition stayed there and now distinguishes card(G_d) (window) from card(G_r) (merge accumulator).
| Each row of the cost table (from | ||
| [`scripts/study_saturation.py --phase optimizer-cost`](../scripts/study_saturation.py)) | ||
| is measured on one instance, at whatever key count the benchmark fed it | ||
| (`measured_keys`). A key is one distinct entry an instance stores exactly, | ||
| such as one group's running sum in an exact accumulator. Turning a row into a | ||
| deployment's cost needs two properties of the family: |
There was a problem hiding this comment.
@zzylol Did you ever review this section? This was present earlier. I believe this should help with Hydra's cost model
There was a problem hiding this comment.
Reviewed. The shape/law part is right, and it is in fact the Hydra skeleton: Shared + Fixed is I_d = 1. Changes in 332ed6a:
- the section uses
I_d/card(G_d)(andcard(G_r)for probes), matching §3; - the side-by-side table no longer repeats every phase (its latency row was the old serial
G·(c_qry + (n−1)·c_mrg)with no compaction, and its merge rows ignored roll-ups). It now gives what each family puts into §3's table:m/c_mrgbasis,I_d,I_r, probes, key tracker, whether it rolls up; - the HydraKLL column is marked as in the code but not in the evaluation, pointing to the Hydra section, since the
subpopulationsceiling here conflicts with the Hydra section saying one ceiling is insufficient.
One thing I didn't change: KLL is still treated as Fixed at the 1e6-values size. Worth confirming that still matches the #174 study.
There was a problem hiding this comment.
Checked the KLL point myself (3060337). The old text was wrong in both directions:
- The cost table's KLL
mis not "the size measured at 1e6 values". It is sketch-bench's nominal footprintkll_footprint = 4·k·sizeof(T): 1600 B atk = 50, 6400 B atk = 200. approxbench reports the same bytes at 1e3, 1e6 and 1e8 items (run on clnode109 with the final030 binary). So "Fixed" is exactly what the model gets, and nothing grows withn. - It is also not what
asap_sketchlib::KLLallocates: the lib preallocates max capacity over 61 levels (aqpbm-planeval'skll_max_capacityreplica), which is 4640 B atk = 50(2.9×), 8032 B atk = 200(1.25×) and 22240 B atk = 800(0.87×).
The doc now says this. Whether the cost table should switch KLL's footprint to the allocated capacity is a sketch-bench change outside this PR; it would raise small-k KLL memory about 3×.
| | Phase | CPU | Memory | | ||
| | --- | --- | --- | | ||
| | Ingest | `lambda * a * c_ins,h` | `a * m_h` | | ||
| | Merge a RAQE | `(L/x - 1) * c_mrg,h / T` | `m_h` when `L/x > 1` | | ||
| | Query a RAQE | `C_r * c_probe,h(G_r) / T` | `C_r * output_bytes` | | ||
| | Storage | `0` | `ceil((max L - x)/y + 1) * m_h` | | ||
|
|
There was a problem hiding this comment.
@milindsrivastava1997 : This should be merged with the above tables. We should just have one cost model table with all phases and CPU and memory costs both.
There was a problem hiding this comment.
Merged in 332ed6a: the proposed Hydra model is now §3's table with I_d = I_r = 1 and c_qry → c_probe,h(G_r), shown as an instance rather than a separate formula. Differences from the old Hydra table:
- added compaction (
k_D > 1workers producek_Dpartial grids:(k_D − 1)·c_mrg,h, holdingk_D·m_h) andk_Dcopies of open windows at ingest; - symbols follow the doc: lookback
S_i(notL, which is the SLA),card(G_r)/card(G_d)(notC_r/C_d), noceilon storage; - the key tracker is a part, so its cost adds in every phase, as the text already required.
The anchor#proposed-hydra-cost-modelis unchanged, so the link from the candidates doc still works.
|
Also fixed in 332ed6a: the cost-by-use §3/§4 formulas were out of sync with this PR's roll-ups. The code already prices them (
Ingest, compaction and storage stay at 🤖 Generated with Claude Code |
|
Two review rounds (code, docs, Hydra), fixes pushed in 81c46b7 (code) and 00fd9a9 (docs). Code
Docs
Open design questions (not changed in code)
🤖 Generated with Claude Code |
|
Decision on open question 3 (from @zzylol): the AutoSketch-vs-ASAP eval keeps roll-ups on. ASAP may serve an ungrouped RQE from a grouped deployment on the same stream; that sharing is part of what the planner offers. The results on main (#138) were planned before roll-ups, so they will be rerun once this PR lands. 🤖 Generated with Claude Code |
Keep #190 to roll-ups. docs/rqe_optimizer_hydra.md and its links move to a separate PR stacked on this one; the docs here only say Hydra is future work and doesn't roll up. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis
A deployment grouped by G_d now serves a RAQE grouped by any subset G_r when its family merges across groups (exact sum/min/max, HLL, UnivMon cardinality, KLL, DD). The read is priced at G_r: card(G_d)·L/x − card(G_r) merges, with card(G_r) accumulators, probes and output. Accuracy reads G_r's shape and items, and KLL's merge curve at the average fan-out times L/x. Fine candidates take windows from coarser RAQEs. validate_facts rejects a label set with more groups than a superset. Closes #187 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Merge the cost-by-use CPU and memory tables into one table per phase (ingest, compaction, query job, storage) in terms of I_d and I_r, so roll-ups (G_r ⊂ G_d) are priced as the code prices them: ℓ = card(G_r)·c_qry + (I_d·n − I_r)·c_mrg, accumulators at G_r. The snapshot-AUC section now lists only how it differs (k = 1, query memory always held, per-RAQE serial latency) instead of repeating the table; the shape/size-law side-by-side gives each family's I_d, I_r and tracker; the proposed Hydra cost model is §3's table with I_d = I_r = 1, adding compaction and using the doc's symbols. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis
…at 1e6 The cost table's KLL m is kll_footprint = 4·k·sizeof(T), independent of the item count (1600 B at k = 50 for 1e3, 1e6 and 1e8 items), so it is not "the size measured at 1e6 values" and does not grow with n. Say so, and how it compares with asap_sketchlib's preallocated capacity. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis
Dominance pruning compared ingest, query latency, merge memory and storage. Cost by use also bills compaction (per window in every chain, per slide in CPU and byte-seconds) and query output memory, which differs between top-k configs answering different k. A direct RAQE never merges at query time, so a candidate with cheaper ingest but costlier merges could prune one that is cheaper by use. Compare those too; the new test fails without it. Also: workload_scenarios no longer says groupings never share; the exact-increase exclusion states its real (conservative) reason; merged_instance_count notes roll-ups; a circular doc comment; stale §N citations now name the docs. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis
- candidates: eligibility allows subset groupings for families that merge across groups; generation shows the roll-up RAQE sets; accuracy is read at G_r with KLL's fan-out merge count, and past the measured merge counts the guarantee applies, floored by the last measurement; pruning lists every compared cost; the MILP lives only in the cost doc; one roll-up section. - cost model: top-k output and query CPU scale with k; k_D sums parts; mergeable_across_groups listed; KLL footprint points at #191/#192; Hydra moved out. - v1 and TODO: phases include compaction and the query job; TODOs match cost by use. - rqe_optimizer_hydra.md (new): Hydra in sketch-bench, accuracy bounds per combination (CMS, CountSketch, HLL, UnivMon, KLL) checked against the paper's Theorem 2 and the library (correlated inner seeds, fixed column hash, few rows), cost model, candidates and MILP edges, sketch-bench requirements and open decisions. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis
…ason UnivMon's counters add exactly on merge, but each level's heavy-hitter heap is rebuilt from the union of the two heaps, so a key heavy only in the union is lost: a merged (or rolled-up) answer now reads merge curves, as KLL's does. exact-increase's accumulator merge joins two pieces of one counter in time, so merging two groups' counters is wrong; say so instead of "unconfirmed". Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis
validate_facts compared every pair of label sets in a metric's cardinality map, so an inconsistent entry no RAQE groups by rejected the whole workload. The roll-up merge count only reads RAQE groupings and the full label set, so check X ⊂ Y among those. Docs: why rate/increase, top-k and Hydra don't roll up. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis
Every Hydra measurement predates the admission requirements, so they are all rerun; hydra-kll goes in that run with the other four variants. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis
Keep #190 to roll-ups. docs/rqe_optimizer_hydra.md and its links move to a separate PR stacked on this one; the docs here only say Hydra is future work and doesn't roll up. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis
Tests that fail without their rule: - a DDSketch roll-up sizes its accumulator at the coarse grouping; - validate_facts rejects a RAQE grouping with more groups than series; - pruning drops a heap whose larger answered k only adds output memory. Docs: the KLL footprint paragraph now that #192 is in; univmon-cardinality reads merge curves like KLL (and has no guarantee past them); the TODO's done list covers roll-ups and the pruning costs; "temporal pre-merging"; enumerate.rs names the MILP variables z and u as the docs do. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis
ad069f1 to
6b1637a
Compare
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis
Closes #187.
A deployment grouped by
G_dcan now serve a RAQE grouped by a subsetG_rwhen its family merges across groups. The read merges eachG_rgroup's fine states into one (a roll-up), and the planner chooses it only when it is SLA-valid and cheapest. ASAPQuery can't run roll-ups yet; see Interface change.Design decisions (please approve)
These are the decisions from the design discussion on #187 (Q&A, Q1 change). Please approve each one or push back. Example used throughout: KLL by
(service, endpoint)(50 groups) serving p99 by(service)(5 groups),L = 1h,x = 1m.Scope
MultipleIncreaseper series, so there is no engine bug there.is_eligible.Q1. Which families roll up? New
FamilyProperties::mergeable_across_groups.is_eligible: the groupings are equal, or the family ismergeable_across_groupsandr.grouping ⊂ d.grouping. A coarse deployment never serves a finer RAQE.Q2. Candidate windows. For families that roll up, a grouping's windows, slides and heaps come from every RAQE with the same capability, metric and filter whose grouping is a subset of it. That way a fine
xcan divide a coarse RAQE'sL. Other families keep their exact-grouping inputs, so top-k candidates don't grow. No new groupings are invented.Q3. Cost of a roll-up read. Each merge folds two states into one. A query reads
card(G_d)·L/xstates and ends withcard(G_r):card(G)·(L/x − 1)·c_mrg / Tmerges · c_mrg / Tcard(G)·m(L)card(G_r)·m(L)card(G)·…card(G_r)·…card(G)·c_qry + card(G)·(L/x−1)·c_mrgcard(G_r)·c_qry + merges·c_mrgG_dG_r = G_d, each row reduces to the old formula.L == x) roll-up still mergescard(G_d) − card(G_r)states.c_mrg.Q4. Accuracy of a roll-up. It reads
data_shape[G_r]and N= λ·L / card(G_r). KLL and univmon-cardinality merge lossily (#131; UnivMon rebuilds its heavy-hitter heaps from the union on merge), so they read their merge curves at⌈card(G_d)/card(G_r)⌉ · L/xmerged instances. That's the average fan-out; the worst group's is in #189. Past the measured merge counts, KLL takes its guarantee. univmon-cardinality has no guarantee, so it has no accuracy there; it isn't in the study orDEPLOYABLE_FAMILIEStoday, so nothing changes in practice.Q5. Plan output. No new field. A RAQE rolls up exactly when its grouping differs from its deployment's.
Q6. Fact validation.
validate_factsrequirescard(X) ≤ card(Y)for every pairX ⊂ Yamong the label sets a plan can use: every RAQE grouping on the metric, plus its full label set. Label sets no RAQE groups by are not checked. This replaces the "more groups than series" check. Otherwise bad facts would make the merge count negative, and the MILP would treat a roll-up as a discount.Interface change (ASAPQuery)
grouping_labelsdiffers from its deployment's. Until ASAPQuery can merge fine groups into coarse ones at read time, it should error on such a plan when it bumps the pinned rev.validate_factsrejects more inputs: a RAQE grouping with more groups than another RAQE grouping, or than the full label set, that is a superset of it.Tests
tests/mixed_grouping.rs(new, end to end):(service, endpoint)and(service)share one fine KLL deployment;candidates.rs:analytical_cost_model.rs: roll-up merge count, accumulators and latency, including a direct roll-up.saturation.rs: KLL roll-up reads the coarse shape, coarse N and the fan-out × windows merge curve.lib.rs: subset-cardinality validation.candidates.rs(pruning):retains_candidate_with_cheaper_compactionandprunes_the_candidate_with_larger_query_output.analytical_cost_model.rs:a_dd_roll_up_sizes_its_accumulator_at_the_coarse_grouping.saturation.rs:univmon_cardinality_merges_lossily_hll_and_dd_do_not.lib.rs:rejects_a_used_label_set_with_more_groups_than_a_used_superset,rejects_a_grouping_with_more_groups_than_seriesandignores_label_sets_no_raqe_uses. These replace the oldrejects_more_groups_than_series.cargo test(104 + 2),clippy --workspace -D warningsandfmtall pass, rebased on main (#192).examples/workload_scenarios.rsagainst main: the only change is the mixed-grouping sweep's "both" case (p99 by(service)plus p99 by(service, endpoint)), which now runs on 1 DDSketch deployment instead of 2:Every other sweep is byte-identical.
Also in this PR (found in review)
validate_factschecks group counts only among the label sets a plan can use: every RAQE grouping on the metric, plus its full label set.rqe_optimizer_candidates.mdandrqe_optimizer_cost_model.md, with one cost table covering every phase, roll-ups included. They now say why rate/increase, top-k and Hydra don't roll up.Open questions
Raqedoesn't record the counted key yet.examples/small_problem.rsaddslatency_p99_1h_by_service. With--milp --saturation-dir study_final030k, it rolls up onto the per-endpoint DD deployment:🤖 Generated with Claude Code