Skip to content

rqe-optimizer: admit Hydra (HLL, UnivMon cardinality, KLL) as candidates, with measured accuracy by coverage - #201

Merged
zzylol merged 11 commits into
mainfrom
rqe-optimizer/hydra
Oct 11, 2026
Merged

zzylol merged 11 commits into
mainfrom
rqe-optimizer/hydra

Conversation

@zzylol

@zzylol zzylol commented Oct 10, 2026 •

Copy link
Copy Markdown
Collaborator

Implements PR 7 of the Hydra plan (docs/rqe_optimizer_hydra.md §3–§4, #193) on top of #196, which the runner needs for coverage and schema streams. Reads hydra_saturation.csv in the format PR 6 will write (the "eval-shaped Hydra measurements" amendment).

What changed

  • Families (lib.rs): hydra-hll and hydra-univmon-cardinality (Cardinality, relative_error) and hydra-kll (Quantile, mean_rank_err) get FamilyProperties: one fixed-size grid, DeltaSet key tracker, mergeable_across_groups = false, and a new answers_any_subgrouping. They are in Capability::families but not in DEPLOYABLE_FAMILIES. hydra-kll and hydra-univmon-cardinality merge lossily; hydra-hll merges exactly.
  • Candidates (candidates.rs): Hydra candidates are built only on a metric with MetricFacts::hydra_dataset, which holds the dataset and the metric's full schema Λ (every schema label, without x). Grids are built only at Λ, since the study measures cost and accuracy at that width: one Hydra candidate per (stream, variant, config, x, y). A Hydra cost row is used only if measured_at.dataset equals the metric's dataset, so rows without a dataset are never candidates. Windows come from RAQEs with ∅ ≠ G_r ⊆ Λ. An edge is eligible for ∅ ≠ G_r ⊆ Λ, with today's window rules and the Hydra accuracy. It is never treated as a roll-up. Λ is a label set, not an ordered list, because the eval datasets fix label order per schema. The old subpopulations ceiling (measured_at_group_count) is removed, since the measured lookup replaces it.
  • Docs (rqe_optimizer_candidates.md, rqe_optimizer_cost_model.md, runner module doc): Hydra is held to the worst case (max over covered groups, worst over N and seeds), while per-group sketches read seed-mean curves. Hydra is built only at the full schema because that is what is measured.
  • Cost: no formula changes. The existing fixed-size-sketch pricing already matches §3 with I_d = I_r = 1: ingest is λ·(x/y)·c_ins per record, memory is one grid per window, a query makes card(G_r) probes, merges are L/x − 1, and the tracker is priced at Λ as today. Only the module doc is updated.
  • MeasuredAt (aqpbm-core): new optional dataset: Option<String> and schema_width: Option<u32> (serde default, skipped if None). PR 6 sets them on Hydra rows. Existing tables load unchanged. MeasuredAt derives no JSON schema, so there is nothing to regenerate.
  • Accuracy (saturation.rs): hydra_saturation.csv under the saturation dir is optional. The lookup key is (variant, config params, dataset, grouping as a label set). Raqe::accuracy_covers_share picks the column: None reads err_max; τ ≥ 0.05 reads err_max_cov_0.05; 0.01 ≤ τ < 0.05 reads err_max_cov_0.01; τ < 0.01 reads err_max, which is always the stricter choice. The value is the worst over records and seeds, and an empty cell leaves it unknown. Lossy merges of L/x windows take the worse of the bracketing merge_shards; past the largest the value is unknown. A missing file, dataset or grouping means no accuracy, so the edge is not eligible. read_csv now uses the csv crate because grouping fields are quoted and contain commas. It checks the header for each caller's required columns and fails on a row that is shorter than the header. A config param that isn't name=number is also an error now, where before the row was silently ineligible.
  • Runner (autosketch_vs_asap.rs): ASAP, PerQuery and AutoSketch all plan with undeployable families allowed, so every method sees the same non-Hydra families. AutoSketch still skips Hydra itself via one_fixed_size_sketch_for_all_groups (rqe-optimizer: AutoSketch baseline should rank shared fixed-size sketches (HydraKLL) by deployment cost #159). On a schema stream, card(full schema) is the product of the label fan-outs along child_of (the generator's grouping_cardinality: http 10⁴, flows 3·10⁶), capped at the stream's series count so the high-cardinality filter doesn't drop the stream. RQEs carry covers_share. Streams of http map to hydra_http, or hydra_http_latency for quantiles; flows maps to hydra_flows. A stream is one value of one metric, so one dataset per MetricFacts is enough. On workloads with a Hydra dataset, each result reports hydra_rqes.

Tests

Each new test was mutation-checked: 13 mutations, all killed.

  • hydra_serves_any_non_empty_subset_of_its_schema covers the subset rule, including non-prefix subsets, ∅, labels outside Λ, and that accuracy still applies.
  • hydra_grids_are_built_at_the_full_schema_from_its_dataset_rows checks that only the full schema (without x) gets a grid, with the right windows. It uses two Hydra rows from other datasets (hydra_other, and one with no dataset) and checks that only the hydra_test row is used. It also checks that a metric with no dataset gets no grid.
  • flows_and_http_latency_get_a_full_schema_hydra_grid (runner) builds the table from the generator's schemas and adds fixture Hydra rows per dataset. flows gets hydra-hll at {dst_subnet, dst_port, proto}, and http latency gets hydra-kll at all four http labels. flows does not pick up a row measured on hydra_http. flows' card(Λ) is capped at its 1e6 series, and the stream is kept.
  • a_malformed_hydra_csv_is_an_error checks three cases: a missing column, a short row, and a non-numeric config param.
  • The aqpbm-core cost-row serde test checks that a row with dataset and schema_width serializes and round-trips, and that both fields are left out when None.
  • hydra_reads_the_column_its_coverage_selects covers coverage column selection.
  • hydra_takes_the_worst_over_records_and_seeds covers worst over N and seeds, and that an empty cell makes the value unknown.
  • a_lossy_hydra_merge_reads_the_worse_bracketing_shard_count covers merges 2→max(1,4), 8→16, 32→unknown, and that HLL reads shards = 1.
  • without_a_hydra_measurement_no_hydra_deployment_has_accuracy covers a missing CSV, dataset or grouping.
  • a_hydra_grid_beats_per_group_hll_on_memory_only_when_accurate is a tiny end-to-end MILP from an inline CSV. Hydra wins on memory at error 0.01 and HLL wins at 0.2.
  • The runner tests check the dataset map, that covers_share is threaded through, and hydra_rqes = 0 without the CSV.

Evidence

Round 1 (39a4d16):

  • cargo fmt --check, cargo clippy --workspace --all-targets -D warnings: clean. cargo test --workspace: all pass. cargo test -p rqe-optimizer --example autosketch_vs_asap: 8 pass.
  • Each new rule was mutated and the tests caught every mutation: without the dataset check, without the series cap, with a lenient config parse, and without the short-row check.
  • I regenerated tables with the base branch's generator; they are byte-identical to this branch's. I ran base (561ee86) and this branch with --runs 1 against the committed saturation dir, which has no Hydra CSV or Hydra rows. With *_secs timings stripped, classic shared1 is identical, including AutoSketch now that it allows undeployable families. All shared1 is identical apart from the added "hydra_rqes": 0. Objectives: classic asap 3.31634 / autosketch 3410.17 / perquery 9.553133; all asap 5.052212 / autosketch 3417.90 / perquery 11.755537.

Round 0:

  • cargo fmt --check, cargo clippy -p rqe-optimizer --all-targets -D warnings: clean.
  • cargo test -p rqe-optimizer: lib 111 (base 104, all 104 kept), integration 2, runner 7 (base 6).
  • Base (561ee86) and this branch were run on freshly generated tables (--runs 1) with the committed saturation dir, which has no Hydra CSV and no Hydra cost rows. Classic shared1 is identical once *_secs timings are stripped. All shared1 is identical apart from the added "hydra_rqes": 0. Objectives match: classic asap 3.31634 / perquery 9.553133 (cpu); all asap 5.052212.
  • Wiring smoke test, not a result: fixture Hydra cost rows and a fixture CSV on the all table. Hydra candidates appear (793 → 843) and are eligible with the expected column (t14 τ=0.01 → cov_0.01, t15 τ=0.05 → cov_0.05). With the fixture's 45× per-record insert they are not chosen.

Notes

  • PR 6's Hydra rows are not here yet. Until they land, no Hydra candidate is eligible in the eval.
  • Review round 2 (902cd73): validate_facts rejects a Hydra schema outside the labels or without a nonzero cardinality; the Hydra accuracy fold skips empty cells (unknown only if all are empty); read_csv rejects rows longer than the header. Tests for each, checked by mutation.

Roll-up ablation (b0319d1)

  • Runner: new method asap-norollup ("ASAP (no roll-ups)"). It is ASAP's MILP with the same candidates, accuracy, frontier sweep and SLA grid, but allowed keeps each RQE to deployments at its own grouping or Hydra grids. Hydra is not a roll-up. Every result (all methods) gets rolled_up_rqes: the RQEs served by a finer non-Hydra deployment. Sanity checks now cover asap ≤ asap-norollup ≤ perquery, both unbounded and at each SLA, with the existing 1e-6 tolerance. The frontiers use different bounds per method, so they are not compared point by point. No rqe-optimizer/src changes.

  • Plots: both scripts draw the ablation as a dashed violet line with ▲ markers. fig_rollup_ablation.png (synthetic script) shows grouped bars of ASAP vs. ASAP (no roll-ups) unbounded cost per workload, in CPU and Fargate panels, with absolute values labeled. summary_synthetic.md and summary_sla.md gain the method's rows and a roll-up section: RQEs rolled up, deployments with/without, cost with/without, and saving %. Results without the method still plot.

  • Size: run_autosketch_vs_asap_synthetic.sh passes the existing --no-chosen. rolled_up_rqes is kept either way.

  • Tests: forbidding_roll_ups_costs_more_only_where_one_applies covers two cases. With {region} and {region, service}, ASAP rolls up 1 RQE on 1 deployment and is strictly cheaper than no-roll-ups, which has 0 rolled up and 2 deployments. With {region} and {service}, the two costs are equal and nothing is rolled up. The Hydra-grid test checks that a grid finer than its RQE is not a roll-up. Mutation-checked: dropping the Hydra exemption, or never treating a deployment as a roll-up, fails these tests.

  • Evidence: fmt, clippy -p rqe-optimizer --all-targets --examples -D warnings, cargo test -p rqe-optimizer (113+2) and the example tests (10) all pass. Smoke run (not committed): fresh --synthetic tables, --runs 1, the 2026-10-08 saturation dir, 0 sanity violations. Unbounded, ASAP vs. no roll-ups:

    • mixed + multi-grouping: cpu 3.53 vs 3.65 (42 RQEs rolled up, 14 vs 30 deployments); fargate 0.233 vs 0.245.
    • mixed + multi-grouping, r = 8: cpu 3.63 vs 3.75.
    • mixed + multi-grouping, m = 8: cpu 28.3 vs 29.2.
    • mixed (classic, every variant): equal, with 0 rolled up.

    The committed tables predate schemas, so on them the two methods are equal. The committed results were not regenerated.

🤖 Generated with Claude Code

https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

zzylol and others added 3 commits October 10, 2026 04:02
The generator gets a metric-level ordered label schema (labels with
cardinality, child_of fan-out and skew; per-value data shape) and
per-RQE `grouping` and `covers_share`. The 10 templates are now the
"classic" set, byte-for-byte the committed tables; "all" adds templates
11, 12, 14, 15, 16, 17 on `http` and `flows`, whose streams each carry
several groupings so roll-ups (and later Hydra) can share. The plan runs
both sets over the grid.

The runner builds each schema stream's MetricFacts from the schema
(labels = schema + x; cardinality per used grouping and the full set),
takes Raqe.grouping_labels from the RQE's grouping, maps the
`cardinality` capability to HLL, and reports covers_share and roll-up
deployment groupings in each choice. Classic output is unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis
Review of #196. http's labels are now Zipf-skewed (region 0.5, service
1.1, endpoint 1.1 within its service, status 2.0) and the flows DDoS
burst targets a tail subnet, so covers_share selects a few groups. The
generator computes each group's share (product of its labels' shares)
and, per RQE, the covered groups (share >= tau, plus always the largest)
and the smallest covered group's share and items per window. The runner
reads a covered RQE's accuracy (ASAP and AutoSketch) at that group's
items instead of the mean group's; covers_share = null is unchanged.

The schema test reads an inline HLL curve fixture instead of the
gitignored study curves. The synthetic plot keys runs by template set
too, and both plot scripts title classic "mixed" and all "mixed +
multi-grouping". README: the coverage rule, the N the runner uses, and
which schema fields are descriptive. label_set.keys_per_window is
commented as the universe K.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis
… {service, endpoint}

- covers_share = null covers every group, so the runner reads its accuracy
  at the smallest group (min_covered_share), like a tau RQE's smallest
  covered group. Classic RQEs (no covers_share) are unchanged.
- Template 17 (negative control) groups by {service, endpoint}: its
  smallest group holds 4.5e4 items in 5m, on the curves.
- The generator asserts every schema RQE's covered_min_N >= 1e3 (drops the
  vacuous `assert covered`); a test checks the old full-label-set template
  fails it.
- Covered facts go through validate_facts.
- README: null semantics, template 17, and the numbers' provenance (fresh
  `all` table, committed saturation inputs).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis
zzylol added a commit that referenced this pull request Oct 10, 2026
Round-1 review of #201:
- One Hydra grid per (stream, variant, config, window, slide), over the
  metric's full schema only (MetricFacts::hydra_dataset now carries the
  dataset and the schema), the width the study measures.
- MeasuredAt gains optional `dataset` and `schema_width`; a Hydra row is a
  candidate on a metric only if measured on its dataset.
- Runner: card(full schema) = product of label fan-outs, capped at the
  stream's series; AutoSketch also allows undeployable families (it skips
  Hydra itself, #159).
- read_csv checks required columns and short rows; config params must be
  numbers.
- Docs: Hydra held to the worst case, built only at the full schema.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis
Per sketch-bench#189, KLL/DD/HLL error does not fall as N grows, so the
largest covered group can be harder than the smallest. The generator also
emits max_covered_share / covered_max_N (the largest group, always covered);
the runner reads each covered RQE's accuracy at both the smallest and the
largest covered group and keeps the worse (direction-aware), unknown if
either read is. Classic RQEs and Hydra's lookup are unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis
zzylol added a commit that referenced this pull request Oct 10, 2026
Round-1 review of #201:
- One Hydra grid per (stream, variant, config, window, slide), over the
  metric's full schema only (MetricFacts::hydra_dataset now carries the
  dataset and the schema), the width the study measures.
- MeasuredAt gains optional `dataset` and `schema_width`; a Hydra row is a
  candidate on a metric only if measured on its dataset.
- Runner: card(full schema) = product of label fan-outs, capped at the
  stream's series; AutoSketch also allows undeployable families (it skips
  Hydra itself, #159).
- read_csv checks required columns and short rows; config params must be
  numbers.
- Docs: Hydra held to the worst case, built only at the full schema.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis
@zzylol
zzylol force-pushed the rqe-optimizer/hydra branch from 902cd73 to c11dce3 Compare October 10, 2026 11:35
zzylol added a commit that referenced this pull request Oct 10, 2026
After rebasing onto #196's two-ended coverage read: #201's runner test RQEs
need max_covered_share, and #196's coverage test schema needs the label
cardinality the full-schema Hydra grid reads.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis
zzylol added a commit that referenced this pull request Oct 10, 2026
… for the §6.3 eval

# Conflicts:
#	aqpbm-core/src/atomic_costs.rs
zzylol and others added 5 commits October 10, 2026 12:25
…_port)

(dst_subnet, dst_port) has 1e6 groups; answering each every minute takes
about 49 s on one core, which set every plan's batch latency above the
largest SLA (10 s) and collapsed the multi-grouping frontier. (dst_subnet,
proto) has 3e3 groups and keeps three groupings over one stream.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis
…tes, with measured accuracy by coverage

- FamilyProperties for hydra-hll, hydra-univmon-cardinality and hydra-kll:
  one fixed-size grid, DeltaSet key tracker, no roll-up, and the new
  answers_any_subgrouping. Added to Capability::families, not to
  DEPLOYABLE_FAMILIES.
- Candidates: on a metric with a hydra_dataset, one grid per schema Λ (each
  RAQE grouping and the union, where card(Λ) is known); eligible for any
  RAQE with ∅ ≠ G_r ⊆ Λ. The subpopulations ceiling is gone.
- Accuracy: hydra_saturation.csv under the saturation dir (optional), keyed
  by variant, config, dataset and grouping; the column the RAQE's
  accuracy_covers_share selects; worst over records and seeds; lossy merges
  bracketed by merge_shards, unknown past the largest.
- Runner: ASAP and PerQuery plan with undeployable families; RQEs carry
  covers_share; http/flows streams name their Hydra dataset; results report
  hydra_rqes for workloads Hydra may serve.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis
Round-1 review of #201:
- One Hydra grid per (stream, variant, config, window, slide), over the
  metric's full schema only (MetricFacts::hydra_dataset now carries the
  dataset and the schema), the width the study measures.
- MeasuredAt gains optional `dataset` and `schema_width`; a Hydra row is a
  candidate on a metric only if measured on its dataset.
- Runner: card(full schema) = product of label fan-outs, capped at the
  stream's series; AutoSketch also allows undeployable families (it skips
  Hydra itself, #159).
- read_csv checks required columns and short rows; config params must be
  numbers.
- Docs: Hydra held to the worst case, built only at the full schema.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis
…ect long CSV rows

validate_facts reports a Hydra schema outside the metric's labels or
without a nonzero cardinality, instead of a later panic. The Hydra
accuracy fold skips empty cells (no covered group scored in that run)
and is unknown only when every row left the column empty. read_csv
rejects rows longer than the header too.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis
After rebasing onto #196's two-ended coverage read: #201's runner test RQEs
need max_covered_share, and #196's coverage test schema needs the label
cardinality the full-schema Hydra grid reads.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis
@zzylol
zzylol force-pushed the rqe-optimizer/hydra branch from 2493109 to b03bd49 Compare October 10, 2026 12:25
zzylol added a commit that referenced this pull request Oct 10, 2026
… for the §6.3 eval

# Conflicts:
#	aqpbm-core/src/atomic_costs.rs
The runner adds the method `asap-norollup` ("ASAP (no roll-ups)"): ASAP's
MILP, candidates, frontier sweep and SLA grid, with each RQE allowed only
deployments at its own grouping or Hydra grids. Every result counts
`rolled_up_rqes` (RQEs served by a finer non-Hydra deployment), and sanity
checks asap <= asap-norollup <= perquery unbounded and at each SLA.

The plot scripts draw it as a dashed violet line, add
fig_rollup_ablation.png (unbounded cost with and without roll-ups, absolute,
CPU and Fargate panels) and a roll-up section in both summaries; results
without the method still plot. The run script passes --no-chosen.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis
zzylol added a commit that referenced this pull request Oct 10, 2026
… for the §6.3 eval

# Conflicts:
#	aqpbm-core/src/atomic_costs.rs
# Conflicts:
#	aqpbm-core/src/atomic_costs.rs
#	docs/rqe_optimizer_candidates.md
#	docs/rqe_optimizer_cost_model.md
#	rqe-optimizer/examples/autosketch_vs_asap.rs
#	rqe-optimizer/src/candidates.rs
@zzylol
zzylol changed the base branch from eval/hydra-workload to main October 11, 2026 17:45
@zzylol
zzylol merged commit 8e38969 into main Oct 11, 2026
2 checks passed
zzylol added a commit that referenced this pull request Oct 11, 2026
…n (§6.3) (#203)

* polars: tag Hydra baseline subset keys by column (#74)

The polars exact baselines keyed a label subset by its `;`-joined values,
asap_sketchlib 0.2.2's subkey. 0.3.0's Hydra tags each value with its
column (`label0:a;label1:b`), so on label columns that share a value
domain the baseline merged a key1 group with the same key2 group and
scored nonzero against itself. subset_key now builds the 0.3.0 key over
hydra_shared's `label{i}` schema (escaping `\`, `:` and `;`), empty
labels keep their column, and every lookup goes through it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* eval: multi-grouping anomaly-detection templates (http, flows)

The generator gets a metric-level ordered label schema (labels with
cardinality, child_of fan-out and skew; per-value data shape) and
per-RQE `grouping` and `covers_share`. The 10 templates are now the
"classic" set, byte-for-byte the committed tables; "all" adds templates
11, 12, 14, 15, 16, 17 on `http` and `flows`, whose streams each carry
several groupings so roll-ups (and later Hydra) can share. The plan runs
both sets over the grid.

The runner builds each schema stream's MetricFacts from the schema
(labels = schema + x; cardinality per used grouping and the full set),
takes Raqe.grouping_labels from the RQE's grouping, maps the
`cardinality` capability to HLL, and reports covers_share and roll-up
deployment groupings in each choice. Classic output is unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* polars: write subset key tags without a temporary String; fix test comment

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* hydra: interleaved merge split; merge exactness and bound tests

`--merge-split contiguous|interleaved` (default contiguous, unchanged) picks
how merge runs split the stream into shards. Interleaved reorders the stream
so the existing contiguous `partition` deals it round-robin by record, at the
contiguous split's shard sizes, so every row's folds take it unchanged.

Tests: hydra-cms, hydra-cs and hydra-hll grids folded from 4 and 7 shards,
under both splits, answer every probed subpopulation exactly as the single
stream; and every scored group on a fixed 20k-record stream lands inside its
hydra doc §2.2 bound (CMS also never under).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* datagen: child_of columns for hierarchical labels

A `child_of: <index>` column with `fan_out: f` nests under an earlier
string column: its value is the parent row's value, a `.`, and a child
index in [0, f) drawn from the column's own distribution. Bad indices,
a fan_out that disagrees with the domain, and fields that would not mean
what they say on a child are refused before drawing.

Adds configs/datagen/hydra_hier.yaml (region 4 -> service 25 -> endpoint
25, status 4, user_id value) for the hydra rows, docs, and tests.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* hydra: score any target grouping; per-group error output

Every hydra-* row (lib and polars, all five variants and hydra-univmon's
four statistics) takes `--group-columns J,K,...` (default `0`, today's
behavior) in place of the hardcoded SCORED_LABEL_COLUMN. A subpopulation
probe is now positioned (`Vec<Option<String>>`, Some at the grouped
columns), so `hydra_shared::query` asks any non-empty subset of the key
columns and the polars baselines look up `subset_key` with that subset's
mask (`group_key`, generalizing `prefix_key`).

Each grouped comparator reports its per-group error in its own metric
(relative error; mean rank error for hydra-kll) and summarises it as
err_mean, err_p50, err_p90, err_max, groups_scored, plus schema_width,
records and fanned_mass. `--per-group-out PATH` writes the rows as
`group_key,n_q,error`.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* datagen: whole uniform bounds for child_of; 2- and 3-label hierarchy cuts

A uniform draw is continuous, so `uniform{0.0, 2.9}` with fan_out 2 drew
child indices 0, 1 and 2; child_of now requires whole uniform bounds.
hydra_hier_d2.yaml and hydra_hier_d3.yaml cut hydra_hier.yaml to its first
2 and 3 labels for the Hydra study's schema-width sweep.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* eval: skewed http labels, per-RQE group coverage, fresh-checkout test

Review of #196. http's labels are now Zipf-skewed (region 0.5, service
1.1, endpoint 1.1 within its service, status 2.0) and the flows DDoS
burst targets a tail subnet, so covers_share selects a few groups. The
generator computes each group's share (product of its labels' shares)
and, per RQE, the covered groups (share >= tau, plus always the largest)
and the smallest covered group's share and items per window. The runner
reads a covered RQE's accuracy (ASAP and AutoSketch) at that group's
items instead of the mean group's; covers_share = null is unchanged.

The schema test reads an inline HLL curve fixture instead of the
gitignored study curves. The synthetic plot keys runs by template set
too, and both plot scripts title classic "mixed" and all "mixed +
multi-grouping". README: the coverage rule, the N the runner uses, and
which schema fields are descriptive. label_set.keys_per_window is
commented as the universe K.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* eval: read null-coverage RQEs at their smallest group; template 17 at {service, endpoint}

- covers_share = null covers every group, so the runner reads its accuracy
  at the smallest group (min_covered_share), like a tau RQE's smallest
  covered group. Classic RQEs (no covers_share) are unchanged.
- Template 17 (negative control) groups by {service, endpoint}: its
  smallest group holds 4.5e4 items in 5m, on the curves.
- The generator asserts every schema RQE's covered_min_N >= 1e3 (drops the
  vacuous `assert covered`); a test checks the old full-label-set template
  fails it.
- Covered facts go through validate_facts.
- README: null semantics, template 17, and the numbers' provenance (fresh
  `all` table, committed saturation inputs).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* datagen: the hydra_hier usage example passes --dtype

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* hydra: list every group in --per-group-out; escape keys; refuse no accuracy

- Group keys escape `\`, `:` and `;` as polars_shared::subset_key does.
- Groups that cannot be scored (zero true statistic, e.g. zero entropy) are
  written with an empty `error`, so n_q sums to the stream's records.
- --per-group-out with no accuracy measurement is refused instead of
  writing an empty file.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* hydra: KLL CDF and UnivMon sum rows; KLL cell merge buffer

- hydra-kll-cdf (lib, polars): asks the KLL grid HydraQuery::Cdf(x) at each
  group's own values at phi = 0.01..0.99, scored by absolute CDF error
  (SubpopCdfErrorGT, comparator subpop-cdf-error). The polars baseline is the
  exact share at or below x over the same sorted groups.
- hydra-univmon-sum (lib, polars): each record inserted with count = value,
  so the UnivMon grid's L1 is the group's sum, scored by relative error
  against exact per-group sums (SubpopSumGT, comparator subpop-sum).
  Integer value dtypes only; f64/str and values outside [0, i32::MAX] are
  refused by name.
- memory_hydra_kll counts each cell as kll_lib_bytes does, merge buffer of
  cell_k items included.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* hydra-kll: the footprint's merge buffer is the steady state; tie tests

The grid clones its cells from one template, and a cloned Vec keeps no
capacity, so a cell's merge buffer is empty at build and grows to about
cell_k items at its first compaction. Say the footprint counts it at that
bound. The CDF test now asks a group of ties (at or below) and a mixed-tie
group on a grid wide enough that it shares no cell.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* hydra: record merge_split; review fixes

- `merge_split` (contiguous|interleaved) is a record field next to
  `merge_shards` (BenchSection and MergeMetrics; schema regenerated).
  Merge records stamp it; a record without it split contiguously, so
  existing results parse, and flatten refuses two different splits in
  one row as it does two shard counts.
- `MergeSplit` moves to aqpbm-core (it is a record type) and is parsed by
  clap directly instead of matching strings.
- The interleaved copy of the stream is built only when a merge runs.
- README states the round-robin exactly (contiguous shard sizes, so a
  short last shard shifts the pattern); MergeMetrics docs name the split.
- hydra-cs bound test: note its margin (passes from beta ~4.8).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* study: --phase hydra on the eval's datasets; Hydra cost rows

configs/datagen/hydra_http.yaml, hydra_http_latency.yaml and hydra_flows.yaml
follow the §6.3 eval's http and flows schemas (labels in schema order,
cardinalities, child_of, skews; user_id Zipf 0.8, latency Pareto a = 2,
src_ip uniform). The flows DDoS burst is not generated: datagen draws each
column independently.

--phase hydra runs HYDRA_SWEEPS: the eval's variants on those datasets (every
grouping, W 1024/4096/16384, N 1e5..1e7, 1/4/16 interleaved shards, 3 seeds)
and every statistic on hydra_hier{_d2,_d3,}.yaml, one run per grouping with
--per-group-out, into hydra_saturation.csv with the coverage columns
err_max_cov_0.01/0.05 (share >= tau plus the largest group). --jobs under a
memory budget, --resume, and --hydra-shard i/n to split it across nodes.

--phase optimizer-cost adds Hydra rows (eval variants x W x dataset);
MeasuredAt gains dataset and schema_width, the reducer reading the width off
the row's scores.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* hydra-univmon-sum: the ground truth refuses negative values too

The lib row inserts each value as a UnivMon count and refuses negatives;
the truth (and so the polars baseline) accepted them, so a spec with
negatives ran on polars and failed on lib.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* study: Hydra cost rows at 3 runs; their test; datagen path from the script

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* docs: Hydra cost rows' run count

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* study_saturation: atomic hydra table rewrite, stricter resume, schema checks

--phase hydra --resume now writes the kept rows to hydra_saturation.csv.tmp
and os.replace()s it over the table, and keeps a row only when it ends in a
line end and every column parses (integers, numbers or "" for the error
columns, non-empty text). The CSV format is unchanged, so tables written by
003eb31 resume as before. The eval-spec test also checks each label's
cardinality, skew and child_of against PR #196's SCHEMAS.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* study_saturation: --hydra-variants/--hydra-ws narrow the Hydra cost rows too

A hydra-univmon-cardinality grid at W=16384 merged from 16 shards
outgrows a 251 GB node, and the cost phase can't resume; these flags let
the cost phase leave a variant or width out, as they do the hydra phase.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* eval: hold covered RQEs to the worse of their smallest and largest group

Per sketch-bench#189, KLL/DD/HLL error does not fall as N grows, so the
largest covered group can be harder than the smallest. The generator also
emits max_covered_share / covered_max_N (the largest group, always covered);
the runner reads each covered RQE's accuracy at both the smallest and the
largest covered group and keeps the worse (direction-aware), unknown if
either read is. Classic RQEs and Hydra's lookup are unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* aqpbm-datagen: scale_by/scale_range; hydra_http_latency scales latency per service

A f64 column with `scale_by: <label column>` and `scale_range: [lo, hi]`
multiplies each row by a factor fixed per label value, log-uniform in
[lo, hi] from a hash of the value and the column's seed. Validation:
earlier string label column, f64 only, 0 < lo <= hi. Absent fields keep
every existing spec's output byte-identical.

hydra_http_latency.yaml scales latency by service in [1, 10], so groups
differ in latency scale and a Hydra-KLL mixture is no longer the queried
group's own CDF.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* eval: template 12 groups by (dst_subnet, proto), not (dst_subnet, dst_port)

(dst_subnet, dst_port) has 1e6 groups; answering each every minute takes
about 49 s on one core, which set every plan's batch latency above the
largest SLA (10 s) and collapsed the multi-grouping frontier. (dst_subnet,
proto) has 3e3 groups and keeps three groupings over one stream.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* rqe-optimizer: admit Hydra (HLL, UnivMon cardinality, KLL) as candidates, with measured accuracy by coverage

- FamilyProperties for hydra-hll, hydra-univmon-cardinality and hydra-kll:
  one fixed-size grid, DeltaSet key tracker, no roll-up, and the new
  answers_any_subgrouping. Added to Capability::families, not to
  DEPLOYABLE_FAMILIES.
- Candidates: on a metric with a hydra_dataset, one grid per schema Λ (each
  RAQE grouping and the union, where card(Λ) is known); eligible for any
  RAQE with ∅ ≠ G_r ⊆ Λ. The subpopulations ceiling is gone.
- Accuracy: hydra_saturation.csv under the saturation dir (optional), keyed
  by variant, config, dataset and grouping; the column the RAQE's
  accuracy_covers_share selects; worst over records and seeds; lossy merges
  bracketed by merge_shards, unknown past the largest.
- Runner: ASAP and PerQuery plan with undeployable families; RQEs carry
  covers_share; http/flows streams name their Hydra dataset; results report
  hydra_rqes for workloads Hydra may serve.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* rqe-optimizer: Hydra grids at the full schema from their dataset's rows

Round-1 review of #201:
- One Hydra grid per (stream, variant, config, window, slide), over the
  metric's full schema only (MetricFacts::hydra_dataset now carries the
  dataset and the schema), the width the study measures.
- MeasuredAt gains optional `dataset` and `schema_width`; a Hydra row is a
  candidate on a metric only if measured on its dataset.
- Runner: card(full schema) = product of label fan-outs, capped at the
  stream's series; AutoSketch also allows undeployable families (it skips
  Hydra itself, #159).
- read_csv checks required columns and short rows; config params must be
  numbers.
- Docs: Hydra held to the worst case, built only at the full schema.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* rqe-optimizer: validate the Hydra schema, skip empty Hydra cells, reject long CSV rows

validate_facts reports a Hydra schema outside the metric's labels or
without a nonzero cardinality, instead of a later panic. The Hydra
accuracy fold skips empty cells (no covered group scored in that run)
and is unknown only when every row left the column empty. read_csv
rejects rows longer than the header too.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* eval: test fixtures carry max_covered_share and label cardinalities

After rebasing onto #196's two-ended coverage read: #201's runner test RQEs
need max_covered_share, and #196's coverage test schema needs the label
cardinality the full-schema Hydra grid reads.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* eval: roll-up ablation, ASAP without roll-ups

The runner adds the method `asap-norollup` ("ASAP (no roll-ups)"): ASAP's
MILP, candidates, frontier sweep and SLA grid, with each RQE allowed only
deployments at its own grouping or Hydra grids. Every result counts
`rolled_up_rqes` (RQEs served by a finer non-Hydra deployment), and sanity
checks asap <= asap-norollup <= perquery unbounded and at each SLA.

The plot scripts draw it as a dashed violet line, add
fig_rollup_ablation.png (unbounded cost with and without roll-ups, absolute,
CPU and Fargate panels) and a roll-up section in both summaries; results
without the method still plot. The run script passes --no-chosen.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* eval: AutoSketch vs. ASAP with Hydra candidates and a roll-up ablation

Results of the §6.3 synthetic eval (ProjectASAP/ASAPQuery#777) rerun with
roll-ups (#190) and their ablation, Hydra candidates measured on the eval's
own data (#195–#202), the multi-grouping template set (#196), #192's KLL
footprint and a re-measured cost table. 8 workloads, 0 RQEs dropped, 0
sanity violations.

Roll-ups serve 40% of the multi-grouping RQEs from a finer deployment and
halve ASAP's deployments, but save only 3.3–3.6% (CPU) / 4.4–4.8%
(Fargate). Hydra is eligible for many RQEs but never chosen: its insert
fans out to every label subset (1.9–5.1 µs per record vs. 3.85 ns for a
per-group HLL).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* eval: cube template and a multi-grouping-only workload

Template 16 is now p99 latency GROUP BY CUBE(region, service, endpoint),
all 7 groupings, replacing {service}, {region, service}, {service, status}.
A new `multigroup` template set runs the multi-grouping templates alone, so
the roll-up ablation is not diluted by the classic set.

Roll-ups now save 60% on multi-grouping (0.551 -> 0.221 vCPU, 26 -> 5
deployments) and 8% on mixed + multi-grouping. Classic is unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* eval: each workload's chosen plans, gzipped

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant