Skip to content

eval: multi-grouping anomaly-detection templates (http, flows) for Hydra and roll-ups - #196

Open
zzylol wants to merge 5 commits into
mainfrom
eval/hydra-workload
Open

zzylol wants to merge 5 commits into
mainfrom
eval/hydra-workload

Conversation

@zzylol

@zzylol zzylol commented Oct 10, 2026 •

Copy link
Copy Markdown
Collaborator

What

  • Generator (scripts/export_autosketch_eval_table.py): a metric-level ordered label schema (SCHEMAS: labels with cardinality, child_of fan-out and Zipf skew; each value's data shape) and per-RQE grouping (label list) and covers_share (null = every group). One stream carries several groupings. The 10 templates are now the classic set; all = classic + templates 11, 12, 14, 15, 16, 17 on metrics http and flows, as the Hydra spec describes; template 17 (negative control: every group) groups by {service, endpoint}, whose smallest group holds 4.5e4 items in 5m. The tables carry schemas per workload. synthetic_plan() runs both sets over the grid (8 tables).
  • Runner (autosketch_vs_asap.rs): for a schema stream, MetricFacts.labels = schema labels + x, cardinality for every used grouping and the full set (= series), data shape per grouping from the families' grid_param/grid_K; Raqe.grouping_labels from the RQE's grouping; capability cardinality maps to HLL. covers_share stays in the runner (Workload::covers_share), not on Raqe, and is reported in each choice, as is deployment_grouping when a deployment is finer than its RQE (a roll-up). Classic tables have neither field, so their output is unchanged.
  • Coverage (review round 1, spec amendment): http labels are Zipf-skewed (region 0.5, service 1.1, endpoint 1.1 within its service, status 2.0) and the flows DDoS burst targets a tail subnet (999), so template 11 covers it. group_coverage() computes each group's share (product of its labels' shares; labels are independent given their parents) and the rule "share >= τ, plus always the largest group"; each schema RQE carries covered_groups, min_covered_share and covered_min_N; the generator fails if any covered_min_N is below the curves' first N (1e3). The runner reads a covered RQE's accuracy (ASAP and AutoSketch) at the smallest covered group's items per window, by scaling the series count in a per-RQE copy of its stream's facts (only items per group reads it); covers_share = null covers every group, so it is read at the smallest group of all (review round 2). Classic RQEs (no covers_share) keep the mean group's N. The per-RQE facts also go through validate_facts. The counted value's skew is unchanged.
  • Worst covered group (sketch-bench#189: KLL/DD/HLL error does not fall as N grows, so the largest group can be the hardest): the generator also writes max_covered_share and covered_max_N (the largest group, always covered). The runner builds a second per-RQE facts copy at the largest covered group and holds ASAP, PerQuery and AutoSketch to the worse of the two reads in the family's direction (Workload::worse_read), unknown if either read is. Classic RQEs and Hydra's lookup are untouched; rqe-optimizer/src is unchanged.
  • Plots: the synthetic script keys runs by (template set, shared, metrics); both scripts title classic "mixed" and all "mixed + multi-grouping".
  • Results README: a "Template sets" section with the schema, rates, per-template groups, mean and smallest-covered items, which N the runner uses, and which schema fields are descriptive. label_set.keys_per_window is commented as the universe K.

Why

The 10 templates read each stream at one grouping, so roll-ups (#190) never trigger and Hydra has nothing to share. These anomaly-detection templates read one stream at many groupings.

Numbers: http region 4 x service 25 x endpoint 25/service x status 4 (1e4 combos, skewed as above), 2e6 samples/s like classic, user_id Zipf 0.8 over 1e6, latency Pareto a = 2. flows dst_subnet 1e3 (Zipf 1.1) x dst_port 1e3 (1.2) x proto 3, src_ip uniform over 1e6 plus a 5% DDoS burst from 1e4 sources at subnet 999; 4e6 samples/s (twice classic) so {dst_subnet, dst_port} (1e6 groups) holds 1,200 items per group in 5m, above the HLL curves' first N (1e3). Windows 5m and 15m, T = 1m: 54 new RQEs, 104 in all.

Tests

  • python3 scripts/test_export_autosketch_eval_table.py (25 pass; new: a template whose smallest group is below 1e3 items (17 at the full label set) fails generation; label shares sum to 1 and the burst lifts subnet 999 to >= 5%; coverage equals a brute force over {region, status} at τ = 0.01/0.05/0.2, is the largest group alone at τ = 0.5, all groups at null; every schema RQE has >= 1 covered group and covered_min_N >= 1e3): classic tables equal the committed templatesall tables (main's) byte for byte after the rename; all = classic + 54; child labels count with their parents (endpoint 625); dropping labels never adds groups; every grouping has a series per group and >= 1e3 items in 5m; groupings are ordered, non-empty schema subsets (15 for template 15); entries carry grouping, coverage and shape; metric copies name data_i/http. Mutation-checked (no child closure; all without the new templates; grouping on classic entries; fixed shape): each fails a test.
  • cargo test -p rqe-optimizer --example autosketch_vs_asap: new a_schema_stream_carries_every_grouping_and_rolls_up (two HLL RQEs, {region} and {region, service}: facts as above, and ASAP serves both from one {region, service} deployment). It now builds a temp saturation dir from the committed cost table and inline HLL curves, so it passes in a fresh checkout without the gitignored curves; it also checks the covered RQE's facts are scaled to 5% of the stream, and a null-coverage RQE's to its smallest group (each fails by mutation: scaling removed, null guard restored) and fails if the runner ignores grouping.
  • Worst group: a_covered_rqe_is_held_to_the_worse_of_its_smallest_and_largest_group (RQE covering a 1% and a 50% group): with error rising in N, the smallest group meets a 3% target and the largest does not, and ASAP and AutoSketch read the largest; with error falling in N, the reverse. A classic RQE reads the mean group as before. Mutation: always smallest / always largest each fail it. Python: group_coverage returns the largest share, and every schema RQE has covered_max_N = rate·window·max_covered_share.
  • cargo fmt --check, cargo clippy -p rqe-optimizer --all-targets -D warnings, cargo test -p rqe-optimizer, the example's tests and the Python tests: clean, in a fresh worktree without the curve CSVs.

Evidence

Inputs: #194's curves + the committed cost table; --runs 1.

  • classic, r = 1: main's runner on main's table vs this branch's runner on the new templatesclassic table: outputs identical apart from timings and the workload name (50 RQEs, 0 dropped).

  • all, r = 1, p95, freshly generated (--synthetic) with the committed saturation inputs (rqe-optimizer/results/autosketch-vs-asap-inputs/saturation; the committed templatesall tables are classic): 104 RQEs, 0 dropped, no sanity violations. CPU weights: ASAP 5.05 vCPU (15 active deployments), PerQuery 11.8, AutoSketch 3.42e3. ASAP serves 40 of the 54 new RQEs by roll-up: templates 14, 15 and 17 from HLL at {region, service, status} or the full label set, {dst_port} from {dst_port, proto}, p99 by {service} from DDSketch at {region, service} or {service, status}. AutoSketch and PerQuery: 0 roll-ups.

  • Worst group (e7d8a03 vs 561ee86, all shared1, scratch saturation inputs without Hydra data): 104 RQEs, 0 dropped in both. Costs unchanged: CPU weights ASAP 5.779 (15 deployments), PerQuery 13.14, AutoSketch 3575; Fargate 0.6114 / 1.247 / 169. Same chosen configs for all three methods; only two DDSketch p99 RQEs' read rises (0.009645 to 0.009650, target 0.05). classic shared1: identical apart from timings.

The committed results were not rerun; their templatesall files are today's classic, so the plot scripts would now title them "mixed + multi-grouping" until they are regenerated under templatesclassic.

🤖 Generated with Claude Code

https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

zzylol and others added 3 commits October 10, 2026 04:02
The generator gets a metric-level ordered label schema (labels with
cardinality, child_of fan-out and skew; per-value data shape) and
per-RQE `grouping` and `covers_share`. The 10 templates are now the
"classic" set, byte-for-byte the committed tables; "all" adds templates
11, 12, 14, 15, 16, 17 on `http` and `flows`, whose streams each carry
several groupings so roll-ups (and later Hydra) can share. The plan runs
both sets over the grid.

The runner builds each schema stream's MetricFacts from the schema
(labels = schema + x; cardinality per used grouping and the full set),
takes Raqe.grouping_labels from the RQE's grouping, maps the
`cardinality` capability to HLL, and reports covers_share and roll-up
deployment groupings in each choice. Classic output is unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis
Review of #196. http's labels are now Zipf-skewed (region 0.5, service
1.1, endpoint 1.1 within its service, status 2.0) and the flows DDoS
burst targets a tail subnet, so covers_share selects a few groups. The
generator computes each group's share (product of its labels' shares)
and, per RQE, the covered groups (share >= tau, plus always the largest)
and the smallest covered group's share and items per window. The runner
reads a covered RQE's accuracy (ASAP and AutoSketch) at that group's
items instead of the mean group's; covers_share = null is unchanged.

The schema test reads an inline HLL curve fixture instead of the
gitignored study curves. The synthetic plot keys runs by template set
too, and both plot scripts title classic "mixed" and all "mixed +
multi-grouping". README: the coverage rule, the N the runner uses, and
which schema fields are descriptive. label_set.keys_per_window is
commented as the universe K.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis
… {service, endpoint}

- covers_share = null covers every group, so the runner reads its accuracy
  at the smallest group (min_covered_share), like a tau RQE's smallest
  covered group. Classic RQEs (no covers_share) are unchanged.
- Template 17 (negative control) groups by {service, endpoint}: its
  smallest group holds 4.5e4 items in 5m, on the curves.
- The generator asserts every schema RQE's covered_min_N >= 1e3 (drops the
  vacuous `assert covered`); a test checks the old full-label-set template
  fails it.
- Covered facts go through validate_facts.
- README: null semantics, template 17, and the numbers' provenance (fresh
  `all` table, committed saturation inputs).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis
zzylol added a commit that referenced this pull request Oct 10, 2026
… checks

--phase hydra --resume now writes the kept rows to hydra_saturation.csv.tmp
and os.replace()s it over the table, and keeps a row only when it ends in a
line end and every column parses (integers, numbers or "" for the error
columns, non-empty text). The CSV format is unchanged, so tables written by
003eb31 resume as before. The eval-spec test also checks each label's
cardinality, skew and child_of against PR #196's SCHEMAS.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis
Per sketch-bench#189, KLL/DD/HLL error does not fall as N grows, so the
largest covered group can be harder than the smallest. The generator also
emits max_covered_share / covered_max_N (the largest group, always covered);
the runner reads each covered RQE's accuracy at both the smallest and the
largest covered group and keeps the worse (direction-aware), unknown if
either read is. Classic RQEs and Hydra's lookup are unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis
zzylol added a commit that referenced this pull request Oct 10, 2026
After rebasing onto #196's two-ended coverage read: #201's runner test RQEs
need max_covered_share, and #196's coverage test schema needs the label
cardinality the full-schema Hydra grid reads.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis
…_port)

(dst_subnet, dst_port) has 1e6 groups; answering each every minute takes
about 49 s on one core, which set every plan's batch latency above the
largest SLA (10 s) and collapsed the multi-grouping frontier. (dst_subnet,
proto) has 3e3 groups and keeps three groupings over one stream.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis
zzylol added a commit that referenced this pull request Oct 10, 2026
After rebasing onto #196's two-ended coverage read: #201's runner test RQEs
need max_covered_share, and #196's coverage test schema needs the label
cardinality the full-schema Hydra grid reads.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis
zzylol added a commit that referenced this pull request Oct 11, 2026
* polars: tag Hydra baseline subset keys by column (#74)

The polars exact baselines keyed a label subset by its `;`-joined values,
asap_sketchlib 0.2.2's subkey. 0.3.0's Hydra tags each value with its
column (`label0:a;label1:b`), so on label columns that share a value
domain the baseline merged a key1 group with the same key2 group and
scored nonzero against itself. subset_key now builds the 0.3.0 key over
hydra_shared's `label{i}` schema (escaping `\`, `:` and `;`), empty
labels keep their column, and every lookup goes through it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* polars: write subset key tags without a temporary String; fix test comment

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* hydra: interleaved merge split; merge exactness and bound tests

`--merge-split contiguous|interleaved` (default contiguous, unchanged) picks
how merge runs split the stream into shards. Interleaved reorders the stream
so the existing contiguous `partition` deals it round-robin by record, at the
contiguous split's shard sizes, so every row's folds take it unchanged.

Tests: hydra-cms, hydra-cs and hydra-hll grids folded from 4 and 7 shards,
under both splits, answer every probed subpopulation exactly as the single
stream; and every scored group on a fixed 20k-record stream lands inside its
hydra doc §2.2 bound (CMS also never under).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* datagen: child_of columns for hierarchical labels

A `child_of: <index>` column with `fan_out: f` nests under an earlier
string column: its value is the parent row's value, a `.`, and a child
index in [0, f) drawn from the column's own distribution. Bad indices,
a fan_out that disagrees with the domain, and fields that would not mean
what they say on a child are refused before drawing.

Adds configs/datagen/hydra_hier.yaml (region 4 -> service 25 -> endpoint
25, status 4, user_id value) for the hydra rows, docs, and tests.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* hydra: score any target grouping; per-group error output

Every hydra-* row (lib and polars, all five variants and hydra-univmon's
four statistics) takes `--group-columns J,K,...` (default `0`, today's
behavior) in place of the hardcoded SCORED_LABEL_COLUMN. A subpopulation
probe is now positioned (`Vec<Option<String>>`, Some at the grouped
columns), so `hydra_shared::query` asks any non-empty subset of the key
columns and the polars baselines look up `subset_key` with that subset's
mask (`group_key`, generalizing `prefix_key`).

Each grouped comparator reports its per-group error in its own metric
(relative error; mean rank error for hydra-kll) and summarises it as
err_mean, err_p50, err_p90, err_max, groups_scored, plus schema_width,
records and fanned_mass. `--per-group-out PATH` writes the rows as
`group_key,n_q,error`.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* datagen: whole uniform bounds for child_of; 2- and 3-label hierarchy cuts

A uniform draw is continuous, so `uniform{0.0, 2.9}` with fan_out 2 drew
child indices 0, 1 and 2; child_of now requires whole uniform bounds.
hydra_hier_d2.yaml and hydra_hier_d3.yaml cut hydra_hier.yaml to its first
2 and 3 labels for the Hydra study's schema-width sweep.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* datagen: the hydra_hier usage example passes --dtype

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* hydra: list every group in --per-group-out; escape keys; refuse no accuracy

- Group keys escape `\`, `:` and `;` as polars_shared::subset_key does.
- Groups that cannot be scored (zero true statistic, e.g. zero entropy) are
  written with an empty `error`, so n_q sums to the stream's records.
- --per-group-out with no accuracy measurement is refused instead of
  writing an empty file.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* hydra: KLL CDF and UnivMon sum rows; KLL cell merge buffer

- hydra-kll-cdf (lib, polars): asks the KLL grid HydraQuery::Cdf(x) at each
  group's own values at phi = 0.01..0.99, scored by absolute CDF error
  (SubpopCdfErrorGT, comparator subpop-cdf-error). The polars baseline is the
  exact share at or below x over the same sorted groups.
- hydra-univmon-sum (lib, polars): each record inserted with count = value,
  so the UnivMon grid's L1 is the group's sum, scored by relative error
  against exact per-group sums (SubpopSumGT, comparator subpop-sum).
  Integer value dtypes only; f64/str and values outside [0, i32::MAX] are
  refused by name.
- memory_hydra_kll counts each cell as kll_lib_bytes does, merge buffer of
  cell_k items included.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* hydra-kll: the footprint's merge buffer is the steady state; tie tests

The grid clones its cells from one template, and a cloned Vec keeps no
capacity, so a cell's merge buffer is empty at build and grows to about
cell_k items at its first compaction. Say the footprint counts it at that
bound. The CDF test now asks a group of ties (at or below) and a mixed-tie
group on a grid wide enough that it shares no cell.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* hydra: record merge_split; review fixes

- `merge_split` (contiguous|interleaved) is a record field next to
  `merge_shards` (BenchSection and MergeMetrics; schema regenerated).
  Merge records stamp it; a record without it split contiguously, so
  existing results parse, and flatten refuses two different splits in
  one row as it does two shard counts.
- `MergeSplit` moves to aqpbm-core (it is a record type) and is parsed by
  clap directly instead of matching strings.
- The interleaved copy of the stream is built only when a merge runs.
- README states the round-robin exactly (contiguous shard sizes, so a
  short last shard shifts the pattern); MergeMetrics docs name the split.
- hydra-cs bound test: note its margin (passes from beta ~4.8).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* study: --phase hydra on the eval's datasets; Hydra cost rows

configs/datagen/hydra_http.yaml, hydra_http_latency.yaml and hydra_flows.yaml
follow the §6.3 eval's http and flows schemas (labels in schema order,
cardinalities, child_of, skews; user_id Zipf 0.8, latency Pareto a = 2,
src_ip uniform). The flows DDoS burst is not generated: datagen draws each
column independently.

--phase hydra runs HYDRA_SWEEPS: the eval's variants on those datasets (every
grouping, W 1024/4096/16384, N 1e5..1e7, 1/4/16 interleaved shards, 3 seeds)
and every statistic on hydra_hier{_d2,_d3,}.yaml, one run per grouping with
--per-group-out, into hydra_saturation.csv with the coverage columns
err_max_cov_0.01/0.05 (share >= tau plus the largest group). --jobs under a
memory budget, --resume, and --hydra-shard i/n to split it across nodes.

--phase optimizer-cost adds Hydra rows (eval variants x W x dataset);
MeasuredAt gains dataset and schema_width, the reducer reading the width off
the row's scores.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* hydra-univmon-sum: the ground truth refuses negative values too

The lib row inserts each value as a UnivMon count and refuses negatives;
the truth (and so the polars baseline) accepted them, so a spec with
negatives ran on polars and failed on lib.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* study: Hydra cost rows at 3 runs; their test; datagen path from the script

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* docs: Hydra cost rows' run count

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* study_saturation: atomic hydra table rewrite, stricter resume, schema checks

--phase hydra --resume now writes the kept rows to hydra_saturation.csv.tmp
and os.replace()s it over the table, and keeps a row only when it ends in a
line end and every column parses (integers, numbers or "" for the error
columns, non-empty text). The CSV format is unchanged, so tables written by
003eb31 resume as before. The eval-spec test also checks each label's
cardinality, skew and child_of against PR #196's SCHEMAS.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* study_saturation: --hydra-variants/--hydra-ws narrow the Hydra cost rows too

A hydra-univmon-cardinality grid at W=16384 merged from 16 shards
outgrows a 251 GB node, and the cost phase can't resume; these flags let
the cost phase leave a variant or width out, as they do the hydra phase.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* aqpbm-datagen: scale_by/scale_range; hydra_http_latency scales latency per service

A f64 column with `scale_by: <label column>` and `scale_range: [lo, hi]`
multiplies each row by a factor fixed per label value, log-uniform in
[lo, hi] from a hash of the value and the column's seed. Validation:
earlier string label column, f64 only, 0 < lo <= hi. Absent fields keep
every existing spec's output byte-identical.

hydra_http_latency.yaml scales latency by service in [1, 10], so groups
differ in latency scale and a Hydra-KLL mixture is no longer the queried
group's own CDF.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant