Skip to content

docs(planner): plan the AutoSketch vs. ASAPQuery planner evaluation (§6.3) - #777

Open
zzylol wants to merge 41 commits into
mainfrom
docs/autosketch-vs-planner-eval
Open

zzylol wants to merge 41 commits into
mainfrom
docs/autosketch-vs-planner-eval

Conversation

@zzylol

@zzylol zzylol commented Oct 4, 2026 •

Copy link
Copy Markdown
Contributor

Before this PR

The §6.3 comparison against AutoSketch had no written plan. The only AutoSketch implementation was an Algorithm 4 adaptation in ASAPQuery-backend (#545 protocol, #547 runner). It is execution-based, CMS-only, and not connected to the RQE MILP that sketch-bench #129 added.

After this PR

docs/evaluation/autosketch-vs-planner.md defines the planner-level comparison, run offline on sketch-bench's rqe-optimizer from sketch-bench measurements.

Validation

Documentation only. Factual claims are checked against sketch-bench (rqe-optimizer at #141's head, the saturation data) and ASAPQuery-backend autosketch_comparison.rs.

🤖 Generated with Claude Code

zzylol and others added 5 commits October 4, 2026 21:14
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…cisions

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…s for KLL/top-k

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…n the whole dataset

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
zzylol and others added 2 commits October 5, 2026 02:16
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
zzylol and others added 4 commits October 5, 2026 13:11
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
zzylol added a commit to ProjectASAP/sketch-bench that referenced this pull request Oct 5, 2026
…anner workload

Measure frequency and top-k saturation curves at K = 10, 100, 1e4 and 1e6
(theta 0..2, N up to 1e8, 3 seeds) and top-k accuracy after merging 4/16/64
shards, for the synthetic PromQL workload of ProjectASAP/ASAPQuery#777.
Commit the summary CSVs, the N_sat tables and the scripts that finish and
check an interrupted cost phase. Cost results are pending.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
zzylol added a commit to ProjectASAP/sketch-bench that referenced this pull request Oct 5, 2026
* rqe-optimizer: price plans per EC2 machine family

Score retained memory per active deployment, (x + max S) / y instances,
and add milp::minimize_cost, which minimizes the hourly price on one EC2
machine family in fractional instances (whichever of CPU or memory binds).
Prices are a committed on-demand snapshot from scripts/fetch_ec2_pricing.py.
Dominance pruning also compares retained memory so it cannot drop a
candidate that is cheaper under the new objective.

Part of ProjectASAP/ASAPQuery#777.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* rqe-optimizer: label small_problem MILP output for both objectives

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* rqe-optimizer: scale the MILP so tiny costs and latencies solve exactly

HiGHS's absolute gap (1e-6) and feasibility tolerance (1e-7) exceed
real plan costs (~1e-6 $/hour) and query latencies (µs). Rows are now
scaled by the plan's demand without sharing, and per-RQE memory and
latency bounds become f64 exclusions of the choices over them.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
zzylol and others added 6 commits October 5, 2026 10:43
…tatus

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ties

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
zzylol added a commit to ProjectASAP/sketch-bench that referenced this pull request Oct 5, 2026
The user dropped FewestPlans from the evaluation (ProjectASAP/ASAPQuery#777).
Remove milp::minimize_deployments and MilpBounds::max_active_deployments,
FewestPlans from the runner and its sanity checks, and its column from the
plots and summaries. milp.rs and small_problem.rs are back to #138's versions.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
zzylol and others added 2 commits October 5, 2026 15:27
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
zzylol added a commit to ProjectASAP/sketch-bench that referenced this pull request Oct 5, 2026
Runner (examples/autosketch_vs_asap.rs) for the traces workloads of
ProjectASAP/ASAPQuery#777: ASAP (joint MILP with one absolute latency SLA
per sweep point), PerQuery-CostAware and AutoSketch-Adapted, scored by
objectives::score and priced per EC2 family, with sanity checks.

Accuracy now depends on the query's merge count m = S/x: a cost entry may
carry "{metric}@m{b}" keys, eligibility reads the smallest measured b >= m,
and a merge larger than every measured b is ineligible.

Results for Alibaba 2022, Google 2011 and BOOM with figures and summary.
The earlier example and scaling workloads are dropped (#777 §6).

Rebased onto main, which now has #137, #135 and #136.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
zzylol added a commit to ProjectASAP/sketch-bench that referenced this pull request Oct 5, 2026
…nd PerQuery

- Runner: a synthetic subcommand reading a table from
  export_autosketch_eval_table.py --synthetic (--target, --replicas,
  --no-chosen) and a wider synthetic SLA grid. PerQuery-CostAware now
  solves each RQE under the same SLA as ASAP, with a sanity check that
  ASAP costs no more.
- export_autosketch_eval_table.py --synthetic: the 10 PromQL templates of
  ProjectASAP/ASAPQuery#777 §6 as RQEs, with per-RQE accuracy lookups.
- run/plot scripts for the synthetic sweep.

Rebased onto #138 (f293b73), which dropped the example and scaling
workloads; FewestPlans is not included.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
zzylol added a commit to ProjectASAP/sketch-bench that referenced this pull request Oct 6, 2026
ProjectASAP/ASAPQuery#777 section 6:
- templates=dashboard (D1-D5, 33 RQEs): per-series p50/p90/p99 over
  {1m, 5m, 15m, 1h, 6h, 24h}, p50/p90/p99 by label_0, sum by label_0 of
  rate over the six windows, top-3 by label_0 of rate over 5m and 1h, and
  the p99/p50 ratio over 5m and 1h as two quantile RQEs on D1's stream.
- Shared replicas (table dimension shared=r): r replicas of a template set
  on the same streams; replica i draws 3 of the set's windows, 3 of the five
  quantiles and an interval (10/60/300 s) from a seeded RNG, and identical
  RQEs are kept once (temporal ranges carry their interval, e.g. 1h@60s).
  The runner's --replicas stays the disjoint control.
- The plan adds the dashboard and shared r in {1..64} for the default mix
  and the dashboard, at every strictness level (111 runs in all). The 40
  existing tables are unchanged.
- Summaries: dashboard tables, a shared-replicas table, and
  fig_synthetic_shared_replicas.png (cost relative to ASAP vs. r).

The runner is unchanged since ce4b885.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@zzylol
zzylol marked this pull request as ready for review October 6, 2026 19:59
Comment thread docs/evaluation/autosketch-vs-planner.md Outdated
Comment thread docs/evaluation/autosketch-vs-planner.md Outdated
Comment thread docs/evaluation/autosketch-vs-planner.md Outdated
Comment thread docs/evaluation/autosketch-vs-planner.md Outdated
Comment thread docs/evaluation/autosketch-vs-planner.md Outdated
Comment thread docs/evaluation/autosketch-vs-planner.md Outdated
Comment thread docs/evaluation/autosketch-vs-planner.md Outdated
zzylol and others added 5 commits October 6, 2026 21:19
…e cost model

Addresses review on #777:
- cost: sketch-bench #145's per-phase CPU and memory with a weighted
  objective; CPU only first, then Fargate prices. EC2 machine families,
  the peak-provisioned model and the CPU timeline are dropped.
- capabilities: sum and rate/increase map to the exact accumulators
  (#144), not frequency; strictness applies to quantile and top-k only.
- inputs: read the export_rqe_optimizer_costs table (measured_at, merge
  accuracy, size sweep) instead of the saturation tables; exact
  accumulators are benchmarked for cost but need no saturation point.
- grid: keep template set, shared replicas {1, 8, 64}, C {1e2, 1e3, 1e4},
  strictness and SLA {0.01, 0.1, 1, 10} ms; move the rest to planner
  sensitivity. Disjoint replicas are dropped (spatial filters unsupported).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… workload

Top-k, KLL and DDSketch are measured on Zipf s=1.1 over 100k keys (quantiles
on the Zipf ranks). The synthetic workload now uses that distribution instead
of Zipf θ=1.0 and Pareto a=2, so the cost table needs no separate run.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- data: C = 1e4 series (label_0, the top-k key), J = 10 jobs, 200
  samples/s per series, Zipf s=1.1 over 10,000 keys (the cost table's
  data at CARDINALITY=10000); fixed, so the comparison does not sweep it.
- windows W = {15m, 1h, 6h, 24h}: every per-series summary sees >= 1e5
  items; per-summary input sizes listed per grouping.
- queries: grouped templates use by (job); top-k is
  topk(k, sum by (label_0) (...)) with k in {100, 200, 300}, served by
  one heap-K deployment for k <= K; quantile syntax fixed. Both template
  sets are enumerated query by query (all checked with promql-parser).
- grid: C removed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…h-bench #156)

Keep the fixed synthetic data (Zipf s=1.1 over 10,000 keys, quantiles on
the Zipf ranks) and top-k with k in {100, 200, 300} served by heap-K
deployments. Everything else follows sketch-bench #156/#162:
- sketch accuracy from the saturation curves at n(S, G); CPU, memory and
  exact rows from the cost table; grid configs only.
- merging: every pane saturated, then read at n(S, G); KLL/top-k merge
  penalty is #158. AutoSketch uses the same lookup with m = 1.
- DDSketch scored in value relative error (#162), with its own targets.
- a targeted saturation run at the workload's data shape (theta 1.1,
  K 1e4, Zipf-rank quantiles, CMS-heap heaps 100/200/300) so lookups land
  on measured points.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
k = 32 is the heap size sketch-bench measures top-k at (CMS_HEAP_TOP_K),
so top-k needs no heap parameter, no k on the RQE and no extra saturation
points. 58 RQEs per replica for the 10 templates (50 distinct), 25 for the
dashboard (21 distinct); query lists regenerated and parser-checked.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
zzylol added a commit to ProjectASAP/sketch-bench that referenced this pull request Oct 7, 2026
ProjectASAP/ASAPQuery#777 (paper §6.3), consolidating #138, #139 and #141.

- rqe-optimizer/examples/autosketch_vs_asap.rs: ASAP (milp::minimize),
  PerQuery-CostAware and AutoSketch-Adapted on one synthetic table. Costs and
  accuracy come from --saturation-dir, as the planner reads them:
  SaturationCurves::accuracy for ASAP (merge curves for KLL and top-k,
  #179), autosketch_accuracy for AutoSketch, and the one-study cost table
  (#178). Each run is scored by analytical_cost_model::score at #777's two
  weight settings and the SLA grid.
- scripts/export_autosketch_eval_table.py --synthetic: one workload-only
  table per grid point, plus plan.tsv. The data constants come from
  study_saturation's COST_*.
- scripts/run_autosketch_vs_asap_synthetic.sh and
  plot_autosketch_vs_asap_synthetic.py: run the grid, then plot.
- scripts/profiling_time.py and the synthetic results README: ASAP's
  one-time profiling time, and how to reproduce.

The trace workloads are left out for now.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
zzylol added a commit to ProjectASAP/sketch-bench that referenced this pull request Oct 7, 2026
… workloads

ProjectASAP/ASAPQuery#777 (paper §6.3), consolidating #138, #139 and #141.

- rqe-optimizer/examples/autosketch_vs_asap.rs (traces and synthetic): ASAP
  (milp::minimize), PerQuery-CostAware and AutoSketch-Adapted. Costs and
  accuracy come from --saturation-dir, as the planner reads them:
  SaturationCurves::accuracy for ASAP (merge curves for KLL and top-k,
  #179), autosketch_accuracy for AutoSketch, and the one-study cost table
  (#178). Each run is scored by analytical_cost_model::score at #777's two
  weight settings and the SLA grid.
- scripts/run_autosketch_vs_asap.sh and plot_autosketch_vs_asap.py: the
  trace workloads alibaba_v2022 and google_2011. boom is left out for now.
- scripts/export_autosketch_eval_table.py --synthetic: one workload-only
  table per grid point, plus plan.tsv. The data constants come from
  study_saturation's COST_*. The traces path is unchanged from main.
- scripts/run_autosketch_vs_asap_synthetic.sh and
  plot_autosketch_vs_asap_synthetic.py: run the synthetic grid, then plot.
- scripts/profiling_time.py and the results README: ASAP's one-time
  profiling time, and how to reproduce.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
zzylol added a commit to ProjectASAP/sketch-bench that referenced this pull request Oct 7, 2026
… workloads

ProjectASAP/ASAPQuery#777 (paper §6.3), consolidating #138, #139 and #141.

- rqe-optimizer/examples/autosketch_vs_asap.rs (traces and synthetic): ASAP
  (milp::minimize), PerQuery-CostAware and AutoSketch-Adapted. Costs and
  accuracy come from --saturation-dir, as the planner reads them:
  SaturationCurves::accuracy for ASAP (merge curves for KLL and top-k,
  #179), autosketch_accuracy for AutoSketch, and the one-study cost table
  (#178). Each run is scored by analytical_cost_model::score at #777's two
  weight settings and the SLA grid.
- scripts/run_autosketch_vs_asap.sh and plot_autosketch_vs_asap.py: the
  trace workloads alibaba_v2022 and google_2011. boom is left out for now.
- scripts/export_autosketch_eval_table.py --synthetic: one workload-only
  table per grid point, plus plan.tsv. The data constants come from
  study_saturation's COST_*. The traces path is unchanged from main.
- scripts/run_autosketch_vs_asap_synthetic.sh and
  plot_autosketch_vs_asap_synthetic.py: run the synthetic grid, then plot.
- scripts/profiling_time.py and the results README: ASAP's one-time
  profiling time, and how to reproduce.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
zzylol added a commit to ProjectASAP/sketch-bench that referenced this pull request Oct 7, 2026
… workloads

ProjectASAP/ASAPQuery#777 (paper §6.3), consolidating #138, #139 and #141.

- rqe-optimizer/examples/autosketch_vs_asap.rs (traces and synthetic): ASAP
  (milp::minimize), PerQuery-CostAware and AutoSketch-Adapted. Costs and
  accuracy come from --saturation-dir, as the planner reads them:
  SaturationCurves::accuracy for ASAP (merge curves for KLL and top-k,
  #179), autosketch_accuracy for AutoSketch, and the one-study cost table
  (#178). Each run is scored by analytical_cost_model::score at #777's two
  weight settings and the SLA grid.
- scripts/run_autosketch_vs_asap.sh and plot_autosketch_vs_asap.py: the
  trace workloads alibaba_v2022 and google_2011. boom is left out for now.
- scripts/export_autosketch_eval_table.py --synthetic: one workload-only
  table per grid point, plus plan.tsv. The data constants come from
  study_saturation's COST_*. The traces path is unchanged from main.
- scripts/run_autosketch_vs_asap_synthetic.sh and
  plot_autosketch_vs_asap_synthetic.py: run the synthetic grid, then plot.
- scripts/profiling_time.py and the results README: ASAP's one-time
  profiling time, and how to reproduce.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
zzylol added a commit to ProjectASAP/sketch-bench that referenced this pull request Oct 7, 2026
… workloads

ProjectASAP/ASAPQuery#777 (paper §6.3), consolidating #138, #139 and #141.

- rqe-optimizer/examples/autosketch_vs_asap.rs (traces and synthetic): ASAP
  (milp::minimize), PerQuery-CostAware and AutoSketch-Adapted. Costs and
  accuracy come from --saturation-dir, as the planner reads them:
  SaturationCurves::accuracy for ASAP (merge curves for KLL and top-k,
  #179), autosketch_accuracy for AutoSketch, and the one-study cost table
  (#178). Each run is scored by analytical_cost_model::score at #777's two
  weight settings and the SLA grid.
- scripts/run_autosketch_vs_asap.sh and plot_autosketch_vs_asap.py: the
  trace workloads alibaba_v2022 and google_2011. boom is left out for now.
- scripts/export_autosketch_eval_table.py --synthetic: one workload-only
  table per grid point, plus plan.tsv. The data constants come from
  study_saturation's COST_*. The traces path is unchanged from main.
- scripts/run_autosketch_vs_asap_synthetic.sh and
  plot_autosketch_vs_asap_synthetic.py: run the synthetic grid, then plot.
- scripts/profiling_time.py and the results README: ASAP's one-time
  profiling time, and how to reproduce.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
zzylol added a commit to ProjectASAP/sketch-bench that referenced this pull request Oct 7, 2026
… workloads

ProjectASAP/ASAPQuery#777 (paper §6.3), consolidating #138, #139 and #141.

- rqe-optimizer/examples/autosketch_vs_asap.rs (traces and synthetic): ASAP
  (milp::minimize), PerQuery-CostAware and AutoSketch-Adapted. Costs and
  accuracy come from --saturation-dir, as the planner reads them:
  SaturationCurves::accuracy for ASAP (merge curves for KLL and top-k,
  #179), autosketch_accuracy for AutoSketch, and the one-study cost table
  (#178). Each run is scored by analytical_cost_model::score at #777's two
  weight settings and the SLA grid.
- scripts/run_autosketch_vs_asap.sh and plot_autosketch_vs_asap.py: the
  trace workloads alibaba_v2022 and google_2011. boom is left out for now.
- scripts/export_autosketch_eval_table.py --synthetic: one workload-only
  table per grid point, plus plan.tsv. The data constants come from
  study_saturation's COST_*. The traces path is unchanged from main.
- scripts/run_autosketch_vs_asap_synthetic.sh and
  plot_autosketch_vs_asap_synthetic.py: run the synthetic grid, then plot.
- scripts/profiling_time.py and the results README: ASAP's one-time
  profiling time, and how to reproduce.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
zzylol and others added 4 commits October 7, 2026 23:27
…tudy

- §5: a single p95 level in each family's metric (error ≤ 0.05, top-k
  precision ≥ 0.95) for synthetic and traces; the three strictness levels
  and fitted trace targets are dropped, and so is the grid's strictness
  dimension.
- §2/§6: costs and accuracy come from one study on asap_sketchlib 0.3.0:
  the cost table from --phase optimizer-cost (#174), a full-cross grid
  with the cost shape (#186), KLL merge curves (#179), top-k heap m·k
  (#182), per-query k with curves at 10/32/100 (#185) and the theoretical
  fallback (#180). Quantiles see Pareto a = 2, the cost shape.
- §6: traces are Alibaba 2022 and Google 2011; BOOM is left out for now.
- §8/§10: statuses as of 2026-10-07; the top-k m·k read is an assumption.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis
…∈ {1, 8}

Shared replicas on one metric stop adding RQEs past r = 8, so r = 64 is
dropped. m ∈ {1, 8, 16} copies of the template set, each on its own
metric, grow RQEs as 21 · m and drive the planning-time figure.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis
zzylol added a commit to ProjectASAP/sketch-bench that referenced this pull request Oct 8, 2026
… workloads

ProjectASAP/ASAPQuery#777 (paper §6.3), consolidating #138, #139 and #141.

- rqe-optimizer/examples/autosketch_vs_asap.rs (traces and synthetic): ASAP
  (milp::minimize), PerQuery-CostAware and AutoSketch-Adapted. Costs and
  accuracy come from --saturation-dir, as the planner reads them:
  SaturationCurves::accuracy for ASAP (merge curves for KLL and top-k,
  #179), autosketch_accuracy for AutoSketch, and the one-study cost table
  (#178). Each run is scored by analytical_cost_model::score at #777's two
  weight settings and the SLA grid.
- scripts/run_autosketch_vs_asap.sh and plot_autosketch_vs_asap.py: the
  trace workloads alibaba_v2022 and google_2011. boom is left out for now.
- scripts/export_autosketch_eval_table.py --synthetic: one workload-only
  table per grid point, plus plan.tsv. The data constants come from
  study_saturation's COST_*. The traces path is unchanged from main.
- scripts/run_autosketch_vs_asap_synthetic.sh and
  plot_autosketch_vs_asap_synthetic.py: run the synthetic grid, then plot.
- scripts/profiling_time.py and the results README: ASAP's one-time
  profiling time, and how to reproduce.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
zzylol added a commit to ProjectASAP/sketch-bench that referenced this pull request Oct 8, 2026
… workloads

ProjectASAP/ASAPQuery#777 (paper §6.3), consolidating #138, #139 and #141.

- rqe-optimizer/examples/autosketch_vs_asap.rs (traces and synthetic): ASAP
  (milp::minimize), PerQuery-CostAware and AutoSketch-Adapted. Costs and
  accuracy come from --saturation-dir, as the planner reads them:
  SaturationCurves::accuracy for ASAP (merge curves for KLL and top-k,
  #179), autosketch_accuracy for AutoSketch, and the one-study cost table
  (#178). Each run is scored by analytical_cost_model::score at #777's two
  weight settings and the SLA grid.
- scripts/run_autosketch_vs_asap.sh and plot_autosketch_vs_asap.py: the
  trace workloads alibaba_v2022 and google_2011. boom is left out for now.
- scripts/export_autosketch_eval_table.py --synthetic: one workload-only
  table per grid point, plus plan.tsv. The data constants come from
  study_saturation's COST_*. The traces path is unchanged from main.
- scripts/run_autosketch_vs_asap_synthetic.sh and
  plot_autosketch_vs_asap_synthetic.py: run the synthetic grid, then plot.
- scripts/profiling_time.py and the results README: ASAP's one-time
  profiling time, and how to reproduce.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- §4: plans priced by use, w_cpu * AUC(CPU) + w_mem * AUC(memory), CPU
  elastic: ingest on ceil(rho) workers split by sample, compaction at
  window close, query jobs; memory as ingest, storage, compaction and
  query, each counted once; latency is a batch's longest chain (model in
  sketch-bench #188). A peak-billed cost model was dropped.
- §5: latency is reported, in two versions: no SLA, with each method's
  cost-latency frontier, and a batch latency SLA over {100 ms .. 10 s}.
- §3: PerQuery is the same MILP over each RQE's own candidates (no sharing).
- §6: the evaluated set is the mixed template set, with shared r in {1, 8}
  and metrics m in {1, 8, 16}; the dashboard set is not evaluated.
- §7: AutoSketch's planning time is search plus its measured benchmark
  (approxbench accuracy runs at 1e8 items); figures are the frontier, cost
  vs. SLA and planning time; absolute costs only.
- §8-§10: PRs (#188, ASAPQuery #812, #138 stacked on #188), decisions Q3,
  Q5-Q7, limitations (elastic CPU, compaction merge cost).
- §11: results on the synthetic mixed set.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis
zzylol added a commit to ProjectASAP/sketch-bench that referenced this pull request Oct 9, 2026
… workloads

ProjectASAP/ASAPQuery#777 (paper §6.3), consolidating #138, #139 and #141.

- rqe-optimizer/examples/autosketch_vs_asap.rs (traces and synthetic): ASAP
  (milp::minimize), PerQuery-CostAware and AutoSketch-Adapted. Costs and
  accuracy come from --saturation-dir, as the planner reads them:
  SaturationCurves::accuracy for ASAP (merge curves for KLL and top-k,
  #179), autosketch_accuracy for AutoSketch, and the one-study cost table
  (#178). Each run is scored by analytical_cost_model::score at #777's two
  weight settings and the SLA grid.
- scripts/run_autosketch_vs_asap.sh and plot_autosketch_vs_asap.py: the
  trace workloads alibaba_v2022 and google_2011. boom is left out for now.
- scripts/export_autosketch_eval_table.py --synthetic: one workload-only
  table per grid point, plus plan.tsv. The data constants come from
  study_saturation's COST_*. The traces path is unchanged from main.
- scripts/run_autosketch_vs_asap_synthetic.sh and
  plot_autosketch_vs_asap_synthetic.py: run the synthetic grid, then plot.
- scripts/profiling_time.py and the results README: ASAP's one-time
  profiling time, and how to reproduce.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
zzylol added a commit to ProjectASAP/sketch-bench that referenced this pull request Oct 9, 2026
…ontier and SLA versions) (#138)

* rqe-optimizer: evaluate AutoSketch vs ASAP on the synthetic and trace workloads

ProjectASAP/ASAPQuery#777 (paper §6.3), consolidating #138, #139 and #141.

- rqe-optimizer/examples/autosketch_vs_asap.rs (traces and synthetic): ASAP
  (milp::minimize), PerQuery-CostAware and AutoSketch-Adapted. Costs and
  accuracy come from --saturation-dir, as the planner reads them:
  SaturationCurves::accuracy for ASAP (merge curves for KLL and top-k,
  #179), autosketch_accuracy for AutoSketch, and the one-study cost table
  (#178). Each run is scored by analytical_cost_model::score at #777's two
  weight settings and the SLA grid.
- scripts/run_autosketch_vs_asap.sh and plot_autosketch_vs_asap.py: the
  trace workloads alibaba_v2022 and google_2011. boom is left out for now.
- scripts/export_autosketch_eval_table.py --synthetic: one workload-only
  table per grid point, plus plan.tsv. The data constants come from
  study_saturation's COST_*. The traces path is unchanged from main.
- scripts/run_autosketch_vs_asap_synthetic.sh and
  plot_autosketch_vs_asap_synthetic.py: run the synthetic grid, then plot.
- scripts/profiling_time.py and the results README: ASAP's one-time
  profiling time, and how to reproduce.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* rqe-optimizer: trace key queries keep one accumulator per key; report the accuracy source

From review on #138:
- A traces key query (sum by key) keeps one exact accumulator per key: its
  groups are the window's keys, not the table's one sketch. A stream's
  card(g) is the largest over its RQEs.
- AutoSketch's benchmark lower bound uses one query of one instance, not
  the stream's q_r of them.
- --weights that names none of the weight settings is an error.
- The chosen report lists each RQE's accuracy source (#180's
  accuracy_with_source).
- Drop a duplicated BoomTest; capability_of's comment matches the code.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* eval: one accuracy level, p95; plots join one RQE set and charge AutoSketch's benchmark

- One accuracy level for every workload, synthetic and traces: 95% in
  each family's metric (rank, relative value and relative error at most
  0.05; top-k precision at least 0.95). The strictness dimension is gone.
- fig_objective_vs_latency joins only points over one RQE set (solid for
  the full set, dashed for the subset tight SLAs keep).
- fig_planning_time adds AutoSketch's charged benchmark time (lower bound
  and 60 s per probe), as #777 section 7 counts it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* eval: summaries show Σ window/slide; synthetic summary to its own file

AutoSketch-Adapted's ingest cost is window / slide overlapping sketches per
deployment (window = the query's lookback), so the column explains its gap
to ASAP. The synthetic and trace scripts both wrote summary.md into one
results directory; the synthetic one is now summary_synthetic.md.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* eval: RQEs carry their top-k k (#185)

Raqe now has topk_k. The tables name it only in the query id (topk32_...),
so the runner reads it there and stops on a top-k query that names none.
On the p95 synthetic and trace tables every top-k query is k = 32, and the
plans are unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* eval: p95 results for the synthetic grid and the traces

Runs on clnode155 over the asap_sketchlib 0.3.0 study (#174, #178–#186):
four synthetic workloads and alibaba_v2022 / google_2011, raw JSON,
summaries and figures. The README gives the no-SLA headline, including
Σ window/slide behind AutoSketch-Adapted's gap. profiling-time.json is
refilled from the study's raw records (0.77 h; the K = 1e4 column refill
of #186 is its own run).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* eval: figures and tables give each method's absolute cost

Cost figures label every bar and point with its value; the summaries and
the results README list each method's objective instead of its ratio to
ASAP. The by-dimension figure's legend moves above the panels.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* eval: metrics dimension, one dashboard per metric (21 · m RQEs)

Shared replicas on one metric stop adding RQEs past r = 8 (52 at r = 8,
57 at r = 64), so the grid keeps shared r ∈ {1, 8} and adds metrics
m ∈ {1, 8, 16}: copy i of the template set reads its own metric data_i,
so nothing is shared across copies and RQEs grow as 21 · m (168, 336).
The planning-time figure plots this dimension. The m = 8 and 16 runs are
added; the r = 64 result is removed. Existing tables are unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* eval: AutoSketch benchmarks each config once per metric

Synthetic tables name the metric each RQE reads (data, or data_i in the
metrics dimension), and AutoSketch-Adapted's benchmark time counts a
probed config once per metric, on that metric's data: 540, 4320 and
8640 s (60 s per probe) for m = 1, 8, 16. Trace tables name none and
read one metric per dataset, as before. All five synthetic workloads are
rerun with the current runner on four nodes; objectives are unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* eval: cost vs. total estimated latency for every workload

One column per workload (five synthetic, two traces), one row per weight
setting, one line per method across the SLA grid, each point labeled with
its cost. Total latency is the sum of the RQEs' estimated latencies.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* eval: finer SLA grid, {0.003, 0.01, 0.03, 0.1, 0.3, 1, 3, 10} ms and none

Checks whether AutoSketch-Adapted's memory objective makes it miss an
SLA: it never does. Its plans never merge, so each RQE's latency is within
1.4% of the lowest any deployment reaches (equal for quantile, exact and
trace RQEs), and RQEs no method can meet are excluded for all methods.
All workloads rerun on four nodes; figures and summaries regenerated.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* eval: AutoSketch probe list and measured benchmark time; traces with 6h/24h windows

- The runner lists AutoSketch-Adapted's distinct probed (metric, config)
  pairs with the data shape each is benchmarked on, and drops the
  benchmark-time lower bound. scripts/autosketch_benchmark_time.py runs
  each probe's approxbench accuracy benchmark (1e8 items, exact baseline
  and scoring) and times it, as AutoSketch benchmarks probed configs.
- Trace streams with more keys than samples per second are left out and
  listed (rounding their rate up to their key count inflated every
  method's cost).
- The trace table is regenerated from ASAPQuery #812's skew summary
  (6h and 24h windows, a day of Alibaba).
- Plots: one line per method over the full RQE set; AutoSketch's planning
  time with its measured benchmark.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* eval: cost by use (sketch-bench #188); Pareto frontier sweep; mixed set by default

- The runner prices every plan by use (usage::usage_cost): w_cpu * AUC(CPU)
  + w_mem * AUC(memory), CPU elastic, latency the longest chain, reported.
- ASAP is milp::minimize_usage_cost over every RQE; PerQuery the same over
  each RQE's own candidates (separate copies via `allowed`, no sharing);
  AutoSketch's plan is priced the same way.
- No SLA grid. Each method's cost-latency frontier: ASAP and PerQuery are
  solved unbounded and for 12 bounds log-spaced from the tightest feasible
  one to the unbounded latency, plus AutoSketch's latency.
- The synthetic grid is the mixed template set: shared r in {1, 8} and
  metrics m in {1, 8, 16} (50, 92, 400, 800 RQEs).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* eval: version 1 (no SLA) results on the mixed set, with AutoSketch's measured benchmark

- Four workloads of the mixed template set (50, 92 with r = 8, 400 with
  m = 8, 800 with m = 16 RQEs), cost by use (#188), each method's cheapest
  plan and ASAP's and PerQuery's cost-latency frontiers (12 bounds plus
  AutoSketch's latency), CPU-only and Fargate weights, --runs 3, one node
  per workload. No RQE dropped, no sanity violation.
- AutoSketch's benchmark time is measured: approxbench's accuracy run (1e8
  items, exact baseline, scoring) once per distinct probed config and
  shape, serially on idle clnode155 (21-37 s each), charged once per
  probed (metric, config): 212, 1697 and 3394 s for m = 1, 8, 16.
- plot_autosketch_vs_asap_synthetic.py: fig_frontier.png (cost vs. query
  latency, every point labeled), fig_planning_time.png (AutoSketch as
  search + measured benchmark; 60 s per probe only as a labeled
  reference), summary_synthetic.md. The SLA-era figures, the dashboard
  results and plot_autosketch_vs_asap_workloads.py are removed.
- README: the cost model, version 1, how to reproduce, and the headline
  numbers.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* eval: version 2, a batch latency SLA, on the synthetic mixed set

The runner adds version 2 of the evaluation beside version 1 (no latency
constraint and the frontiers, unchanged in `results`): at each SLA of a
grid (default 100, 300, 1000, 3000, 10000 ms; --slas-ms overrides), ASAP
and PerQuery plan the cheapest plan whose batch latency (the longest
chain) is at most the SLA, via minimize_usage_cost's latency bound, which
is exact with elastic CPU. An SLA below a method's tightest feasible
bound gives an infeasible record. AutoSketch's one plan is recorded at
each SLA with meets_sla. Records go to `sla_results`, with sanity checks
that every plan meets its SLA and that ASAP costs no more than PerQuery.

Results on clnode109 (--runs 3) for the mixed set, r = 8 and m = 8, 16:
no drops, no sanity violations; the tightest feasible SLA is 92 ms, and
AutoSketch (latency 1011 ms) meets only 3000 and 10000 ms. AutoSketch's
planning time is search plus its measured benchmark
(scripts/autosketch_benchmark_time.py: approxbench accuracy runs at 1e8
items, 20-33 s each, serial and alone on the node): 197 s at m = 1,
3159 s at m = 16. scripts/plot_autosketch_vs_asap_sla.py draws
fig_cost_vs_sla.png and summary_sla.md.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* eval: remove the old trace results and figures

They came from the earlier model (per-RQE SLA over a µs grid, mean-CPU
cost); the traces are to be rerun on cost by use with both latency
versions.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* eval: trim the PR to what it needs

- Result JSONs without per-RQE choices (`chosen`): the figures and
  summaries don't read them, and they were nearly all of the 2.9M lines;
  rerun the runner for them.
- Machine files folded into the READMEs (all Intel Xeon E5-2683 v3,
  56 cores); the synthetic run script no longer writes them.
- The SLA-era trace scripts (plot_autosketch_vs_asap.py,
  run_autosketch_vs_asap.sh) are removed; the traces are to be rerun on
  cost by use.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* eval: keep only what the three figures need

The PR's results are the frontier (version 1), cost vs. SLA (version 2)
and planning time on the synthetic mixed set. The trace table goes back
to the base version (the traces are to be rerun on cost by use, with
their new table), and the ASAP profiling-time script and data, which no
figure uses, are removed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* eval: commit the inputs of the three figures

rqe-optimizer/results/autosketch-vs-asap-inputs/ holds what the figures
were computed from: the sketch-bench study on asap_sketchlib 0.3.0 (the
cost table, sketches and exact aggregations, and the accuracy curves the
runner reads as --saturation-dir) and the synthetic mixed-set tables with
plan.tsv. The READMEs point the reproduce steps at them.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* eval: the raw approxbench records behind the figures' study

rqe-optimizer/results/autosketch-vs-asap-inputs/raw/ holds the gzipped
approxbench records (920 KB) that ../saturation/ was reduced from, with a
README mapping each file to its machine, arguments, record count, run time
and output. Rerunning them is 0.77 h of machine time over 5 machines in
parallel (about 20-30 minutes of wall time).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
zzylol and others added 2 commits October 10, 2026 14:20
…10-10)

Results of sketch-bench #203: roll-ups serve 40% of the multi-grouping
RQEs and halve ASAP's deployments but save only 3.3–4.8%; Hydra is
eligible but never chosen (its insert fans out to every label subset,
1.9–5.1 µs per record vs. 3.85 ns), so it is not implemented in
ASAPQuery. Adds the multi-grouping template set and the no-roll-up method.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis
…rkload (2026-10-11)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants