Repository navigation
Conversation
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…cisions Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…s for KLL/top-k Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…n the whole dataset Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This was referenced Oct 4, 2026
Merged
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
3 tasks
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
zzylol
added a commit
to ProjectASAP/sketch-bench
that referenced
this pull request
Oct 5, 2026
…anner workload Measure frequency and top-k saturation curves at K = 10, 100, 1e4 and 1e6 (theta 0..2, N up to 1e8, 3 seeds) and top-k accuracy after merging 4/16/64 shards, for the synthetic PromQL workload of ProjectASAP/ASAPQuery#777. Commit the summary CSVs, the N_sat tables and the scripts that finish and check an interrupted cost phase. Cost results are pending. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
zzylol
added a commit
to ProjectASAP/sketch-bench
that referenced
this pull request
Oct 5, 2026
* rqe-optimizer: price plans per EC2 machine family Score retained memory per active deployment, (x + max S) / y instances, and add milp::minimize_cost, which minimizes the hourly price on one EC2 machine family in fractional instances (whichever of CPU or memory binds). Prices are a committed on-demand snapshot from scripts/fetch_ec2_pricing.py. Dominance pruning also compares retained memory so it cannot drop a candidate that is cheaper under the new objective. Part of ProjectASAP/ASAPQuery#777. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * rqe-optimizer: label small_problem MILP output for both objectives Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * rqe-optimizer: scale the MILP so tiny costs and latencies solve exactly HiGHS's absolute gap (1e-6) and feasibility tolerance (1e-7) exceed real plan costs (~1e-6 $/hour) and query latencies (µs). Rows are now scaled by the plan's demand without sharing, and per-RQE memory and latency bounds become f64 exclusions of the choices over them. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
…tatus Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ties Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
zzylol
added a commit
to ProjectASAP/sketch-bench
that referenced
this pull request
Oct 5, 2026
The user dropped FewestPlans from the evaluation (ProjectASAP/ASAPQuery#777). Remove milp::minimize_deployments and MilpBounds::max_active_deployments, FewestPlans from the runner and its sanity checks, and its column from the plots and summaries. milp.rs and small_problem.rs are back to #138's versions. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
zzylol
added a commit
to ProjectASAP/sketch-bench
that referenced
this pull request
Oct 5, 2026
Runner (examples/autosketch_vs_asap.rs) for the traces workloads of ProjectASAP/ASAPQuery#777: ASAP (joint MILP with one absolute latency SLA per sweep point), PerQuery-CostAware and AutoSketch-Adapted, scored by objectives::score and priced per EC2 family, with sanity checks. Accuracy now depends on the query's merge count m = S/x: a cost entry may carry "{metric}@m{b}" keys, eligibility reads the smallest measured b >= m, and a merge larger than every measured b is ineligible. Results for Alibaba 2022, Google 2011 and BOOM with figures and summary. The earlier example and scaling workloads are dropped (#777 §6). Rebased onto main, which now has #137, #135 and #136. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
zzylol
added a commit
to ProjectASAP/sketch-bench
that referenced
this pull request
Oct 5, 2026
…nd PerQuery - Runner: a synthetic subcommand reading a table from export_autosketch_eval_table.py --synthetic (--target, --replicas, --no-chosen) and a wider synthetic SLA grid. PerQuery-CostAware now solves each RQE under the same SLA as ASAP, with a sanity check that ASAP costs no more. - export_autosketch_eval_table.py --synthetic: the 10 PromQL templates of ProjectASAP/ASAPQuery#777 §6 as RQEs, with per-RQE accuracy lookups. - run/plot scripts for the synthetic sweep. Rebased onto #138 (f293b73), which dropped the example and scaling workloads; FewestPlans is not included. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
zzylol
added a commit
to ProjectASAP/sketch-bench
that referenced
this pull request
Oct 6, 2026
ProjectASAP/ASAPQuery#777 section 6: - templates=dashboard (D1-D5, 33 RQEs): per-series p50/p90/p99 over {1m, 5m, 15m, 1h, 6h, 24h}, p50/p90/p99 by label_0, sum by label_0 of rate over the six windows, top-3 by label_0 of rate over 5m and 1h, and the p99/p50 ratio over 5m and 1h as two quantile RQEs on D1's stream. - Shared replicas (table dimension shared=r): r replicas of a template set on the same streams; replica i draws 3 of the set's windows, 3 of the five quantiles and an interval (10/60/300 s) from a seeded RNG, and identical RQEs are kept once (temporal ranges carry their interval, e.g. 1h@60s). The runner's --replicas stays the disjoint control. - The plan adds the dashboard and shared r in {1..64} for the default mix and the dashboard, at every strictness level (111 runs in all). The 40 existing tables are unchanged. - Summaries: dashboard tables, a shared-replicas table, and fig_synthetic_shared_replicas.png (cost relative to ASAP vs. r). The runner is unchanged since ce4b885. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
zzylol
marked this pull request as ready for review
October 6, 2026 19:59
…e cost model Addresses review on #777: - cost: sketch-bench #145's per-phase CPU and memory with a weighted objective; CPU only first, then Fargate prices. EC2 machine families, the peak-provisioned model and the CPU timeline are dropped. - capabilities: sum and rate/increase map to the exact accumulators (#144), not frequency; strictness applies to quantile and top-k only. - inputs: read the export_rqe_optimizer_costs table (measured_at, merge accuracy, size sweep) instead of the saturation tables; exact accumulators are benchmarked for cost but need no saturation point. - grid: keep template set, shared replicas {1, 8, 64}, C {1e2, 1e3, 1e4}, strictness and SLA {0.01, 0.1, 1, 10} ms; move the rest to planner sensitivity. Disjoint replicas are dropped (spatial filters unsupported). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… workload Top-k, KLL and DDSketch are measured on Zipf s=1.1 over 100k keys (quantiles on the Zipf ranks). The synthetic workload now uses that distribution instead of Zipf θ=1.0 and Pareto a=2, so the cost table needs no separate run. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- data: C = 1e4 series (label_0, the top-k key), J = 10 jobs, 200
samples/s per series, Zipf s=1.1 over 10,000 keys (the cost table's
data at CARDINALITY=10000); fixed, so the comparison does not sweep it.
- windows W = {15m, 1h, 6h, 24h}: every per-series summary sees >= 1e5
items; per-summary input sizes listed per grouping.
- queries: grouped templates use by (job); top-k is
topk(k, sum by (label_0) (...)) with k in {100, 200, 300}, served by
one heap-K deployment for k <= K; quantile syntax fixed. Both template
sets are enumerated query by query (all checked with promql-parser).
- grid: C removed.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…h-bench #156) Keep the fixed synthetic data (Zipf s=1.1 over 10,000 keys, quantiles on the Zipf ranks) and top-k with k in {100, 200, 300} served by heap-K deployments. Everything else follows sketch-bench #156/#162: - sketch accuracy from the saturation curves at n(S, G); CPU, memory and exact rows from the cost table; grid configs only. - merging: every pane saturated, then read at n(S, G); KLL/top-k merge penalty is #158. AutoSketch uses the same lookup with m = 1. - DDSketch scored in value relative error (#162), with its own targets. - a targeted saturation run at the workload's data shape (theta 1.1, K 1e4, Zipf-rank quantiles, CMS-heap heaps 100/200/300) so lookups land on measured points. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
k = 32 is the heap size sketch-bench measures top-k at (CMS_HEAP_TOP_K), so top-k needs no heap parameter, no k on the RQE and no extra saturation points. 58 RQEs per replica for the 10 templates (50 distinct), 25 for the dashboard (21 distinct); query lists regenerated and parser-checked. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
zzylol
added a commit
to ProjectASAP/sketch-bench
that referenced
this pull request
Oct 7, 2026
ProjectASAP/ASAPQuery#777 (paper §6.3), consolidating #138, #139 and #141. - rqe-optimizer/examples/autosketch_vs_asap.rs: ASAP (milp::minimize), PerQuery-CostAware and AutoSketch-Adapted on one synthetic table. Costs and accuracy come from --saturation-dir, as the planner reads them: SaturationCurves::accuracy for ASAP (merge curves for KLL and top-k, #179), autosketch_accuracy for AutoSketch, and the one-study cost table (#178). Each run is scored by analytical_cost_model::score at #777's two weight settings and the SLA grid. - scripts/export_autosketch_eval_table.py --synthetic: one workload-only table per grid point, plus plan.tsv. The data constants come from study_saturation's COST_*. - scripts/run_autosketch_vs_asap_synthetic.sh and plot_autosketch_vs_asap_synthetic.py: run the grid, then plot. - scripts/profiling_time.py and the synthetic results README: ASAP's one-time profiling time, and how to reproduce. The trace workloads are left out for now. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
zzylol
added a commit
to ProjectASAP/sketch-bench
that referenced
this pull request
Oct 7, 2026
… workloads ProjectASAP/ASAPQuery#777 (paper §6.3), consolidating #138, #139 and #141. - rqe-optimizer/examples/autosketch_vs_asap.rs (traces and synthetic): ASAP (milp::minimize), PerQuery-CostAware and AutoSketch-Adapted. Costs and accuracy come from --saturation-dir, as the planner reads them: SaturationCurves::accuracy for ASAP (merge curves for KLL and top-k, #179), autosketch_accuracy for AutoSketch, and the one-study cost table (#178). Each run is scored by analytical_cost_model::score at #777's two weight settings and the SLA grid. - scripts/run_autosketch_vs_asap.sh and plot_autosketch_vs_asap.py: the trace workloads alibaba_v2022 and google_2011. boom is left out for now. - scripts/export_autosketch_eval_table.py --synthetic: one workload-only table per grid point, plus plan.tsv. The data constants come from study_saturation's COST_*. The traces path is unchanged from main. - scripts/run_autosketch_vs_asap_synthetic.sh and plot_autosketch_vs_asap_synthetic.py: run the synthetic grid, then plot. - scripts/profiling_time.py and the results README: ASAP's one-time profiling time, and how to reproduce. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
zzylol
added a commit
to ProjectASAP/sketch-bench
that referenced
this pull request
Oct 7, 2026
… workloads ProjectASAP/ASAPQuery#777 (paper §6.3), consolidating #138, #139 and #141. - rqe-optimizer/examples/autosketch_vs_asap.rs (traces and synthetic): ASAP (milp::minimize), PerQuery-CostAware and AutoSketch-Adapted. Costs and accuracy come from --saturation-dir, as the planner reads them: SaturationCurves::accuracy for ASAP (merge curves for KLL and top-k, #179), autosketch_accuracy for AutoSketch, and the one-study cost table (#178). Each run is scored by analytical_cost_model::score at #777's two weight settings and the SLA grid. - scripts/run_autosketch_vs_asap.sh and plot_autosketch_vs_asap.py: the trace workloads alibaba_v2022 and google_2011. boom is left out for now. - scripts/export_autosketch_eval_table.py --synthetic: one workload-only table per grid point, plus plan.tsv. The data constants come from study_saturation's COST_*. The traces path is unchanged from main. - scripts/run_autosketch_vs_asap_synthetic.sh and plot_autosketch_vs_asap_synthetic.py: run the synthetic grid, then plot. - scripts/profiling_time.py and the results README: ASAP's one-time profiling time, and how to reproduce. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
zzylol
added a commit
to ProjectASAP/sketch-bench
that referenced
this pull request
Oct 7, 2026
… workloads ProjectASAP/ASAPQuery#777 (paper §6.3), consolidating #138, #139 and #141. - rqe-optimizer/examples/autosketch_vs_asap.rs (traces and synthetic): ASAP (milp::minimize), PerQuery-CostAware and AutoSketch-Adapted. Costs and accuracy come from --saturation-dir, as the planner reads them: SaturationCurves::accuracy for ASAP (merge curves for KLL and top-k, #179), autosketch_accuracy for AutoSketch, and the one-study cost table (#178). Each run is scored by analytical_cost_model::score at #777's two weight settings and the SLA grid. - scripts/run_autosketch_vs_asap.sh and plot_autosketch_vs_asap.py: the trace workloads alibaba_v2022 and google_2011. boom is left out for now. - scripts/export_autosketch_eval_table.py --synthetic: one workload-only table per grid point, plus plan.tsv. The data constants come from study_saturation's COST_*. The traces path is unchanged from main. - scripts/run_autosketch_vs_asap_synthetic.sh and plot_autosketch_vs_asap_synthetic.py: run the synthetic grid, then plot. - scripts/profiling_time.py and the results README: ASAP's one-time profiling time, and how to reproduce. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
zzylol
added a commit
to ProjectASAP/sketch-bench
that referenced
this pull request
Oct 7, 2026
… workloads ProjectASAP/ASAPQuery#777 (paper §6.3), consolidating #138, #139 and #141. - rqe-optimizer/examples/autosketch_vs_asap.rs (traces and synthetic): ASAP (milp::minimize), PerQuery-CostAware and AutoSketch-Adapted. Costs and accuracy come from --saturation-dir, as the planner reads them: SaturationCurves::accuracy for ASAP (merge curves for KLL and top-k, #179), autosketch_accuracy for AutoSketch, and the one-study cost table (#178). Each run is scored by analytical_cost_model::score at #777's two weight settings and the SLA grid. - scripts/run_autosketch_vs_asap.sh and plot_autosketch_vs_asap.py: the trace workloads alibaba_v2022 and google_2011. boom is left out for now. - scripts/export_autosketch_eval_table.py --synthetic: one workload-only table per grid point, plus plan.tsv. The data constants come from study_saturation's COST_*. The traces path is unchanged from main. - scripts/run_autosketch_vs_asap_synthetic.sh and plot_autosketch_vs_asap_synthetic.py: run the synthetic grid, then plot. - scripts/profiling_time.py and the results README: ASAP's one-time profiling time, and how to reproduce. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
zzylol
added a commit
to ProjectASAP/sketch-bench
that referenced
this pull request
Oct 7, 2026
… workloads ProjectASAP/ASAPQuery#777 (paper §6.3), consolidating #138, #139 and #141. - rqe-optimizer/examples/autosketch_vs_asap.rs (traces and synthetic): ASAP (milp::minimize), PerQuery-CostAware and AutoSketch-Adapted. Costs and accuracy come from --saturation-dir, as the planner reads them: SaturationCurves::accuracy for ASAP (merge curves for KLL and top-k, #179), autosketch_accuracy for AutoSketch, and the one-study cost table (#178). Each run is scored by analytical_cost_model::score at #777's two weight settings and the SLA grid. - scripts/run_autosketch_vs_asap.sh and plot_autosketch_vs_asap.py: the trace workloads alibaba_v2022 and google_2011. boom is left out for now. - scripts/export_autosketch_eval_table.py --synthetic: one workload-only table per grid point, plus plan.tsv. The data constants come from study_saturation's COST_*. The traces path is unchanged from main. - scripts/run_autosketch_vs_asap_synthetic.sh and plot_autosketch_vs_asap_synthetic.py: run the synthetic grid, then plot. - scripts/profiling_time.py and the results README: ASAP's one-time profiling time, and how to reproduce. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…tudy - §5: a single p95 level in each family's metric (error ≤ 0.05, top-k precision ≥ 0.95) for synthetic and traces; the three strictness levels and fitted trace targets are dropped, and so is the grid's strictness dimension. - §2/§6: costs and accuracy come from one study on asap_sketchlib 0.3.0: the cost table from --phase optimizer-cost (#174), a full-cross grid with the cost shape (#186), KLL merge curves (#179), top-k heap m·k (#182), per-query k with curves at 10/32/100 (#185) and the theoretical fallback (#180). Quantiles see Pareto a = 2, the cost shape. - §6: traces are Alibaba 2022 and Google 2011; BOOM is left out for now. - §8/§10: statuses as of 2026-10-07; the top-k m·k read is an assumption. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis
…∈ {1, 8}
Shared replicas on one metric stop adding RQEs past r = 8, so r = 64 is
dropped. m ∈ {1, 8, 16} copies of the template set, each on its own
metric, grow RQEs as 21 · m and drive the planning-time figure.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis
This was referenced Oct 8, 2026
zzylol
added a commit
to ProjectASAP/sketch-bench
that referenced
this pull request
Oct 8, 2026
… workloads ProjectASAP/ASAPQuery#777 (paper §6.3), consolidating #138, #139 and #141. - rqe-optimizer/examples/autosketch_vs_asap.rs (traces and synthetic): ASAP (milp::minimize), PerQuery-CostAware and AutoSketch-Adapted. Costs and accuracy come from --saturation-dir, as the planner reads them: SaturationCurves::accuracy for ASAP (merge curves for KLL and top-k, #179), autosketch_accuracy for AutoSketch, and the one-study cost table (#178). Each run is scored by analytical_cost_model::score at #777's two weight settings and the SLA grid. - scripts/run_autosketch_vs_asap.sh and plot_autosketch_vs_asap.py: the trace workloads alibaba_v2022 and google_2011. boom is left out for now. - scripts/export_autosketch_eval_table.py --synthetic: one workload-only table per grid point, plus plan.tsv. The data constants come from study_saturation's COST_*. The traces path is unchanged from main. - scripts/run_autosketch_vs_asap_synthetic.sh and plot_autosketch_vs_asap_synthetic.py: run the synthetic grid, then plot. - scripts/profiling_time.py and the results README: ASAP's one-time profiling time, and how to reproduce. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
zzylol
added a commit
to ProjectASAP/sketch-bench
that referenced
this pull request
Oct 8, 2026
… workloads ProjectASAP/ASAPQuery#777 (paper §6.3), consolidating #138, #139 and #141. - rqe-optimizer/examples/autosketch_vs_asap.rs (traces and synthetic): ASAP (milp::minimize), PerQuery-CostAware and AutoSketch-Adapted. Costs and accuracy come from --saturation-dir, as the planner reads them: SaturationCurves::accuracy for ASAP (merge curves for KLL and top-k, #179), autosketch_accuracy for AutoSketch, and the one-study cost table (#178). Each run is scored by analytical_cost_model::score at #777's two weight settings and the SLA grid. - scripts/run_autosketch_vs_asap.sh and plot_autosketch_vs_asap.py: the trace workloads alibaba_v2022 and google_2011. boom is left out for now. - scripts/export_autosketch_eval_table.py --synthetic: one workload-only table per grid point, plus plan.tsv. The data constants come from study_saturation's COST_*. The traces path is unchanged from main. - scripts/run_autosketch_vs_asap_synthetic.sh and plot_autosketch_vs_asap_synthetic.py: run the synthetic grid, then plot. - scripts/profiling_time.py and the results README: ASAP's one-time profiling time, and how to reproduce. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- §4: plans priced by use, w_cpu * AUC(CPU) + w_mem * AUC(memory), CPU
elastic: ingest on ceil(rho) workers split by sample, compaction at
window close, query jobs; memory as ingest, storage, compaction and
query, each counted once; latency is a batch's longest chain (model in
sketch-bench #188). A peak-billed cost model was dropped.
- §5: latency is reported, in two versions: no SLA, with each method's
cost-latency frontier, and a batch latency SLA over {100 ms .. 10 s}.
- §3: PerQuery is the same MILP over each RQE's own candidates (no sharing).
- §6: the evaluated set is the mixed template set, with shared r in {1, 8}
and metrics m in {1, 8, 16}; the dashboard set is not evaluated.
- §7: AutoSketch's planning time is search plus its measured benchmark
(approxbench accuracy runs at 1e8 items); figures are the frontier, cost
vs. SLA and planning time; absolute costs only.
- §8-§10: PRs (#188, ASAPQuery #812, #138 stacked on #188), decisions Q3,
Q5-Q7, limitations (elastic CPU, compaction merge cost).
- §11: results on the synthetic mixed set.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis
zzylol
added a commit
to ProjectASAP/sketch-bench
that referenced
this pull request
Oct 9, 2026
… workloads ProjectASAP/ASAPQuery#777 (paper §6.3), consolidating #138, #139 and #141. - rqe-optimizer/examples/autosketch_vs_asap.rs (traces and synthetic): ASAP (milp::minimize), PerQuery-CostAware and AutoSketch-Adapted. Costs and accuracy come from --saturation-dir, as the planner reads them: SaturationCurves::accuracy for ASAP (merge curves for KLL and top-k, #179), autosketch_accuracy for AutoSketch, and the one-study cost table (#178). Each run is scored by analytical_cost_model::score at #777's two weight settings and the SLA grid. - scripts/run_autosketch_vs_asap.sh and plot_autosketch_vs_asap.py: the trace workloads alibaba_v2022 and google_2011. boom is left out for now. - scripts/export_autosketch_eval_table.py --synthetic: one workload-only table per grid point, plus plan.tsv. The data constants come from study_saturation's COST_*. The traces path is unchanged from main. - scripts/run_autosketch_vs_asap_synthetic.sh and plot_autosketch_vs_asap_synthetic.py: run the synthetic grid, then plot. - scripts/profiling_time.py and the results README: ASAP's one-time profiling time, and how to reproduce. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
zzylol
added a commit
to ProjectASAP/sketch-bench
that referenced
this pull request
Oct 9, 2026
…ontier and SLA versions) (#138) * rqe-optimizer: evaluate AutoSketch vs ASAP on the synthetic and trace workloads ProjectASAP/ASAPQuery#777 (paper §6.3), consolidating #138, #139 and #141. - rqe-optimizer/examples/autosketch_vs_asap.rs (traces and synthetic): ASAP (milp::minimize), PerQuery-CostAware and AutoSketch-Adapted. Costs and accuracy come from --saturation-dir, as the planner reads them: SaturationCurves::accuracy for ASAP (merge curves for KLL and top-k, #179), autosketch_accuracy for AutoSketch, and the one-study cost table (#178). Each run is scored by analytical_cost_model::score at #777's two weight settings and the SLA grid. - scripts/run_autosketch_vs_asap.sh and plot_autosketch_vs_asap.py: the trace workloads alibaba_v2022 and google_2011. boom is left out for now. - scripts/export_autosketch_eval_table.py --synthetic: one workload-only table per grid point, plus plan.tsv. The data constants come from study_saturation's COST_*. The traces path is unchanged from main. - scripts/run_autosketch_vs_asap_synthetic.sh and plot_autosketch_vs_asap_synthetic.py: run the synthetic grid, then plot. - scripts/profiling_time.py and the results README: ASAP's one-time profiling time, and how to reproduce. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * rqe-optimizer: trace key queries keep one accumulator per key; report the accuracy source From review on #138: - A traces key query (sum by key) keeps one exact accumulator per key: its groups are the window's keys, not the table's one sketch. A stream's card(g) is the largest over its RQEs. - AutoSketch's benchmark lower bound uses one query of one instance, not the stream's q_r of them. - --weights that names none of the weight settings is an error. - The chosen report lists each RQE's accuracy source (#180's accuracy_with_source). - Drop a duplicated BoomTest; capability_of's comment matches the code. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis * eval: one accuracy level, p95; plots join one RQE set and charge AutoSketch's benchmark - One accuracy level for every workload, synthetic and traces: 95% in each family's metric (rank, relative value and relative error at most 0.05; top-k precision at least 0.95). The strictness dimension is gone. - fig_objective_vs_latency joins only points over one RQE set (solid for the full set, dashed for the subset tight SLAs keep). - fig_planning_time adds AutoSketch's charged benchmark time (lower bound and 60 s per probe), as #777 section 7 counts it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis * eval: summaries show Σ window/slide; synthetic summary to its own file AutoSketch-Adapted's ingest cost is window / slide overlapping sketches per deployment (window = the query's lookback), so the column explains its gap to ASAP. The synthetic and trace scripts both wrote summary.md into one results directory; the synthetic one is now summary_synthetic.md. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis * eval: RQEs carry their top-k k (#185) Raqe now has topk_k. The tables name it only in the query id (topk32_...), so the runner reads it there and stops on a top-k query that names none. On the p95 synthetic and trace tables every top-k query is k = 32, and the plans are unchanged. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis * eval: p95 results for the synthetic grid and the traces Runs on clnode155 over the asap_sketchlib 0.3.0 study (#174, #178–#186): four synthetic workloads and alibaba_v2022 / google_2011, raw JSON, summaries and figures. The README gives the no-SLA headline, including Σ window/slide behind AutoSketch-Adapted's gap. profiling-time.json is refilled from the study's raw records (0.77 h; the K = 1e4 column refill of #186 is its own run). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis * eval: figures and tables give each method's absolute cost Cost figures label every bar and point with its value; the summaries and the results README list each method's objective instead of its ratio to ASAP. The by-dimension figure's legend moves above the panels. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis * eval: metrics dimension, one dashboard per metric (21 · m RQEs) Shared replicas on one metric stop adding RQEs past r = 8 (52 at r = 8, 57 at r = 64), so the grid keeps shared r ∈ {1, 8} and adds metrics m ∈ {1, 8, 16}: copy i of the template set reads its own metric data_i, so nothing is shared across copies and RQEs grow as 21 · m (168, 336). The planning-time figure plots this dimension. The m = 8 and 16 runs are added; the r = 64 result is removed. Existing tables are unchanged. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis * eval: AutoSketch benchmarks each config once per metric Synthetic tables name the metric each RQE reads (data, or data_i in the metrics dimension), and AutoSketch-Adapted's benchmark time counts a probed config once per metric, on that metric's data: 540, 4320 and 8640 s (60 s per probe) for m = 1, 8, 16. Trace tables name none and read one metric per dataset, as before. All five synthetic workloads are rerun with the current runner on four nodes; objectives are unchanged. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis * eval: cost vs. total estimated latency for every workload One column per workload (five synthetic, two traces), one row per weight setting, one line per method across the SLA grid, each point labeled with its cost. Total latency is the sum of the RQEs' estimated latencies. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis * eval: finer SLA grid, {0.003, 0.01, 0.03, 0.1, 0.3, 1, 3, 10} ms and none Checks whether AutoSketch-Adapted's memory objective makes it miss an SLA: it never does. Its plans never merge, so each RQE's latency is within 1.4% of the lowest any deployment reaches (equal for quantile, exact and trace RQEs), and RQEs no method can meet are excluded for all methods. All workloads rerun on four nodes; figures and summaries regenerated. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis * eval: AutoSketch probe list and measured benchmark time; traces with 6h/24h windows - The runner lists AutoSketch-Adapted's distinct probed (metric, config) pairs with the data shape each is benchmarked on, and drops the benchmark-time lower bound. scripts/autosketch_benchmark_time.py runs each probe's approxbench accuracy benchmark (1e8 items, exact baseline and scoring) and times it, as AutoSketch benchmarks probed configs. - Trace streams with more keys than samples per second are left out and listed (rounding their rate up to their key count inflated every method's cost). - The trace table is regenerated from ASAPQuery #812's skew summary (6h and 24h windows, a day of Alibaba). - Plots: one line per method over the full RQE set; AutoSketch's planning time with its measured benchmark. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis * eval: cost by use (sketch-bench #188); Pareto frontier sweep; mixed set by default - The runner prices every plan by use (usage::usage_cost): w_cpu * AUC(CPU) + w_mem * AUC(memory), CPU elastic, latency the longest chain, reported. - ASAP is milp::minimize_usage_cost over every RQE; PerQuery the same over each RQE's own candidates (separate copies via `allowed`, no sharing); AutoSketch's plan is priced the same way. - No SLA grid. Each method's cost-latency frontier: ASAP and PerQuery are solved unbounded and for 12 bounds log-spaced from the tightest feasible one to the unbounded latency, plus AutoSketch's latency. - The synthetic grid is the mixed template set: shared r in {1, 8} and metrics m in {1, 8, 16} (50, 92, 400, 800 RQEs). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis * eval: version 1 (no SLA) results on the mixed set, with AutoSketch's measured benchmark - Four workloads of the mixed template set (50, 92 with r = 8, 400 with m = 8, 800 with m = 16 RQEs), cost by use (#188), each method's cheapest plan and ASAP's and PerQuery's cost-latency frontiers (12 bounds plus AutoSketch's latency), CPU-only and Fargate weights, --runs 3, one node per workload. No RQE dropped, no sanity violation. - AutoSketch's benchmark time is measured: approxbench's accuracy run (1e8 items, exact baseline, scoring) once per distinct probed config and shape, serially on idle clnode155 (21-37 s each), charged once per probed (metric, config): 212, 1697 and 3394 s for m = 1, 8, 16. - plot_autosketch_vs_asap_synthetic.py: fig_frontier.png (cost vs. query latency, every point labeled), fig_planning_time.png (AutoSketch as search + measured benchmark; 60 s per probe only as a labeled reference), summary_synthetic.md. The SLA-era figures, the dashboard results and plot_autosketch_vs_asap_workloads.py are removed. - README: the cost model, version 1, how to reproduce, and the headline numbers. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis * eval: version 2, a batch latency SLA, on the synthetic mixed set The runner adds version 2 of the evaluation beside version 1 (no latency constraint and the frontiers, unchanged in `results`): at each SLA of a grid (default 100, 300, 1000, 3000, 10000 ms; --slas-ms overrides), ASAP and PerQuery plan the cheapest plan whose batch latency (the longest chain) is at most the SLA, via minimize_usage_cost's latency bound, which is exact with elastic CPU. An SLA below a method's tightest feasible bound gives an infeasible record. AutoSketch's one plan is recorded at each SLA with meets_sla. Records go to `sla_results`, with sanity checks that every plan meets its SLA and that ASAP costs no more than PerQuery. Results on clnode109 (--runs 3) for the mixed set, r = 8 and m = 8, 16: no drops, no sanity violations; the tightest feasible SLA is 92 ms, and AutoSketch (latency 1011 ms) meets only 3000 and 10000 ms. AutoSketch's planning time is search plus its measured benchmark (scripts/autosketch_benchmark_time.py: approxbench accuracy runs at 1e8 items, 20-33 s each, serial and alone on the node): 197 s at m = 1, 3159 s at m = 16. scripts/plot_autosketch_vs_asap_sla.py draws fig_cost_vs_sla.png and summary_sla.md. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis * eval: remove the old trace results and figures They came from the earlier model (per-RQE SLA over a µs grid, mean-CPU cost); the traces are to be rerun on cost by use with both latency versions. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis * eval: trim the PR to what it needs - Result JSONs without per-RQE choices (`chosen`): the figures and summaries don't read them, and they were nearly all of the 2.9M lines; rerun the runner for them. - Machine files folded into the READMEs (all Intel Xeon E5-2683 v3, 56 cores); the synthetic run script no longer writes them. - The SLA-era trace scripts (plot_autosketch_vs_asap.py, run_autosketch_vs_asap.sh) are removed; the traces are to be rerun on cost by use. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis * eval: keep only what the three figures need The PR's results are the frontier (version 1), cost vs. SLA (version 2) and planning time on the synthetic mixed set. The trace table goes back to the base version (the traces are to be rerun on cost by use, with their new table), and the ASAP profiling-time script and data, which no figure uses, are removed. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis * eval: commit the inputs of the three figures rqe-optimizer/results/autosketch-vs-asap-inputs/ holds what the figures were computed from: the sketch-bench study on asap_sketchlib 0.3.0 (the cost table, sketches and exact aggregations, and the accuracy curves the runner reads as --saturation-dir) and the synthetic mixed-set tables with plan.tsv. The READMEs point the reproduce steps at them. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis * eval: the raw approxbench records behind the figures' study rqe-optimizer/results/autosketch-vs-asap-inputs/raw/ holds the gzipped approxbench records (920 KB) that ../saturation/ was reduced from, with a README mapping each file to its machine, arguments, record count, run time and output. Rerunning them is 0.77 h of machine time over 5 machines in parallel (about 20-30 minutes of wall time). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
…10-10) Results of sketch-bench #203: roll-ups serve 40% of the multi-grouping RQEs and halve ASAP's deployments but save only 3.3–4.8%; Hydra is eligible but never chosen (its insert fans out to every label subset, 1.9–5.1 µs per record vs. 3.85 ns), so it is not implemented in ASAPQuery. Adds the multi-grouping template set and the no-roll-up method. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis
…rkload (2026-10-11) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Before this PR
The §6.3 comparison against AutoSketch had no written plan. The only AutoSketch implementation was an Algorithm 4 adaptation in ASAPQuery-backend (#545 protocol, #547 runner). It is execution-based, CMS-only, and not connected to the RQE MILP that sketch-bench #129 added.
After this PR
docs/evaluation/autosketch-vs-planner.mddefines the planner-level comparison, run offline on sketch-bench'srqe-optimizerfrom sketch-bench measurements.x = S,y = T), no sharing.traces: ARE 0.05, rank error 0.01, top-k precision 0.95.N_satand KLL/top-k merge curves (Change SimpleStore to have a better indexing for time and label lookups #130, Benchmarking ASAP with Elastic #131, Add experiment scripts to test throughput and cost of ingest path #140), because it merges smaller-window sketches. AutoSketch never merges, so following the paper it reads the measured value at its window size.synthetic, main figure: 10 PromQL templates and a data model withCgroups andsseries per group. A workload grid over query mix, replicas, window set, interval,C,s, θ/a, strictness and SLA.traces, appendix: Alibaba 2022, Google 2011, BOOM, with worst-case parameters fit over each whole trace./api/v1/status/runtimeinfoto QueryEngineRust, but VictoriaMetrics does not support this end point #141 are open, with experiments running.Validation
Documentation only. Factual claims are checked against sketch-bench (
rqe-optimizerat #141's head, the saturation data) and ASAPQuery-backendautosketch_comparison.rs.🤖 Generated with Claude Code