Skip to content

Synthetic workload results: reproduction steps and ASAP's profiling time - #141

Closed
zzylol wants to merge 3 commits into
eval/autosketch-vs-asap-syntheticfrom
eval/two-cost-models
Closed

zzylol wants to merge 3 commits into
eval/autosketch-vs-asap-syntheticfrom
eval/two-cost-models

Conversation

@zzylol

@zzylol zzylol commented Oct 5, 2026 •

Copy link
Copy Markdown
Collaborator

Part of ASAPQuery#777. Stacked on #139.

The two cost models this PR introduced (model A usage, model B peak per EC2 family, the CPU timeline) are gone: #145's weighted objective replaces them, and #138's runner uses it. This PR now keeps:

  • scripts/profiling_time.py and rqe-optimizer/data/profiling-time.json: ASAP's one-time profiling time (#777 §7);
  • rqe-optimizer/results/autosketch-vs-asap-synthetic/README.md: how the reduced grid is generated, run and plotted.

Results are committed here after the full run, which waits for the targeted saturation run (θ = 1.1, K = 1e4, Zipf-rank quantiles) and the cost re-export (#157). A sanity run of all 6 workloads is in #139.

🤖 Generated with Claude Code

zzylol and others added 3 commits October 6, 2026 23:39
Runner for ASAPQuery#777's comparison on main's planner: one metric per
stream of the evaluation table, ASAP by milp::minimize, PerQuery by
minimize per RQE, AutoSketch by autosketch::plan, all scored by
analytical_cost_model::score (#145) at #777's weights (CPU only; Fargate
$/vCPU-h and $/GB-h) and SLA grid {0.01, 0.1, 1, 10} ms and none. Each
RQE's accuracy rides in the cost entries: acc@<id> (with merged values in
merge_accuracy, #154) for ASAP, as@<id> for AutoSketch.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
export_autosketch_eval_table.py --synthetic writes #777's synthetic
workload: fixed data (1e4 series = label_0 keys, 10 jobs, 200 samples/s
per series, Zipf theta 1.1), windows {15m, 1h, 6h, 24h}, the 10 templates
(50 distinct RQEs) and the dashboard set (21), top-k at k = 32. Sum and
rate are served by exact accumulators priced from --exact-costs (the cost
table's exact-sum / exact-increase rows), with zero error at every merge
count. The grid is the default point plus each dimension alone: template
set, shared replicas {1, 8, 64}, strictness. Quantile curves stay on
Pareto a = 2 until the targeted Zipf-rank saturation run.

run_autosketch_vs_asap_synthetic.sh runs plan.tsv; the plot script writes
the objective-vs-latency, per-dimension and planning-time figures and a
summary.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The two cost models are gone (#145's weighted objective replaces them, in
the runner of the PRs below). This keeps ASAP's one-time profiling time
(#777 section 7) and documents how the reduced grid is run; results are
committed after the full run.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@zzylol
zzylol force-pushed the eval/autosketch-vs-asap-synthetic branch from ccf724c to 3bd0b7d Compare October 6, 2026 23:43
@zzylol
zzylol force-pushed the eval/two-cost-models branch from 2cfe0e5 to a6ed461 Compare October 6, 2026 23:43
@zzylol zzylol changed the title Price plans under two cost models and drive the synthetic workload grid Synthetic workload results: reproduction steps and ASAP's profiling time Oct 6, 2026
@zzylol
zzylol force-pushed the eval/autosketch-vs-asap-synthetic branch 2 times, most recently from 115536b to fb1064f Compare October 7, 2026 18:30
zzylol added a commit that referenced this pull request Oct 7, 2026
ProjectASAP/ASAPQuery#777 (paper §6.3), consolidating #138, #139 and #141.

- rqe-optimizer/examples/autosketch_vs_asap.rs: ASAP (milp::minimize),
  PerQuery-CostAware and AutoSketch-Adapted on one synthetic table. Costs and
  accuracy come from --saturation-dir, as the planner reads them:
  SaturationCurves::accuracy for ASAP (merge curves for KLL and top-k,
  #179), autosketch_accuracy for AutoSketch, and the one-study cost table
  (#178). Each run is scored by analytical_cost_model::score at #777's two
  weight settings and the SLA grid.
- scripts/export_autosketch_eval_table.py --synthetic: one workload-only
  table per grid point, plus plan.tsv. The data constants come from
  study_saturation's COST_*.
- scripts/run_autosketch_vs_asap_synthetic.sh and
  plot_autosketch_vs_asap_synthetic.py: run the grid, then plot.
- scripts/profiling_time.py and the synthetic results README: ASAP's
  one-time profiling time, and how to reproduce.

The trace workloads are left out for now.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@zzylol

zzylol commented Oct 7, 2026

Copy link
Copy Markdown
Collaborator Author

Folded into #138, which now holds the runner, the synthetic export/run/plot scripts and #141's profiling script and README in one PR on main (synthetic workload only). Closing.

🤖 Generated with Claude Code

@zzylol zzylol closed this Oct 7, 2026
zzylol added a commit that referenced this pull request Oct 7, 2026
… workloads

ProjectASAP/ASAPQuery#777 (paper §6.3), consolidating #138, #139 and #141.

- rqe-optimizer/examples/autosketch_vs_asap.rs (traces and synthetic): ASAP
  (milp::minimize), PerQuery-CostAware and AutoSketch-Adapted. Costs and
  accuracy come from --saturation-dir, as the planner reads them:
  SaturationCurves::accuracy for ASAP (merge curves for KLL and top-k,
  #179), autosketch_accuracy for AutoSketch, and the one-study cost table
  (#178). Each run is scored by analytical_cost_model::score at #777's two
  weight settings and the SLA grid.
- scripts/run_autosketch_vs_asap.sh and plot_autosketch_vs_asap.py: the
  trace workloads alibaba_v2022 and google_2011. boom is left out for now.
- scripts/export_autosketch_eval_table.py --synthetic: one workload-only
  table per grid point, plus plan.tsv. The data constants come from
  study_saturation's COST_*. The traces path is unchanged from main.
- scripts/run_autosketch_vs_asap_synthetic.sh and
  plot_autosketch_vs_asap_synthetic.py: run the synthetic grid, then plot.
- scripts/profiling_time.py and the results README: ASAP's one-time
  profiling time, and how to reproduce.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
zzylol added a commit that referenced this pull request Oct 7, 2026
… workloads

ProjectASAP/ASAPQuery#777 (paper §6.3), consolidating #138, #139 and #141.

- rqe-optimizer/examples/autosketch_vs_asap.rs (traces and synthetic): ASAP
  (milp::minimize), PerQuery-CostAware and AutoSketch-Adapted. Costs and
  accuracy come from --saturation-dir, as the planner reads them:
  SaturationCurves::accuracy for ASAP (merge curves for KLL and top-k,
  #179), autosketch_accuracy for AutoSketch, and the one-study cost table
  (#178). Each run is scored by analytical_cost_model::score at #777's two
  weight settings and the SLA grid.
- scripts/run_autosketch_vs_asap.sh and plot_autosketch_vs_asap.py: the
  trace workloads alibaba_v2022 and google_2011. boom is left out for now.
- scripts/export_autosketch_eval_table.py --synthetic: one workload-only
  table per grid point, plus plan.tsv. The data constants come from
  study_saturation's COST_*. The traces path is unchanged from main.
- scripts/run_autosketch_vs_asap_synthetic.sh and
  plot_autosketch_vs_asap_synthetic.py: run the synthetic grid, then plot.
- scripts/profiling_time.py and the results README: ASAP's one-time
  profiling time, and how to reproduce.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
zzylol added a commit that referenced this pull request Oct 7, 2026
… workloads

ProjectASAP/ASAPQuery#777 (paper §6.3), consolidating #138, #139 and #141.

- rqe-optimizer/examples/autosketch_vs_asap.rs (traces and synthetic): ASAP
  (milp::minimize), PerQuery-CostAware and AutoSketch-Adapted. Costs and
  accuracy come from --saturation-dir, as the planner reads them:
  SaturationCurves::accuracy for ASAP (merge curves for KLL and top-k,
  #179), autosketch_accuracy for AutoSketch, and the one-study cost table
  (#178). Each run is scored by analytical_cost_model::score at #777's two
  weight settings and the SLA grid.
- scripts/run_autosketch_vs_asap.sh and plot_autosketch_vs_asap.py: the
  trace workloads alibaba_v2022 and google_2011. boom is left out for now.
- scripts/export_autosketch_eval_table.py --synthetic: one workload-only
  table per grid point, plus plan.tsv. The data constants come from
  study_saturation's COST_*. The traces path is unchanged from main.
- scripts/run_autosketch_vs_asap_synthetic.sh and
  plot_autosketch_vs_asap_synthetic.py: run the synthetic grid, then plot.
- scripts/profiling_time.py and the results README: ASAP's one-time
  profiling time, and how to reproduce.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
zzylol added a commit that referenced this pull request Oct 7, 2026
… workloads

ProjectASAP/ASAPQuery#777 (paper §6.3), consolidating #138, #139 and #141.

- rqe-optimizer/examples/autosketch_vs_asap.rs (traces and synthetic): ASAP
  (milp::minimize), PerQuery-CostAware and AutoSketch-Adapted. Costs and
  accuracy come from --saturation-dir, as the planner reads them:
  SaturationCurves::accuracy for ASAP (merge curves for KLL and top-k,
  #179), autosketch_accuracy for AutoSketch, and the one-study cost table
  (#178). Each run is scored by analytical_cost_model::score at #777's two
  weight settings and the SLA grid.
- scripts/run_autosketch_vs_asap.sh and plot_autosketch_vs_asap.py: the
  trace workloads alibaba_v2022 and google_2011. boom is left out for now.
- scripts/export_autosketch_eval_table.py --synthetic: one workload-only
  table per grid point, plus plan.tsv. The data constants come from
  study_saturation's COST_*. The traces path is unchanged from main.
- scripts/run_autosketch_vs_asap_synthetic.sh and
  plot_autosketch_vs_asap_synthetic.py: run the synthetic grid, then plot.
- scripts/profiling_time.py and the results README: ASAP's one-time
  profiling time, and how to reproduce.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
zzylol added a commit that referenced this pull request Oct 7, 2026
… workloads

ProjectASAP/ASAPQuery#777 (paper §6.3), consolidating #138, #139 and #141.

- rqe-optimizer/examples/autosketch_vs_asap.rs (traces and synthetic): ASAP
  (milp::minimize), PerQuery-CostAware and AutoSketch-Adapted. Costs and
  accuracy come from --saturation-dir, as the planner reads them:
  SaturationCurves::accuracy for ASAP (merge curves for KLL and top-k,
  #179), autosketch_accuracy for AutoSketch, and the one-study cost table
  (#178). Each run is scored by analytical_cost_model::score at #777's two
  weight settings and the SLA grid.
- scripts/run_autosketch_vs_asap.sh and plot_autosketch_vs_asap.py: the
  trace workloads alibaba_v2022 and google_2011. boom is left out for now.
- scripts/export_autosketch_eval_table.py --synthetic: one workload-only
  table per grid point, plus plan.tsv. The data constants come from
  study_saturation's COST_*. The traces path is unchanged from main.
- scripts/run_autosketch_vs_asap_synthetic.sh and
  plot_autosketch_vs_asap_synthetic.py: run the synthetic grid, then plot.
- scripts/profiling_time.py and the results README: ASAP's one-time
  profiling time, and how to reproduce.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
zzylol added a commit that referenced this pull request Oct 8, 2026
… workloads

ProjectASAP/ASAPQuery#777 (paper §6.3), consolidating #138, #139 and #141.

- rqe-optimizer/examples/autosketch_vs_asap.rs (traces and synthetic): ASAP
  (milp::minimize), PerQuery-CostAware and AutoSketch-Adapted. Costs and
  accuracy come from --saturation-dir, as the planner reads them:
  SaturationCurves::accuracy for ASAP (merge curves for KLL and top-k,
  #179), autosketch_accuracy for AutoSketch, and the one-study cost table
  (#178). Each run is scored by analytical_cost_model::score at #777's two
  weight settings and the SLA grid.
- scripts/run_autosketch_vs_asap.sh and plot_autosketch_vs_asap.py: the
  trace workloads alibaba_v2022 and google_2011. boom is left out for now.
- scripts/export_autosketch_eval_table.py --synthetic: one workload-only
  table per grid point, plus plan.tsv. The data constants come from
  study_saturation's COST_*. The traces path is unchanged from main.
- scripts/run_autosketch_vs_asap_synthetic.sh and
  plot_autosketch_vs_asap_synthetic.py: run the synthetic grid, then plot.
- scripts/profiling_time.py and the results README: ASAP's one-time
  profiling time, and how to reproduce.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
zzylol added a commit that referenced this pull request Oct 8, 2026
… workloads

ProjectASAP/ASAPQuery#777 (paper §6.3), consolidating #138, #139 and #141.

- rqe-optimizer/examples/autosketch_vs_asap.rs (traces and synthetic): ASAP
  (milp::minimize), PerQuery-CostAware and AutoSketch-Adapted. Costs and
  accuracy come from --saturation-dir, as the planner reads them:
  SaturationCurves::accuracy for ASAP (merge curves for KLL and top-k,
  #179), autosketch_accuracy for AutoSketch, and the one-study cost table
  (#178). Each run is scored by analytical_cost_model::score at #777's two
  weight settings and the SLA grid.
- scripts/run_autosketch_vs_asap.sh and plot_autosketch_vs_asap.py: the
  trace workloads alibaba_v2022 and google_2011. boom is left out for now.
- scripts/export_autosketch_eval_table.py --synthetic: one workload-only
  table per grid point, plus plan.tsv. The data constants come from
  study_saturation's COST_*. The traces path is unchanged from main.
- scripts/run_autosketch_vs_asap_synthetic.sh and
  plot_autosketch_vs_asap_synthetic.py: run the synthetic grid, then plot.
- scripts/profiling_time.py and the results README: ASAP's one-time
  profiling time, and how to reproduce.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
zzylol added a commit that referenced this pull request Oct 9, 2026
… workloads

ProjectASAP/ASAPQuery#777 (paper §6.3), consolidating #138, #139 and #141.

- rqe-optimizer/examples/autosketch_vs_asap.rs (traces and synthetic): ASAP
  (milp::minimize), PerQuery-CostAware and AutoSketch-Adapted. Costs and
  accuracy come from --saturation-dir, as the planner reads them:
  SaturationCurves::accuracy for ASAP (merge curves for KLL and top-k,
  #179), autosketch_accuracy for AutoSketch, and the one-study cost table
  (#178). Each run is scored by analytical_cost_model::score at #777's two
  weight settings and the SLA grid.
- scripts/run_autosketch_vs_asap.sh and plot_autosketch_vs_asap.py: the
  trace workloads alibaba_v2022 and google_2011. boom is left out for now.
- scripts/export_autosketch_eval_table.py --synthetic: one workload-only
  table per grid point, plus plan.tsv. The data constants come from
  study_saturation's COST_*. The traces path is unchanged from main.
- scripts/run_autosketch_vs_asap_synthetic.sh and
  plot_autosketch_vs_asap_synthetic.py: run the synthetic grid, then plot.
- scripts/profiling_time.py and the results README: ASAP's one-time
  profiling time, and how to reproduce.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
zzylol added a commit that referenced this pull request Oct 9, 2026
…ontier and SLA versions) (#138)

* rqe-optimizer: evaluate AutoSketch vs ASAP on the synthetic and trace workloads

ProjectASAP/ASAPQuery#777 (paper §6.3), consolidating #138, #139 and #141.

- rqe-optimizer/examples/autosketch_vs_asap.rs (traces and synthetic): ASAP
  (milp::minimize), PerQuery-CostAware and AutoSketch-Adapted. Costs and
  accuracy come from --saturation-dir, as the planner reads them:
  SaturationCurves::accuracy for ASAP (merge curves for KLL and top-k,
  #179), autosketch_accuracy for AutoSketch, and the one-study cost table
  (#178). Each run is scored by analytical_cost_model::score at #777's two
  weight settings and the SLA grid.
- scripts/run_autosketch_vs_asap.sh and plot_autosketch_vs_asap.py: the
  trace workloads alibaba_v2022 and google_2011. boom is left out for now.
- scripts/export_autosketch_eval_table.py --synthetic: one workload-only
  table per grid point, plus plan.tsv. The data constants come from
  study_saturation's COST_*. The traces path is unchanged from main.
- scripts/run_autosketch_vs_asap_synthetic.sh and
  plot_autosketch_vs_asap_synthetic.py: run the synthetic grid, then plot.
- scripts/profiling_time.py and the results README: ASAP's one-time
  profiling time, and how to reproduce.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* rqe-optimizer: trace key queries keep one accumulator per key; report the accuracy source

From review on #138:
- A traces key query (sum by key) keeps one exact accumulator per key: its
  groups are the window's keys, not the table's one sketch. A stream's
  card(g) is the largest over its RQEs.
- AutoSketch's benchmark lower bound uses one query of one instance, not
  the stream's q_r of them.
- --weights that names none of the weight settings is an error.
- The chosen report lists each RQE's accuracy source (#180's
  accuracy_with_source).
- Drop a duplicated BoomTest; capability_of's comment matches the code.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* eval: one accuracy level, p95; plots join one RQE set and charge AutoSketch's benchmark

- One accuracy level for every workload, synthetic and traces: 95% in
  each family's metric (rank, relative value and relative error at most
  0.05; top-k precision at least 0.95). The strictness dimension is gone.
- fig_objective_vs_latency joins only points over one RQE set (solid for
  the full set, dashed for the subset tight SLAs keep).
- fig_planning_time adds AutoSketch's charged benchmark time (lower bound
  and 60 s per probe), as #777 section 7 counts it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* eval: summaries show Σ window/slide; synthetic summary to its own file

AutoSketch-Adapted's ingest cost is window / slide overlapping sketches per
deployment (window = the query's lookback), so the column explains its gap
to ASAP. The synthetic and trace scripts both wrote summary.md into one
results directory; the synthetic one is now summary_synthetic.md.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* eval: RQEs carry their top-k k (#185)

Raqe now has topk_k. The tables name it only in the query id (topk32_...),
so the runner reads it there and stops on a top-k query that names none.
On the p95 synthetic and trace tables every top-k query is k = 32, and the
plans are unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* eval: p95 results for the synthetic grid and the traces

Runs on clnode155 over the asap_sketchlib 0.3.0 study (#174, #178–#186):
four synthetic workloads and alibaba_v2022 / google_2011, raw JSON,
summaries and figures. The README gives the no-SLA headline, including
Σ window/slide behind AutoSketch-Adapted's gap. profiling-time.json is
refilled from the study's raw records (0.77 h; the K = 1e4 column refill
of #186 is its own run).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* eval: figures and tables give each method's absolute cost

Cost figures label every bar and point with its value; the summaries and
the results README list each method's objective instead of its ratio to
ASAP. The by-dimension figure's legend moves above the panels.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* eval: metrics dimension, one dashboard per metric (21 · m RQEs)

Shared replicas on one metric stop adding RQEs past r = 8 (52 at r = 8,
57 at r = 64), so the grid keeps shared r ∈ {1, 8} and adds metrics
m ∈ {1, 8, 16}: copy i of the template set reads its own metric data_i,
so nothing is shared across copies and RQEs grow as 21 · m (168, 336).
The planning-time figure plots this dimension. The m = 8 and 16 runs are
added; the r = 64 result is removed. Existing tables are unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* eval: AutoSketch benchmarks each config once per metric

Synthetic tables name the metric each RQE reads (data, or data_i in the
metrics dimension), and AutoSketch-Adapted's benchmark time counts a
probed config once per metric, on that metric's data: 540, 4320 and
8640 s (60 s per probe) for m = 1, 8, 16. Trace tables name none and
read one metric per dataset, as before. All five synthetic workloads are
rerun with the current runner on four nodes; objectives are unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* eval: cost vs. total estimated latency for every workload

One column per workload (five synthetic, two traces), one row per weight
setting, one line per method across the SLA grid, each point labeled with
its cost. Total latency is the sum of the RQEs' estimated latencies.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* eval: finer SLA grid, {0.003, 0.01, 0.03, 0.1, 0.3, 1, 3, 10} ms and none

Checks whether AutoSketch-Adapted's memory objective makes it miss an
SLA: it never does. Its plans never merge, so each RQE's latency is within
1.4% of the lowest any deployment reaches (equal for quantile, exact and
trace RQEs), and RQEs no method can meet are excluded for all methods.
All workloads rerun on four nodes; figures and summaries regenerated.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* eval: AutoSketch probe list and measured benchmark time; traces with 6h/24h windows

- The runner lists AutoSketch-Adapted's distinct probed (metric, config)
  pairs with the data shape each is benchmarked on, and drops the
  benchmark-time lower bound. scripts/autosketch_benchmark_time.py runs
  each probe's approxbench accuracy benchmark (1e8 items, exact baseline
  and scoring) and times it, as AutoSketch benchmarks probed configs.
- Trace streams with more keys than samples per second are left out and
  listed (rounding their rate up to their key count inflated every
  method's cost).
- The trace table is regenerated from ASAPQuery #812's skew summary
  (6h and 24h windows, a day of Alibaba).
- Plots: one line per method over the full RQE set; AutoSketch's planning
  time with its measured benchmark.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* eval: cost by use (sketch-bench #188); Pareto frontier sweep; mixed set by default

- The runner prices every plan by use (usage::usage_cost): w_cpu * AUC(CPU)
  + w_mem * AUC(memory), CPU elastic, latency the longest chain, reported.
- ASAP is milp::minimize_usage_cost over every RQE; PerQuery the same over
  each RQE's own candidates (separate copies via `allowed`, no sharing);
  AutoSketch's plan is priced the same way.
- No SLA grid. Each method's cost-latency frontier: ASAP and PerQuery are
  solved unbounded and for 12 bounds log-spaced from the tightest feasible
  one to the unbounded latency, plus AutoSketch's latency.
- The synthetic grid is the mixed template set: shared r in {1, 8} and
  metrics m in {1, 8, 16} (50, 92, 400, 800 RQEs).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* eval: version 1 (no SLA) results on the mixed set, with AutoSketch's measured benchmark

- Four workloads of the mixed template set (50, 92 with r = 8, 400 with
  m = 8, 800 with m = 16 RQEs), cost by use (#188), each method's cheapest
  plan and ASAP's and PerQuery's cost-latency frontiers (12 bounds plus
  AutoSketch's latency), CPU-only and Fargate weights, --runs 3, one node
  per workload. No RQE dropped, no sanity violation.
- AutoSketch's benchmark time is measured: approxbench's accuracy run (1e8
  items, exact baseline, scoring) once per distinct probed config and
  shape, serially on idle clnode155 (21-37 s each), charged once per
  probed (metric, config): 212, 1697 and 3394 s for m = 1, 8, 16.
- plot_autosketch_vs_asap_synthetic.py: fig_frontier.png (cost vs. query
  latency, every point labeled), fig_planning_time.png (AutoSketch as
  search + measured benchmark; 60 s per probe only as a labeled
  reference), summary_synthetic.md. The SLA-era figures, the dashboard
  results and plot_autosketch_vs_asap_workloads.py are removed.
- README: the cost model, version 1, how to reproduce, and the headline
  numbers.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* eval: version 2, a batch latency SLA, on the synthetic mixed set

The runner adds version 2 of the evaluation beside version 1 (no latency
constraint and the frontiers, unchanged in `results`): at each SLA of a
grid (default 100, 300, 1000, 3000, 10000 ms; --slas-ms overrides), ASAP
and PerQuery plan the cheapest plan whose batch latency (the longest
chain) is at most the SLA, via minimize_usage_cost's latency bound, which
is exact with elastic CPU. An SLA below a method's tightest feasible
bound gives an infeasible record. AutoSketch's one plan is recorded at
each SLA with meets_sla. Records go to `sla_results`, with sanity checks
that every plan meets its SLA and that ASAP costs no more than PerQuery.

Results on clnode109 (--runs 3) for the mixed set, r = 8 and m = 8, 16:
no drops, no sanity violations; the tightest feasible SLA is 92 ms, and
AutoSketch (latency 1011 ms) meets only 3000 and 10000 ms. AutoSketch's
planning time is search plus its measured benchmark
(scripts/autosketch_benchmark_time.py: approxbench accuracy runs at 1e8
items, 20-33 s each, serial and alone on the node): 197 s at m = 1,
3159 s at m = 16. scripts/plot_autosketch_vs_asap_sla.py draws
fig_cost_vs_sla.png and summary_sla.md.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* eval: remove the old trace results and figures

They came from the earlier model (per-RQE SLA over a µs grid, mean-CPU
cost); the traces are to be rerun on cost by use with both latency
versions.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* eval: trim the PR to what it needs

- Result JSONs without per-RQE choices (`chosen`): the figures and
  summaries don't read them, and they were nearly all of the 2.9M lines;
  rerun the runner for them.
- Machine files folded into the READMEs (all Intel Xeon E5-2683 v3,
  56 cores); the synthetic run script no longer writes them.
- The SLA-era trace scripts (plot_autosketch_vs_asap.py,
  run_autosketch_vs_asap.sh) are removed; the traces are to be rerun on
  cost by use.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* eval: keep only what the three figures need

The PR's results are the frontier (version 1), cost vs. SLA (version 2)
and planning time on the synthetic mixed set. The trace table goes back
to the base version (the traces are to be rerun on cost by use, with
their new table), and the ASAP profiling-time script and data, which no
figure uses, are removed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* eval: commit the inputs of the three figures

rqe-optimizer/results/autosketch-vs-asap-inputs/ holds what the figures
were computed from: the sketch-bench study on asap_sketchlib 0.3.0 (the
cost table, sketches and exact aggregations, and the accuracy curves the
runner reads as --saturation-dir) and the synthetic mixed-set tables with
plan.tsv. The READMEs point the reproduce steps at them.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* eval: the raw approxbench records behind the figures' study

rqe-optimizer/results/autosketch-vs-asap-inputs/raw/ holds the gzipped
approxbench records (920 KB) that ../saturation/ was reduced from, with a
README mapping each file to its machine, arguments, record count, run time
and output. Rerunning them is 0.77 h of machine time over 5 machines in
parallel (about 20-30 minutes of wall time).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant