Skip to content

Evaluate the synthetic PromQL workload of ASAPQuery#777 - #139

Closed
zzylol wants to merge 4 commits into
eval/autosketch-vs-asapfrom
eval/autosketch-vs-asap-synthetic
Closed

zzylol wants to merge 4 commits into
eval/autosketch-vs-asapfrom
eval/autosketch-vs-asap-synthetic

Conversation

@zzylol

@zzylol zzylol commented Oct 5, 2026 •

Copy link
Copy Markdown
Collaborator

Part of ASAPQuery#777 (§6 "Synthetic workload"). Stacked on #138. Rebuilt for the revised #777.

Why

#777 now fixes the synthetic data and lists every query; the earlier tables (C = 1e3, s = 100, frequency templates, top-k k = 3, a 9-dimension grid) no longer match.

What

export_autosketch_eval_table.py --synthetic, its tests, run_autosketch_vs_asap_synthetic.sh and plot_autosketch_vs_asap_synthetic.py.

How

  • Data (fixed): 1e4 series (label_0, the top-k key), 10 jobs, 200 samples/s per series, Zipf θ = 1.1; windows {15m, 1h, 6h, 24h}, T = 1 m, spatial S = T = 1 s.
  • Queries: the 10 templates (54 queries, 50 distinct RQEs) and the dashboard set (21), as listed in #777; grouped queries by (job), top-k topk(32, sum by (label_0) (...)).
  • Exact accumulators: sum and rate RQEs are served by exact-sum / exact-increase, priced from --exact-costs (a cost table's rows), zero error at every merge count.
  • Grid: default (dashboard, 1 replica, default strictness) plus each dimension alone: 10 templates; shared replicas 8, 64; strictness loose, strict. 6 runs.

Before / After

  • Before: 67-RQE template set over frequency sketches, C = 1e3, s = 100, plus a grid of hundreds of points.
  • After: #777's queries and data; 6 workloads; every summary sees ≥ 1.8e5 items.

Evidence (sanity run, local, --runs 1)

Inputs: saturation curves (out_grid_1e7_cost, out_1e9, #140's K = 1e4 points), merge curves (#131), exact rows from the 2026-10-06 cost export. CPU-only weights, no SLA (vCPU):

workload RQEs ASAP PerQuery AutoSketch planning (ASAP) sanity violations
dashboard, default 21 3.13 7.55 1050 0.08 s 0
10 templates 50 6.10 23.5 8963 0.43 s 0
dashboard, shared 8 52 9.30 32.0 7435 0.49 s 0
dashboard, shared 64 57 9.32 34.2 7561 0.54 s 0
dashboard, loose 21 2.77 6.77 1044 0.13 s 0
dashboard, strict 21 3.14 7.57 1130 0.07 s 0

AutoSketch's cost is dominated by ingest into S/T overlapping windows (up to 1440 for 24 h every 1 m); ASAP tiles and merges, at up to ~5–8 s of serial latency with no SLA.

Verification

  • Unit tests (scripts/test_export_autosketch_eval_table.py, 14 pass): exact families from --exact-costs; 50 distinct RQEs on 8 streams; items per instance from the data model; one query per instance; the 6-point plan; 21 dashboard RQEs; seeded, deduplicated shared replicas.
  • End-to-end: the 6 runs above.

Architectural decisions

  • Template 10 and D5 reuse template 5 / D1's RQEs, so they share rather than duplicate.
  • Disjoint replicas (spatial filters) dropped; the planner rejects filters.

Limitations and follow-up

  • Quantile curves are Pareto a = 2; #777 asks for a targeted run on the Zipf ranks at θ = 1.1, K = 1e4.
  • Shared replicas saturate at ~57 distinct RQEs (3 quantiles × 4 windows × 3 intervals).

Human review — do not complete with an agent

  • The MVP boundary is correct.
  • New conceptual layers or public interfaces are necessary.
  • The before/after description matches the intended product behavior.
  • Human reviewer:
  • Decision and rationale:

🤖 Generated with Claude Code

@zzylol zzylol changed the title Evaluate the synthetic PromQL workload with FewestPlans and SLA-bound strawmen Evaluate the synthetic PromQL workload with an SLA-bound PerQuery baseline Oct 5, 2026
@zzylol
zzylol force-pushed the eval/autosketch-vs-asap branch from b12111f to f293b73 Compare October 5, 2026 15:48
@zzylol
zzylol force-pushed the eval/autosketch-vs-asap-synthetic branch from 4460007 to ccf724c Compare October 5, 2026 15:53
@zzylol
zzylol force-pushed the eval/autosketch-vs-asap branch from f293b73 to 8b59466 Compare October 6, 2026 23:43
@zzylol
zzylol force-pushed the eval/autosketch-vs-asap-synthetic branch from ccf724c to 3bd0b7d Compare October 6, 2026 23:43
@zzylol zzylol changed the title Evaluate the synthetic PromQL workload with an SLA-bound PerQuery baseline Evaluate the synthetic PromQL workload of ASAPQuery#777 Oct 6, 2026
@zzylol
zzylol force-pushed the eval/autosketch-vs-asap branch from 8b59466 to 944f8fc Compare October 7, 2026 17:52
@zzylol
zzylol force-pushed the eval/autosketch-vs-asap branch 2 times, most recently from 3397af3 to c76d555 Compare October 7, 2026 18:15
@zzylol
zzylol force-pushed the eval/autosketch-vs-asap-synthetic branch from 3bd0b7d to 115536b Compare October 7, 2026 18:29
zzylol and others added 4 commits October 7, 2026 18:30
Since #171 every sketch read its plain curve at the lookback's item count,
whatever the window, so a KLL or top-k answer merged from L/x windows got
the unmerged sketch's accuracy. SaturationCurves now loads
saturation_merge_curve.csv and reads it for those two families at
m = L/x: the worse of the measured shard counts either side, the largest
past it, and no accuracy without a merge curve. A later run's point keeps
an earlier run's merge curves when it has none (out_1e9 measures none).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…over

From review on #179:
- Past a curve's last checkpoint, its plateau only if n_sat is at or before
  it: a merge curve stops at 1e7, while its point's n_sat may come from the
  1e9 run.
- A later run keeps an earlier run's merge curves at shard counts it didn't
  measure, moved rather than cloned.
- load refuses a candidate KLL or top-k point with no merge curve, rather
  than leaving every merged deployment silently without accuracy.
- shards must be a whole count >= 1; univmon-topk merges lossily too.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Restacked on #178 and #179. The runner reads the planner's own inputs from
--saturation-dir: the one-study cost table (#174) and the saturation
curves. ASAP's accuracy is SaturationCurves::accuracy (KLL and top-k
merged from L/x read the merge curves, #158), AutoSketch's
autosketch_accuracy. The table gives the workload and each stream's data
shape (grid_param/grid_K). Per-RQE accuracies are no longer carried in cost
entries (acc@/as@), and MILP solutions are read as PlannedDeployment.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Restacked on #138, which reads costs and accuracy from the saturation study
like the planner (#174). The synthetic tables now hold the workload only:
RQEs, streams, targets and each family's data shape. The data is the cost
table's, imported from study_saturation (COST_THETA, COST_KEYS = C,
COST_PARETO_ALPHA for quantiles). Sum and rate queries name exact
accumulators. --synthetic takes no curves, costs or --exact-costs. The
traces mode is unchanged, and per-query CPU from cost records is dropped
because the runner no longer reads it. The run script passes
SATURATION_DIR to the runner.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@zzylol
zzylol force-pushed the eval/autosketch-vs-asap branch from c76d555 to da8259d Compare October 7, 2026 18:30
@zzylol
zzylol force-pushed the eval/autosketch-vs-asap-synthetic branch from 115536b to fb1064f Compare October 7, 2026 18:30
zzylol added a commit that referenced this pull request Oct 7, 2026
ProjectASAP/ASAPQuery#777 (paper §6.3), consolidating #138, #139 and #141.

- rqe-optimizer/examples/autosketch_vs_asap.rs: ASAP (milp::minimize),
  PerQuery-CostAware and AutoSketch-Adapted on one synthetic table. Costs and
  accuracy come from --saturation-dir, as the planner reads them:
  SaturationCurves::accuracy for ASAP (merge curves for KLL and top-k,
  #179), autosketch_accuracy for AutoSketch, and the one-study cost table
  (#178). Each run is scored by analytical_cost_model::score at #777's two
  weight settings and the SLA grid.
- scripts/export_autosketch_eval_table.py --synthetic: one workload-only
  table per grid point, plus plan.tsv. The data constants come from
  study_saturation's COST_*.
- scripts/run_autosketch_vs_asap_synthetic.sh and
  plot_autosketch_vs_asap_synthetic.py: run the grid, then plot.
- scripts/profiling_time.py and the synthetic results README: ASAP's
  one-time profiling time, and how to reproduce.

The trace workloads are left out for now.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@zzylol
zzylol force-pushed the eval/autosketch-vs-asap branch from da8259d to 412f915 Compare October 7, 2026 19:00
@zzylol

zzylol commented Oct 7, 2026

Copy link
Copy Markdown
Collaborator Author

Folded into #138, which now holds the runner, the synthetic export/run/plot scripts and #141's profiling script and README in one PR on main (synthetic workload only). Closing.

🤖 Generated with Claude Code

@zzylol zzylol closed this Oct 7, 2026
zzylol added a commit that referenced this pull request Oct 7, 2026
… workloads

ProjectASAP/ASAPQuery#777 (paper §6.3), consolidating #138, #139 and #141.

- rqe-optimizer/examples/autosketch_vs_asap.rs (traces and synthetic): ASAP
  (milp::minimize), PerQuery-CostAware and AutoSketch-Adapted. Costs and
  accuracy come from --saturation-dir, as the planner reads them:
  SaturationCurves::accuracy for ASAP (merge curves for KLL and top-k,
  #179), autosketch_accuracy for AutoSketch, and the one-study cost table
  (#178). Each run is scored by analytical_cost_model::score at #777's two
  weight settings and the SLA grid.
- scripts/run_autosketch_vs_asap.sh and plot_autosketch_vs_asap.py: the
  trace workloads alibaba_v2022 and google_2011. boom is left out for now.
- scripts/export_autosketch_eval_table.py --synthetic: one workload-only
  table per grid point, plus plan.tsv. The data constants come from
  study_saturation's COST_*. The traces path is unchanged from main.
- scripts/run_autosketch_vs_asap_synthetic.sh and
  plot_autosketch_vs_asap_synthetic.py: run the synthetic grid, then plot.
- scripts/profiling_time.py and the results README: ASAP's one-time
  profiling time, and how to reproduce.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
zzylol added a commit that referenced this pull request Oct 7, 2026
… workloads

ProjectASAP/ASAPQuery#777 (paper §6.3), consolidating #138, #139 and #141.

- rqe-optimizer/examples/autosketch_vs_asap.rs (traces and synthetic): ASAP
  (milp::minimize), PerQuery-CostAware and AutoSketch-Adapted. Costs and
  accuracy come from --saturation-dir, as the planner reads them:
  SaturationCurves::accuracy for ASAP (merge curves for KLL and top-k,
  #179), autosketch_accuracy for AutoSketch, and the one-study cost table
  (#178). Each run is scored by analytical_cost_model::score at #777's two
  weight settings and the SLA grid.
- scripts/run_autosketch_vs_asap.sh and plot_autosketch_vs_asap.py: the
  trace workloads alibaba_v2022 and google_2011. boom is left out for now.
- scripts/export_autosketch_eval_table.py --synthetic: one workload-only
  table per grid point, plus plan.tsv. The data constants come from
  study_saturation's COST_*. The traces path is unchanged from main.
- scripts/run_autosketch_vs_asap_synthetic.sh and
  plot_autosketch_vs_asap_synthetic.py: run the synthetic grid, then plot.
- scripts/profiling_time.py and the results README: ASAP's one-time
  profiling time, and how to reproduce.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
zzylol added a commit that referenced this pull request Oct 7, 2026
… workloads

ProjectASAP/ASAPQuery#777 (paper §6.3), consolidating #138, #139 and #141.

- rqe-optimizer/examples/autosketch_vs_asap.rs (traces and synthetic): ASAP
  (milp::minimize), PerQuery-CostAware and AutoSketch-Adapted. Costs and
  accuracy come from --saturation-dir, as the planner reads them:
  SaturationCurves::accuracy for ASAP (merge curves for KLL and top-k,
  #179), autosketch_accuracy for AutoSketch, and the one-study cost table
  (#178). Each run is scored by analytical_cost_model::score at #777's two
  weight settings and the SLA grid.
- scripts/run_autosketch_vs_asap.sh and plot_autosketch_vs_asap.py: the
  trace workloads alibaba_v2022 and google_2011. boom is left out for now.
- scripts/export_autosketch_eval_table.py --synthetic: one workload-only
  table per grid point, plus plan.tsv. The data constants come from
  study_saturation's COST_*. The traces path is unchanged from main.
- scripts/run_autosketch_vs_asap_synthetic.sh and
  plot_autosketch_vs_asap_synthetic.py: run the synthetic grid, then plot.
- scripts/profiling_time.py and the results README: ASAP's one-time
  profiling time, and how to reproduce.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
zzylol added a commit that referenced this pull request Oct 7, 2026
… workloads

ProjectASAP/ASAPQuery#777 (paper §6.3), consolidating #138, #139 and #141.

- rqe-optimizer/examples/autosketch_vs_asap.rs (traces and synthetic): ASAP
  (milp::minimize), PerQuery-CostAware and AutoSketch-Adapted. Costs and
  accuracy come from --saturation-dir, as the planner reads them:
  SaturationCurves::accuracy for ASAP (merge curves for KLL and top-k,
  #179), autosketch_accuracy for AutoSketch, and the one-study cost table
  (#178). Each run is scored by analytical_cost_model::score at #777's two
  weight settings and the SLA grid.
- scripts/run_autosketch_vs_asap.sh and plot_autosketch_vs_asap.py: the
  trace workloads alibaba_v2022 and google_2011. boom is left out for now.
- scripts/export_autosketch_eval_table.py --synthetic: one workload-only
  table per grid point, plus plan.tsv. The data constants come from
  study_saturation's COST_*. The traces path is unchanged from main.
- scripts/run_autosketch_vs_asap_synthetic.sh and
  plot_autosketch_vs_asap_synthetic.py: run the synthetic grid, then plot.
- scripts/profiling_time.py and the results README: ASAP's one-time
  profiling time, and how to reproduce.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
zzylol added a commit that referenced this pull request Oct 7, 2026
… workloads

ProjectASAP/ASAPQuery#777 (paper §6.3), consolidating #138, #139 and #141.

- rqe-optimizer/examples/autosketch_vs_asap.rs (traces and synthetic): ASAP
  (milp::minimize), PerQuery-CostAware and AutoSketch-Adapted. Costs and
  accuracy come from --saturation-dir, as the planner reads them:
  SaturationCurves::accuracy for ASAP (merge curves for KLL and top-k,
  #179), autosketch_accuracy for AutoSketch, and the one-study cost table
  (#178). Each run is scored by analytical_cost_model::score at #777's two
  weight settings and the SLA grid.
- scripts/run_autosketch_vs_asap.sh and plot_autosketch_vs_asap.py: the
  trace workloads alibaba_v2022 and google_2011. boom is left out for now.
- scripts/export_autosketch_eval_table.py --synthetic: one workload-only
  table per grid point, plus plan.tsv. The data constants come from
  study_saturation's COST_*. The traces path is unchanged from main.
- scripts/run_autosketch_vs_asap_synthetic.sh and
  plot_autosketch_vs_asap_synthetic.py: run the synthetic grid, then plot.
- scripts/profiling_time.py and the results README: ASAP's one-time
  profiling time, and how to reproduce.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
zzylol added a commit that referenced this pull request Oct 8, 2026
… workloads

ProjectASAP/ASAPQuery#777 (paper §6.3), consolidating #138, #139 and #141.

- rqe-optimizer/examples/autosketch_vs_asap.rs (traces and synthetic): ASAP
  (milp::minimize), PerQuery-CostAware and AutoSketch-Adapted. Costs and
  accuracy come from --saturation-dir, as the planner reads them:
  SaturationCurves::accuracy for ASAP (merge curves for KLL and top-k,
  #179), autosketch_accuracy for AutoSketch, and the one-study cost table
  (#178). Each run is scored by analytical_cost_model::score at #777's two
  weight settings and the SLA grid.
- scripts/run_autosketch_vs_asap.sh and plot_autosketch_vs_asap.py: the
  trace workloads alibaba_v2022 and google_2011. boom is left out for now.
- scripts/export_autosketch_eval_table.py --synthetic: one workload-only
  table per grid point, plus plan.tsv. The data constants come from
  study_saturation's COST_*. The traces path is unchanged from main.
- scripts/run_autosketch_vs_asap_synthetic.sh and
  plot_autosketch_vs_asap_synthetic.py: run the synthetic grid, then plot.
- scripts/profiling_time.py and the results README: ASAP's one-time
  profiling time, and how to reproduce.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
zzylol added a commit that referenced this pull request Oct 8, 2026
… workloads

ProjectASAP/ASAPQuery#777 (paper §6.3), consolidating #138, #139 and #141.

- rqe-optimizer/examples/autosketch_vs_asap.rs (traces and synthetic): ASAP
  (milp::minimize), PerQuery-CostAware and AutoSketch-Adapted. Costs and
  accuracy come from --saturation-dir, as the planner reads them:
  SaturationCurves::accuracy for ASAP (merge curves for KLL and top-k,
  #179), autosketch_accuracy for AutoSketch, and the one-study cost table
  (#178). Each run is scored by analytical_cost_model::score at #777's two
  weight settings and the SLA grid.
- scripts/run_autosketch_vs_asap.sh and plot_autosketch_vs_asap.py: the
  trace workloads alibaba_v2022 and google_2011. boom is left out for now.
- scripts/export_autosketch_eval_table.py --synthetic: one workload-only
  table per grid point, plus plan.tsv. The data constants come from
  study_saturation's COST_*. The traces path is unchanged from main.
- scripts/run_autosketch_vs_asap_synthetic.sh and
  plot_autosketch_vs_asap_synthetic.py: run the synthetic grid, then plot.
- scripts/profiling_time.py and the results README: ASAP's one-time
  profiling time, and how to reproduce.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
zzylol added a commit that referenced this pull request Oct 9, 2026
… workloads

ProjectASAP/ASAPQuery#777 (paper §6.3), consolidating #138, #139 and #141.

- rqe-optimizer/examples/autosketch_vs_asap.rs (traces and synthetic): ASAP
  (milp::minimize), PerQuery-CostAware and AutoSketch-Adapted. Costs and
  accuracy come from --saturation-dir, as the planner reads them:
  SaturationCurves::accuracy for ASAP (merge curves for KLL and top-k,
  #179), autosketch_accuracy for AutoSketch, and the one-study cost table
  (#178). Each run is scored by analytical_cost_model::score at #777's two
  weight settings and the SLA grid.
- scripts/run_autosketch_vs_asap.sh and plot_autosketch_vs_asap.py: the
  trace workloads alibaba_v2022 and google_2011. boom is left out for now.
- scripts/export_autosketch_eval_table.py --synthetic: one workload-only
  table per grid point, plus plan.tsv. The data constants come from
  study_saturation's COST_*. The traces path is unchanged from main.
- scripts/run_autosketch_vs_asap_synthetic.sh and
  plot_autosketch_vs_asap_synthetic.py: run the synthetic grid, then plot.
- scripts/profiling_time.py and the results README: ASAP's one-time
  profiling time, and how to reproduce.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
zzylol added a commit that referenced this pull request Oct 9, 2026
…ontier and SLA versions) (#138)

* rqe-optimizer: evaluate AutoSketch vs ASAP on the synthetic and trace workloads

ProjectASAP/ASAPQuery#777 (paper §6.3), consolidating #138, #139 and #141.

- rqe-optimizer/examples/autosketch_vs_asap.rs (traces and synthetic): ASAP
  (milp::minimize), PerQuery-CostAware and AutoSketch-Adapted. Costs and
  accuracy come from --saturation-dir, as the planner reads them:
  SaturationCurves::accuracy for ASAP (merge curves for KLL and top-k,
  #179), autosketch_accuracy for AutoSketch, and the one-study cost table
  (#178). Each run is scored by analytical_cost_model::score at #777's two
  weight settings and the SLA grid.
- scripts/run_autosketch_vs_asap.sh and plot_autosketch_vs_asap.py: the
  trace workloads alibaba_v2022 and google_2011. boom is left out for now.
- scripts/export_autosketch_eval_table.py --synthetic: one workload-only
  table per grid point, plus plan.tsv. The data constants come from
  study_saturation's COST_*. The traces path is unchanged from main.
- scripts/run_autosketch_vs_asap_synthetic.sh and
  plot_autosketch_vs_asap_synthetic.py: run the synthetic grid, then plot.
- scripts/profiling_time.py and the results README: ASAP's one-time
  profiling time, and how to reproduce.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* rqe-optimizer: trace key queries keep one accumulator per key; report the accuracy source

From review on #138:
- A traces key query (sum by key) keeps one exact accumulator per key: its
  groups are the window's keys, not the table's one sketch. A stream's
  card(g) is the largest over its RQEs.
- AutoSketch's benchmark lower bound uses one query of one instance, not
  the stream's q_r of them.
- --weights that names none of the weight settings is an error.
- The chosen report lists each RQE's accuracy source (#180's
  accuracy_with_source).
- Drop a duplicated BoomTest; capability_of's comment matches the code.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* eval: one accuracy level, p95; plots join one RQE set and charge AutoSketch's benchmark

- One accuracy level for every workload, synthetic and traces: 95% in
  each family's metric (rank, relative value and relative error at most
  0.05; top-k precision at least 0.95). The strictness dimension is gone.
- fig_objective_vs_latency joins only points over one RQE set (solid for
  the full set, dashed for the subset tight SLAs keep).
- fig_planning_time adds AutoSketch's charged benchmark time (lower bound
  and 60 s per probe), as #777 section 7 counts it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* eval: summaries show Σ window/slide; synthetic summary to its own file

AutoSketch-Adapted's ingest cost is window / slide overlapping sketches per
deployment (window = the query's lookback), so the column explains its gap
to ASAP. The synthetic and trace scripts both wrote summary.md into one
results directory; the synthetic one is now summary_synthetic.md.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* eval: RQEs carry their top-k k (#185)

Raqe now has topk_k. The tables name it only in the query id (topk32_...),
so the runner reads it there and stops on a top-k query that names none.
On the p95 synthetic and trace tables every top-k query is k = 32, and the
plans are unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* eval: p95 results for the synthetic grid and the traces

Runs on clnode155 over the asap_sketchlib 0.3.0 study (#174, #178–#186):
four synthetic workloads and alibaba_v2022 / google_2011, raw JSON,
summaries and figures. The README gives the no-SLA headline, including
Σ window/slide behind AutoSketch-Adapted's gap. profiling-time.json is
refilled from the study's raw records (0.77 h; the K = 1e4 column refill
of #186 is its own run).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* eval: figures and tables give each method's absolute cost

Cost figures label every bar and point with its value; the summaries and
the results README list each method's objective instead of its ratio to
ASAP. The by-dimension figure's legend moves above the panels.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* eval: metrics dimension, one dashboard per metric (21 · m RQEs)

Shared replicas on one metric stop adding RQEs past r = 8 (52 at r = 8,
57 at r = 64), so the grid keeps shared r ∈ {1, 8} and adds metrics
m ∈ {1, 8, 16}: copy i of the template set reads its own metric data_i,
so nothing is shared across copies and RQEs grow as 21 · m (168, 336).
The planning-time figure plots this dimension. The m = 8 and 16 runs are
added; the r = 64 result is removed. Existing tables are unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* eval: AutoSketch benchmarks each config once per metric

Synthetic tables name the metric each RQE reads (data, or data_i in the
metrics dimension), and AutoSketch-Adapted's benchmark time counts a
probed config once per metric, on that metric's data: 540, 4320 and
8640 s (60 s per probe) for m = 1, 8, 16. Trace tables name none and
read one metric per dataset, as before. All five synthetic workloads are
rerun with the current runner on four nodes; objectives are unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* eval: cost vs. total estimated latency for every workload

One column per workload (five synthetic, two traces), one row per weight
setting, one line per method across the SLA grid, each point labeled with
its cost. Total latency is the sum of the RQEs' estimated latencies.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* eval: finer SLA grid, {0.003, 0.01, 0.03, 0.1, 0.3, 1, 3, 10} ms and none

Checks whether AutoSketch-Adapted's memory objective makes it miss an
SLA: it never does. Its plans never merge, so each RQE's latency is within
1.4% of the lowest any deployment reaches (equal for quantile, exact and
trace RQEs), and RQEs no method can meet are excluded for all methods.
All workloads rerun on four nodes; figures and summaries regenerated.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* eval: AutoSketch probe list and measured benchmark time; traces with 6h/24h windows

- The runner lists AutoSketch-Adapted's distinct probed (metric, config)
  pairs with the data shape each is benchmarked on, and drops the
  benchmark-time lower bound. scripts/autosketch_benchmark_time.py runs
  each probe's approxbench accuracy benchmark (1e8 items, exact baseline
  and scoring) and times it, as AutoSketch benchmarks probed configs.
- Trace streams with more keys than samples per second are left out and
  listed (rounding their rate up to their key count inflated every
  method's cost).
- The trace table is regenerated from ASAPQuery #812's skew summary
  (6h and 24h windows, a day of Alibaba).
- Plots: one line per method over the full RQE set; AutoSketch's planning
  time with its measured benchmark.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* eval: cost by use (sketch-bench #188); Pareto frontier sweep; mixed set by default

- The runner prices every plan by use (usage::usage_cost): w_cpu * AUC(CPU)
  + w_mem * AUC(memory), CPU elastic, latency the longest chain, reported.
- ASAP is milp::minimize_usage_cost over every RQE; PerQuery the same over
  each RQE's own candidates (separate copies via `allowed`, no sharing);
  AutoSketch's plan is priced the same way.
- No SLA grid. Each method's cost-latency frontier: ASAP and PerQuery are
  solved unbounded and for 12 bounds log-spaced from the tightest feasible
  one to the unbounded latency, plus AutoSketch's latency.
- The synthetic grid is the mixed template set: shared r in {1, 8} and
  metrics m in {1, 8, 16} (50, 92, 400, 800 RQEs).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* eval: version 1 (no SLA) results on the mixed set, with AutoSketch's measured benchmark

- Four workloads of the mixed template set (50, 92 with r = 8, 400 with
  m = 8, 800 with m = 16 RQEs), cost by use (#188), each method's cheapest
  plan and ASAP's and PerQuery's cost-latency frontiers (12 bounds plus
  AutoSketch's latency), CPU-only and Fargate weights, --runs 3, one node
  per workload. No RQE dropped, no sanity violation.
- AutoSketch's benchmark time is measured: approxbench's accuracy run (1e8
  items, exact baseline, scoring) once per distinct probed config and
  shape, serially on idle clnode155 (21-37 s each), charged once per
  probed (metric, config): 212, 1697 and 3394 s for m = 1, 8, 16.
- plot_autosketch_vs_asap_synthetic.py: fig_frontier.png (cost vs. query
  latency, every point labeled), fig_planning_time.png (AutoSketch as
  search + measured benchmark; 60 s per probe only as a labeled
  reference), summary_synthetic.md. The SLA-era figures, the dashboard
  results and plot_autosketch_vs_asap_workloads.py are removed.
- README: the cost model, version 1, how to reproduce, and the headline
  numbers.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* eval: version 2, a batch latency SLA, on the synthetic mixed set

The runner adds version 2 of the evaluation beside version 1 (no latency
constraint and the frontiers, unchanged in `results`): at each SLA of a
grid (default 100, 300, 1000, 3000, 10000 ms; --slas-ms overrides), ASAP
and PerQuery plan the cheapest plan whose batch latency (the longest
chain) is at most the SLA, via minimize_usage_cost's latency bound, which
is exact with elastic CPU. An SLA below a method's tightest feasible
bound gives an infeasible record. AutoSketch's one plan is recorded at
each SLA with meets_sla. Records go to `sla_results`, with sanity checks
that every plan meets its SLA and that ASAP costs no more than PerQuery.

Results on clnode109 (--runs 3) for the mixed set, r = 8 and m = 8, 16:
no drops, no sanity violations; the tightest feasible SLA is 92 ms, and
AutoSketch (latency 1011 ms) meets only 3000 and 10000 ms. AutoSketch's
planning time is search plus its measured benchmark
(scripts/autosketch_benchmark_time.py: approxbench accuracy runs at 1e8
items, 20-33 s each, serial and alone on the node): 197 s at m = 1,
3159 s at m = 16. scripts/plot_autosketch_vs_asap_sla.py draws
fig_cost_vs_sla.png and summary_sla.md.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* eval: remove the old trace results and figures

They came from the earlier model (per-RQE SLA over a µs grid, mean-CPU
cost); the traces are to be rerun on cost by use with both latency
versions.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* eval: trim the PR to what it needs

- Result JSONs without per-RQE choices (`chosen`): the figures and
  summaries don't read them, and they were nearly all of the 2.9M lines;
  rerun the runner for them.
- Machine files folded into the READMEs (all Intel Xeon E5-2683 v3,
  56 cores); the synthetic run script no longer writes them.
- The SLA-era trace scripts (plot_autosketch_vs_asap.py,
  run_autosketch_vs_asap.sh) are removed; the traces are to be rerun on
  cost by use.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* eval: keep only what the three figures need

The PR's results are the frontier (version 1), cost vs. SLA (version 2)
and planning time on the synthetic mixed set. The trace table goes back
to the base version (the traces are to be rerun on cost by use, with
their new table), and the ASAP profiling-time script and data, which no
figure uses, are removed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* eval: commit the inputs of the three figures

rqe-optimizer/results/autosketch-vs-asap-inputs/ holds what the figures
were computed from: the sketch-bench study on asap_sketchlib 0.3.0 (the
cost table, sketches and exact aggregations, and the accuracy curves the
runner reads as --saturation-dir) and the synthetic mixed-set tables with
plan.tsv. The READMEs point the reproduce steps at them.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

* eval: the raw approxbench records behind the figures' study

rqe-optimizer/results/autosketch-vs-asap-inputs/raw/ holds the gzipped
approxbench records (920 KB) that ../saturation/ was reduced from, with a
README mapping each file to its machine, arguments, record count, run time
and output. Rerunning them is 0.77 h of machine time over 5 machines in
parallel (about 20-30 minutes of wall time).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant