Skip to content

eval: AutoSketch vs. ASAP with Hydra candidates and a roll-up ablation (§6.3) - #203

Open
zzylol wants to merge 2 commits into
hydra/eval-integrationfrom
eval/hydra-rollup-results
Open

zzylol wants to merge 2 commits into
hydra/eval-integrationfrom
eval/hydra-rollup-results

Conversation

@zzylol

@zzylol zzylol commented Oct 10, 2026

Copy link
Copy Markdown
Collaborator

Results of the §6.3 synthetic eval (ProjectASAP/ASAPQuery#777) with Hydra candidates and a roll-up ablation. Everything is under rqe-optimizer/results/autosketch-vs-asap-hydra/: figures, result JSONs, summaries, inputs, and a README with the analysis.

Base. The base is hydra/eval-integration, which merges the open stack #195 → #199 → #200 → #202 (with #197 and #198) and #196 → #201. Retarget to main once those merge. The accuracy curves come from #138's study, committed in #194.

Findings

  • Roll-ups (rqe-optimizer: let mergeable finer-group deployments serve coarser RAQEs #190) help ASAP only a little. On the multi-grouping set, 40% of RQEs are served from a finer deployment (42 of 104, and the same share at r=8, m=8, m=16), and ASAP's deployments halve (30 → 14; 480 → 224 at m=16). Cost drops by only 3.3–3.6% under CPU weights and 4.4–4.8% under Fargate, at every frontier point. The classic set has one grouping per stream, so roll-ups never apply there. ASAP's ~3× advantage over PerQuery comes from sharing, not from roll-ups.
  • Hydra is never chosen. It is eligible for many RQEs, but neither ASAP nor PerQuery picks it in any workload, weight setting, bound or SLA. Its insert fans out to every label subset: 1.9–5.1 µs per record, against 3.85 ns for a per-group HLL. A check on real traces (Alibaba 2022, Google 2011) agreed. Per #777's plan, Hydra is therefore not implemented in ASAPQuery.
  • Overall: 0 RQEs dropped and 0 sanity violations in all 8 workloads. ASAP < ASAP without roll-ups < PerQuery << AutoSketch.
Workload Weights ASAP ASAP (no roll-ups) PerQuery AutoSketch
mixed (classic), 50 RQEs CPU (vCPU) 3.54 3.54 10.5 3.72e3
mixed + multi-grouping, 104 CPU 3.79 3.94 11.3 3.72e3
mixed + multi-grouping, 104 Fargate ($/h) 0.245 0.257 0.799 176
mixed + multi-grouping, m=16, 1664 CPU 60.7 63.0 181 5.96e4

Figures

  • fig_frontier.png: cost vs. query latency.
  • fig_cost_vs_sla.png: cost vs. batch-latency SLA.
  • fig_planning_time.png: planning time.
  • fig_rollup_ablation.png: ASAP with vs. without roll-ups.

The README has every table, the method definitions, the inputs, the reproduce commands and the caveats.

🤖 Generated with Claude Code

https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis

zzylol and others added 2 commits October 10, 2026 14:18
Results of the §6.3 synthetic eval (ProjectASAP/ASAPQuery#777) rerun with
roll-ups (#190) and their ablation, Hydra candidates measured on the eval's
own data (#195–#202), the multi-grouping template set (#196), #192's KLL
footprint and a re-measured cost table. 8 workloads, 0 RQEs dropped, 0
sanity violations.

Roll-ups serve 40% of the multi-grouping RQEs from a finer deployment and
halve ASAP's deployments, but save only 3.3–3.6% (CPU) / 4.4–4.8%
(Fargate). Hydra is eligible for many RQEs but never chosen: its insert
fans out to every label subset (1.9–5.1 µs per record vs. 3.85 ns for a
per-group HLL).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis
Template 16 is now p99 latency GROUP BY CUBE(region, service, endpoint),
all 7 groupings, replacing {service}, {region, service}, {service, status}.
A new `multigroup` template set runs the multi-grouping templates alone, so
the roll-up ablation is not diluted by the classic set.

Roll-ups now save 60% on multi-grouping (0.551 -> 0.221 vCPU, 26 -> 5
deployments) and 8% on mixed + multi-grouping. Classic is unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis
@zzylol

zzylol commented Oct 11, 2026

Copy link
Copy Markdown
Collaborator Author

Added f8cecf9: template 16 is now p99 latency GROUP BY CUBE(region, service, endpoint) (all 7 groupings), and a new multigroup set runs the multi-grouping templates without the classic ones. Roll-ups save 60% on multi-grouping (0.551 → 0.221 vCPU, 26 → 5 deployments) and 8% on mixed + multi-grouping. Classic is unchanged, and Hydra is still never chosen. README and ASAPQuery#777 §11 (ba6b4e3) are updated.

🤖 Generated with Claude Code

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant