Repository navigation
Conversation
Results of the §6.3 synthetic eval (ProjectASAP/ASAPQuery#777) rerun with roll-ups (#190) and their ablation, Hydra candidates measured on the eval's own data (#195–#202), the multi-grouping template set (#196), #192's KLL footprint and a re-measured cost table. 8 workloads, 0 RQEs dropped, 0 sanity violations. Roll-ups serve 40% of the multi-grouping RQEs from a finer deployment and halve ASAP's deployments, but save only 3.3–3.6% (CPU) / 4.4–4.8% (Fargate). Hydra is eligible for many RQEs but never chosen: its insert fans out to every label subset (1.9–5.1 µs per record vs. 3.85 ns for a per-group HLL). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis
Template 16 is now p99 latency GROUP BY CUBE(region, service, endpoint),
all 7 groupings, replacing {service}, {region, service}, {service, status}.
A new `multigroup` template set runs the multi-grouping templates alone, so
the roll-up ablation is not diluted by the classic set.
Roll-ups now save 60% on multi-grouping (0.551 -> 0.221 vCPU, 26 -> 5
deployments) and 8% on mixed + multi-grouping. Classic is unchanged.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis
Collaborator
Author
|
Added f8cecf9: template 16 is now p99 latency 🤖 Generated with Claude Code |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Results of the §6.3 synthetic eval (ProjectASAP/ASAPQuery#777) with Hydra candidates and a roll-up ablation. Everything is under
rqe-optimizer/results/autosketch-vs-asap-hydra/: figures, result JSONs, summaries, inputs, and a README with the analysis.Base. The base is
hydra/eval-integration, which merges the open stack #195 → #199 → #200 → #202 (with #197 and #198) and #196 → #201. Retarget to main once those merge. The accuracy curves come from #138's study, committed in #194.Findings
Figures
fig_frontier.png: cost vs. query latency.fig_cost_vs_sla.png: cost vs. batch-latency SLA.fig_planning_time.png: planning time.fig_rollup_ablation.png: ASAP with vs. without roll-ups.The README has every table, the method definitions, the inputs, the reproduce commands and the caveats.
🤖 Generated with Claude Code
https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis