Repository navigation
rqe-optimizer: price plans per EC2 machine family - #137
Merged
Merged
Conversation
Score retained memory per active deployment, (x + max S) / y instances, and add milp::minimize_cost, which minimizes the hourly price on one EC2 machine family in fractional instances (whichever of CPU or memory binds). Prices are a committed on-demand snapshot from scripts/fetch_ec2_pricing.py. Dominance pruning also compares retained memory so it cannot drop a candidate that is cheaper under the new objective. Part of ProjectASAP/ASAPQuery#777. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
HiGHS's absolute gap (1e-6) and feasibility tolerance (1e-7) exceed real plan costs (~1e-6 $/hour) and query latencies (µs). Rows are now scaled by the plan's demand without sharing, and per-RQE memory and latency bounds become f64 exclusions of the choices over them. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
zzylol
added a commit
that referenced
this pull request
Oct 5, 2026
Replace the alpha × fastest-latency limit with one absolute SLA per RQE, swept over 0.01 ms to 1 s plus none. RQEs no method can serve within an SLA are excluded from all methods and listed. Drop the runner's demand scaling now that minimize_cost scales internally (#137). Rerun W0, W2 and W1 sequentially on an idle CloudLab node and regenerate figures, summary and README. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
3 tasks
Collaborator
Author
|
@milindsrivastava1997 I reviewed it and the changes include MILP files. I plan to merge this to sketch-bench? |
zzylol
added a commit
that referenced
this pull request
Oct 5, 2026
Runner (examples/autosketch_vs_asap.rs) for the traces workloads of ProjectASAP/ASAPQuery#777: ASAP (joint MILP with one absolute latency SLA per sweep point), PerQuery-CostAware and AutoSketch-Adapted, scored by objectives::score and priced per EC2 family, with sanity checks. Accuracy now depends on the query's merge count m = S/x: a cost entry may carry "{metric}@m{b}" keys, eligibility reads the smallest measured b >= m, and a merge larger than every measured b is ineligible. Results for Alibaba 2022, Google 2011 and BOOM with figures and summary. The earlier example and scaling workloads are dropped (#777 §6). Rebased onto main, which now has #137, #135 and #136. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Before this PR
rqe-optimizerminimized total CPU only. Memory was tracked as peak per-query memory, which does not count the sketch state a deployment keeps over a long lookback. Nothing priced CPU against memory.After this PR
This implements §4 of the AutoSketch-vs-planner evaluation plan (ProjectASAP/ASAPQuery#777, PR 1 in its §8).
Retained memory.
Objectives::retained_memory_bytessums, over active deployments,card × mem_per_instance × (x + max S) / y. That is thex/yopen instances plus the closed ones still inside the longest lookback the deployment serves.Deployment::retained_instance_countgives the per-RQE count.EC2 pricing.
milp::minimize_cost(rqes, deployments, label_sets, bounds, family)minimizesprice_f · n_fsubject ton_f ≥ CPU / vCPU_fandn_f ≥ retained GiB / GiB_f. So it pays for whichever of CPU or memory runs out first. Retained memory is linearized per deployment withR_D ≥ retained_{i,D} · z_{i,D}. The existing latency and peak-memory bounds still apply.minimize_tcokeeps its signature and behavior, and both share one model builder.Prices.
rqe-optimizer/data/ec2-pricing-2026-10-04.jsonis written byscripts/fetch_ec2_pricing.pyfrom the public file behind aws.amazon.com/ec2/pricing/on-demand (us-east-1, Linux, on-demand, published 2026-09-25):Disk is not priced.
MachineFamily::from_pricing_json,instancesandusd_per_hourread the snapshot and score any mapping.Dominance pruning now also requires the replacement to retain no more memory for each RQE the removed candidate serves. Without this, it could drop a candidate that is cheaper under the new objective. A regression test covers a case the old rule pruned.
Example.
small_problem --milp --machine-family NAMEsolves the cost model.Docs.
docs/rqe_sketch_deployment_v1.mdanddocs/rqe_optimizer_TODO.mddescribe the new objective.Coordination: Milind is porting
milp.rsinto ASAPQuery's planner. This adds a function rather than changingminimize_tco, so the port can take it separately.Validation
cargo test -p rqe-optimizer: 18 passed. New tests:minimize_costmatching an exhaustive search on a tiny workload for two families;cargo clippy -p rqe-optimizer --all-targets -- -D warningsandcargo fmt --checkare clean.Ran
small_problem --milpend to end on a freshly exported cost table (scripts/export_rqe_optimizer_costs.sh, 18 rows):On this toy workload, retained state binds in every family. Memory-optimized is cheapest, and the cost plan uses about 40% more CPU than the CPU-only plan (1.18e-2 vs. 8.39e-3).
Update: solver scaling (cdad511)
Problem. The solver works to fixed tolerances that are larger than this model's numbers:
So on small workloads
minimize_costcould stop at a plan several times more expensive than the best one, and could pick a deployment whose latency was slightly over its bound. Neither showed up as an error.Fix, inside
solve, so callers don't change:referenceis what the plan would cost with no sharing, each RQE on its cheapest eligible deployment, in the units being minimized. The CPU, memory and instance-count terms are divided by it, so the optimum is close to 1 and the 1e-6 gap is negligible. The returned plan is still scored on real costs.minimize_tcouses the same path; neither public signature changed.Test.
tiny_magnitudes_match_brute_force_and_respect_latency_boundsuses a workload with tiny costs and a 1e-7 s latency bound. Before this commitminimize_costreturned 3.6e-9 $/hour against a brute-force optimum of 8.0e-10, about 4.5× too expensive. It now matches brute force and respects the bound.🤖 Generated with Claude Code