Skip to content

rqe-optimizer: price plans per EC2 machine family - #137

Merged
zzylol merged 3 commits into
mainfrom
feat/rqe-optimizer-ec2-cost
Oct 5, 2026
Merged

zzylol merged 3 commits into
mainfrom
feat/rqe-optimizer-ec2-cost

Conversation

@zzylol

@zzylol zzylol commented Oct 4, 2026 •

Copy link
Copy Markdown
Collaborator

Before this PR

rqe-optimizer minimized total CPU only. Memory was tracked as peak per-query memory, which does not count the sketch state a deployment keeps over a long lookback. Nothing priced CPU against memory.

After this PR

This implements §4 of the AutoSketch-vs-planner evaluation plan (ProjectASAP/ASAPQuery#777, PR 1 in its §8).

  • Retained memory. Objectives::retained_memory_bytes sums, over active deployments, card × mem_per_instance × (x + max S) / y. That is the x/y open instances plus the closed ones still inside the longest lookback the deployment serves. Deployment::retained_instance_count gives the per-RQE count.

  • EC2 pricing. milp::minimize_cost(rqes, deployments, label_sets, bounds, family) minimizes price_f · n_f subject to n_f ≥ CPU / vCPU_f and n_f ≥ retained GiB / GiB_f. So it pays for whichever of CPU or memory runs out first. Retained memory is linearized per deployment with R_D ≥ retained_{i,D} · z_{i,D}. The existing latency and peak-memory bounds still apply. minimize_tco keeps its signature and behavior, and both share one model builder.

  • Prices. rqe-optimizer/data/ec2-pricing-2026-10-04.json is written by scripts/fetch_ec2_pricing.py from the public file behind aws.amazon.com/ec2/pricing/on-demand (us-east-1, Linux, on-demand, published 2026-09-25):

    Family Instance vCPU GiB $/hour
    compute_optimized c7i.xlarge 4 8 0.1785
    general_purpose m7i.xlarge 4 16 0.2016
    memory_optimized r7i.xlarge 4 32 0.2646

    Disk is not priced. MachineFamily::from_pricing_json, instances and usd_per_hour read the snapshot and score any mapping.

  • Dominance pruning now also requires the replacement to retain no more memory for each RQE the removed candidate serves. Without this, it could drop a candidate that is cheaper under the new objective. A regression test covers a case the old rule pruned.

  • Example. small_problem --milp --machine-family NAME solves the cost model.

  • Docs. docs/rqe_sketch_deployment_v1.md and docs/rqe_optimizer_TODO.md describe the new objective.

Coordination: Milind is porting milp.rs into ASAPQuery's planner. This adds a function rather than changing minimize_tco, so the port can take it separately.

Validation

  • cargo test -p rqe-optimizer: 18 passed. New tests:

    • retained-memory scoring, including a shared deployment sized for its longest lookback;
    • binding-resource sizing;
    • parsing the committed snapshot;
    • minimize_cost matching an exhaustive search on a tiny workload for two families;
    • the chosen plan switching with the family's CPU/memory ratio;
    • the dominance regression.
  • cargo clippy -p rqe-optimizer --all-targets -- -D warnings and cargo fmt --check are clean.

  • Ran small_problem --milp end to end on a freshly exported cost table (scripts/export_rqe_optimizer_costs.sh, 18 rows):

    Objective $/hour Instances Retained memory Total CPU (cpu-sec/sec)
    minimum TCO CPU — — — 8.39e-3
    compute_optimized 0.0234 0.131 1128 MB 1.18e-2
    general_purpose 0.0132 0.066 1128 MB 1.18e-2
    memory_optimized 0.0087 0.033 1128 MB 1.18e-2

    On this toy workload, retained state binds in every family. Memory-optimized is cheapest, and the cost plan uses about 40% more CPU than the CPU-only plan (1.18e-2 vs. 8.39e-3).

Update: solver scaling (cdad511)

Problem. The solver works to fixed tolerances that are larger than this model's numbers:

Tolerance (HiGHS default) Value Our numbers
Absolute MIP gap: stops once within this of the optimum 1e-6 Whole-plan costs of ~1e-6 $/hour
Feasibility tolerance: accepts a constraint violated by up to this 1e-7 Query latencies of ~1e-6 s

So on small workloads minimize_cost could stop at a plan several times more expensive than the best one, and could pick a deployment whose latency was slightly over its bound. Neither showed up as an error.

Fix, inside solve, so callers don't change:

  1. Rescale the objective to about 1. reference is what the plan would cost with no sharing, each RQE on its cheapest eligible deployment, in the units being minimized. The CPU, memory and instance-count terms are divided by it, so the optimum is close to 1 and the 1e-6 gap is negligible. The returned plan is still scored on real costs.
  2. Check per-RQE bounds before solving. Each RQE uses exactly one deployment, so a latency or memory bound only rules out that RQE's choices that exceed it. These are compared directly in f64 and never become solver constraints, so the 1e-7 tolerance can't let a violation through.

minimize_tco uses the same path; neither public signature changed.

Test. tiny_magnitudes_match_brute_force_and_respect_latency_bounds uses a workload with tiny costs and a 1e-7 s latency bound. Before this commit minimize_cost returned 3.6e-9 $/hour against a brute-force optimum of 8.0e-10, about 4.5× too expensive. It now matches brute force and respects the bound.

🤖 Generated with Claude Code

zzylol and others added 2 commits October 4, 2026 22:36
Score retained memory per active deployment, (x + max S) / y instances,
and add milp::minimize_cost, which minimizes the hourly price on one EC2
machine family in fractional instances (whichever of CPU or memory binds).
Prices are a committed on-demand snapshot from scripts/fetch_ec2_pricing.py.
Dominance pruning also compares retained memory so it cannot drop a
candidate that is cheaper under the new objective.

Part of ProjectASAP/ASAPQuery#777.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
HiGHS's absolute gap (1e-6) and feasibility tolerance (1e-7) exceed
real plan costs (~1e-6 $/hour) and query latencies (µs). Rows are now
scaled by the plan's demand without sharing, and per-RQE memory and
latency bounds become f64 exclusions of the choices over them.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
zzylol added a commit that referenced this pull request Oct 5, 2026
Replace the alpha × fastest-latency limit with one absolute SLA per
RQE, swept over 0.01 ms to 1 s plus none. RQEs no method can serve
within an SLA are excluded from all methods and listed. Drop the
runner's demand scaling now that minimize_cost scales internally
(#137). Rerun W0, W2 and W1 sequentially on an idle CloudLab node and
regenerate figures, summary and README.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@zzylol

zzylol commented Oct 5, 2026

Copy link
Copy Markdown
Collaborator Author

@milindsrivastava1997 I reviewed it and the changes include MILP files. I plan to merge this to sketch-bench?

@zzylol
zzylol merged commit e89b304 into main Oct 5, 2026
2 checks passed
@zzylol
zzylol deleted the feat/rqe-optimizer-ec2-cost branch October 5, 2026 14:01
zzylol added a commit that referenced this pull request Oct 5, 2026
Runner (examples/autosketch_vs_asap.rs) for the traces workloads of
ProjectASAP/ASAPQuery#777: ASAP (joint MILP with one absolute latency SLA
per sweep point), PerQuery-CostAware and AutoSketch-Adapted, scored by
objectives::score and priced per EC2 family, with sanity checks.

Accuracy now depends on the query's merge count m = S/x: a cost entry may
carry "{metric}@m{b}" keys, eligibility reads the smallest measured b >= m,
and a merge larger than every measured b is ineligible.

Results for Alibaba 2022, Google 2011 and BOOM with figures and summary.
The earlier example and scaling workloads are dropped (#777 §6).

Rebased onto main, which now has #137, #135 and #136.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant