Skip to content

ci(profile): stop regenerating committed schedules - #360

Draft
andrii-lz wants to merge 2 commits into
mainfrom
codex/profile-bench-skip-regeneration
Draft

ci(profile): stop regenerating committed schedules#360
andrii-lz wants to merge 2 commits into
mainfrom
codex/profile-bench-skip-regeneration

Conversation

@andrii-lz

@andrii-lz andrii-lz commented Aug 6, 2026

Copy link
Copy Markdown
Collaborator

Summary

Stop regenerating tracked schedule tables twice in every profile-benchmark matrix job. The benchmark now consumes the tables committed at each revision; the dedicated schedule-drift CI guard remains the source of truth.

Why

In the fresh slow CI job, the visible schedule-generation step took 1m43s. The merge-base build then ran the same generator again inside its worktree before compiling its profile binary.

Expected improvement

  • Removes two generator invocations per matrix cell.
  • Expected net saving: roughly 2–3 minutes per cell after accounting for dependencies that the following profile build must compile itself.
  • No benchmark/runtime behavior change; both revisions still use their own committed generated tables.

Validation

  • git diff --check

Stack

  1. This PR: remove redundant schedule regeneration.
  2. Add narrow profile feature groups.
  3. Switch each benchmark matrix group to its narrow feature.

@github-actions

github-actions Bot commented Aug 7, 2026

Copy link
Copy Markdown

Documentation blast radius (advisory)

These regions may need doc/spec/book updates based on changed paths.
This is not a merge gate. See docs/documentation.md.

Changed files in this PR: 1

ci-tooling

CI workflows and repo scripts

Code paths touched:

  • .github/workflows/profile-bench.yml

Consider updating:

  • docs/ci-test-timing.md
  • docs/documentation.md
  • specs/ci-test-timing.md

Per-PR checklist: spec Status / acceptance criteria; book owning page; AGENTS.md if contracts changed; archive spec after fold.

@github-actions github-actions Bot added the no-spec PR has no spec file label Aug 7, 2026
@github-actions

github-actions Bot commented Aug 7, 2026

Copy link
Copy Markdown

CI test timing

  • Report generated: 2026-08-07T00:21:53Z.
  • Source: 716c1cd on codex/profile-bench-skip-regeneration.
  • Workflow run: 31133541320.
  • Main baseline: c9ca8e9.

Run summary

Wall s Main wall s Main Δ Ratio Tests Skipped Failed Status
320.0 302.0 +6.0% 1.06x 1278 0 0 ok

Wall time spans 2 parallel nextest slice shards.

Slowest tests

Rank Duration s Test
1 12.0 akita-planner::schedule_params::tests::pruned_mixed_search_matches_unpruned_traversal_and_is_canonical
2 9.2 akita-pcs::akita_e2e::dense_d64_snap_regen_prove_verify_nv24
3 7.3 akita-planner::schedule_params::tests::mixed_nv36_benchmark_policy_selects_minimum_setup_schedule
4 6.5 akita-planner::schedule_params::tests::uniform_suffix_dp_matches_unpruned_exact_cutover_search
5 5.7 akita-sis-estimator::search_mode_parity::parallel_exhaustive_matches_serial_exhaustive_smoke
6 5.4 akita-prover::kernels::linear::tests::chunking::q128_many_blocks_digits_chunk_instead_of_unsafe_block_parallel
7 4.9 akita-pcs::single_poly_e2e::single_dense_nv18
8 4.4 akita-pcs::setup::d64_dense::large_setup_batch_passes
9 4.3 akita-sis-estimator::search_mode_parity::exhaustive_search_is_at_least_as_good_as_local_minimum_smoke
10 3.3 akita-pcs::setup::d128_dense::large_setup_nv_passes
11 3.1 akita-pcs::setup::d64_dense::same_size_passes
12 2.9 akita-prover::protocol::sumcheck::relation_range_image::tests::stage2_large_odd_dense_prefix_matches_padded_reference
13 2.8 akita-prover::protocol::sumcheck::relation_range_image::tests::stage2_large_odd_sparse_boolean_prefix_matches_padded_reference
14 2.8 akita-pcs::batched_aggregated_e2e::aggregated_mixed_dense_and_onehot_under_dense_cfg
15 2.7 akita-pcs::scheme::tests::single::fp128_degree_one_batched_proof_roundtrip_is_stable
16 2.7 akita-pcs::batched_aggregated_e2e::non_zk_aggregated_cases::aggregated_dense_nv17_batch4
17 2.7 akita-pcs::scheme::tests::onehot::multi_group_root_allows_precommitted_arity_above_final_group
18 2.6 akita-pcs::setup::d128_dense::large_setup_batch_passes
19 2.4 akita-pcs::scheme::tests::single::verify_rejects_malformed_v_dimension_without_panicking
20 2.3 akita-pcs::setup::d64_dense::large_setup_nv_passes

Regressions vs main

No per-test regressions above the threshold.

New slow tests

No new tests ≥30s vs main baseline.

@github-actions

github-actions Bot commented Aug 7, 2026

Copy link
Copy Markdown

Benchmark Report

  • Latest run: 694d3fa
  • Message: ci: retrigger checks after GitHub outage
  • Ref: codex/profile-bench-skip-regeneration
  • Workflow run: run 31133541339 attempt 1
  • Report generated: 2026-08-07T00:24:12Z.
  • Main baseline: 0f49d0a from the merge-base benchmarked on this runner.
  • Binary: target/release/examples/profile.
  • Memory: maximum resident set size from /usr/bin/time on the benchmark process.
Status Workload Setup contribution Setup and preparation Setup vector size Prepared NTT cache size Verifier NTT cache size Commit Prove Verify Peak process RSS Proof size
ok Fp32 - nv28Onehot256 - D=128 direct 0.043 s
+4.58% vs main
36.0 MiB
+0.00% vs main
144.0 MiB
+0.00% vs main
1.2 MiB
+0.00% vs main
0.106 s
+14.42% vs main
1.532 s
-0.22% vs main
33.7 ms
+1.23% vs main
460.8 MiB
+0.04% vs main
77,834 bytes
+0.00% vs main
ok Fp64 - nv28Onehot256 - D=128 direct 0.036 s
-3.03% vs main
40.0 MiB
+0.00% vs main
120.0 MiB
+0.00% vs main
0.9 MiB
+0.00% vs main
0.071 s
-2.40% vs main
1.087 s
+0.27% vs main
32.1 ms
+1.10% vs main
562.2 MiB
+0.54% vs main
83,596 bytes
+0.00% vs main
ok Fp128 - nv24Dense - D=64 direct 0.142 s
-0.49% vs main
215.0 MiB
+0.00% vs main
537.5 MiB
+0.00% vs main
1.4 MiB
+0.00% vs main
1.770 s
+3.29% vs main
1.207 s
-0.60% vs main
16.3 ms
+1.81% vs main
1572.4 MiB
+0.32% vs main
83,562 bytes
+0.00% vs main
ok Fp128 - nv26Onehot256 - D=64 - Tensor direct 0.658 s
+0.89% vs main
1024.0 MiB
+0.00% vs main
2560.0 MiB
+0.00% vs main
1.4 MiB
+0.00% vs main
0.117 s
+1.89% vs main
1.255 s
+0.12% vs main
33.6 ms
-1.51% vs main
4136.3 MiB
-0.24% vs main
85,080 bytes
+0.00% vs main
ok Fp128 - nv32Onehot256 - D=64 direct 0.206 s
-0.20% vs main
320.0 MiB
+0.00% vs main
800.0 MiB
+0.00% vs main
1.4 MiB
+0.00% vs main
1.236 s
+0.79% vs main
1.190 s
+0.65% vs main
23.3 ms
-0.34% vs main
1704.8 MiB
+0.24% vs main
84,972 bytes
+0.00% vs main
ok Fp128 - nv32Onehot256 - D_a=256D_b=128D_d=128 - MixedD256ToD64 direct 0.139 s
-1.51% vs main
128.0 MiB
+0.00% vs main
608.5 MiB
+0.00% vs main
1.4 MiB
+0.00% vs main
0.978 s
+0.09% vs main
1.316 s
-0.72% vs main
21.0 ms
+8.10% vs main
1410.0 MiB
-0.11% vs main
85,080 bytes
+0.00% vs main
ok Fp128 - nv30Onehot256 - Batched4 - D=64 direct 0.207 s
+0.12% vs main
320.0 MiB
+0.00% vs main
800.0 MiB
+0.00% vs main
1.4 MiB
+0.00% vs main
1.146 s
-1.24% vs main
1.194 s
-0.39% vs main
24.4 ms
-0.30% vs main
1701.3 MiB
+0.12% vs main
84,970 bytes
+0.00% vs main
ok Fp128 - nv32Onehot256 - Batched4 - D=64 - MultiGroup direct 0.427 s
-0.20% vs main
1032.0 MiB
+0.00% vs main
1075.0 MiB
+0.00% vs main
1.4 MiB
+0.00% vs main
3.215 s
+1.16% vs main
1.645 s
+1.19% vs main
29.8 ms
+5.48% vs main
3030.6 MiB
+0.09% vs main
85,093 bytes
+0.00% vs main
ok Fp128 - nv32Onehot256 - Batched4 - D=64 - MultiGroup recursive 3.903 s
-0.21% vs main
1032.0 MiB
+0.00% vs main
1075.0 MiB
+0.00% vs main
1.4 MiB
+0.00% vs main
3.169 s
+0.57% vs main
3.450 s
-0.34% vs main
26.1 ms
-4.94% vs main
3788.3 MiB
+0.31% vs main
89,873 bytes
+0.00% vs main
ok Fp128 - nv32Onehot256 - Batched4 - D=64 - MultiGroupW8R2 recursive 1.889 s
-0.04% vs main
430.0 MiB
+0.00% vs main
1075.0 MiB
+0.00% vs main
1.4 MiB
+0.00% vs main
3.522 s
+0.43% vs main
10.830 s
+0.19% vs main
37.9 ms
-5.50% vs main
3579.5 MiB
+0.14% vs main
95,904 bytes
+0.00% vs main
ok Fp128 - nv32Onehot256 - D=64 - MultiChunkW2R2 direct 0.207 s
-0.45% vs main
320.0 MiB
+0.00% vs main
800.0 MiB
+0.00% vs main
1.4 MiB
+0.00% vs main
1.254 s
+1.89% vs main
1.556 s
-0.06% vs main
24.7 ms
+1.18% vs main
1846.1 MiB
+0.06% vs main
85,331 bytes
+0.00% vs main
ok Fp128 - nv32Onehot256 - D=64 - MultiChunkW4R2 direct 0.223 s
-0.40% vs main
344.0 MiB
+0.00% vs main
860.0 MiB
+0.00% vs main
1.4 MiB
+0.00% vs main
0.428 s
-0.75% vs main
1.820 s
+0.07% vs main
28.2 ms
+2.66% vs main
2024.0 MiB
+0.21% vs main
85,946 bytes
+0.00% vs main
ok Fp128 - nv32Onehot256 - D=64 - MultiChunkW8R2 direct 0.223 s
+0.70% vs main
344.0 MiB
+0.00% vs main
860.0 MiB
+0.00% vs main
1.4 MiB
+0.00% vs main
0.428 s
+0.17% vs main
2.584 s
+0.54% vs main
30.3 ms
-1.80% vs main
2266.0 MiB
+0.09% vs main
86,215 bytes
+0.00% vs main

Negative deltas are improvements for time, memory, and proof size.

Terminal response component breakdown

Workload Folded response (z) Opening values (e) Inner-commitment values (t) Total terminal response
Fp32 - nv28Onehot256 - D=128 21,406 bytes 3,072 bytes 24,576 bytes 49,054 bytes
Fp64 - nv28Onehot256 - D=128 21,460 bytes 7,168 bytes 28,672 bytes 57,300 bytes
Fp128 - nv24Dense - D=64 21,854 bytes 7,168 bytes 28,672 bytes 57,694 bytes
Fp128 - nv26Onehot256 - D=64 - Tensor 21,848 bytes 7,168 bytes 28,672 bytes 57,688 bytes
Fp128 - nv32Onehot256 - D=64 21,852 bytes 7,168 bytes 28,672 bytes 57,692 bytes
Fp128 - nv32Onehot256 - D_a=256D_b=128D_d=128 - MixedD256ToD64 21,848 bytes 7,168 bytes 28,672 bytes 57,688 bytes
Fp128 - nv30Onehot256 - Batched4 - D=64 21,850 bytes 7,168 bytes 28,672 bytes 57,690 bytes
Fp128 - nv32Onehot256 - Batched4 - D=64 - MultiGroup 21,861 bytes 7,168 bytes 28,672 bytes 57,701 bytes
Fp128 - nv32Onehot256 - Batched4 - D=64 - MultiGroup 21,825 bytes 7,168 bytes 28,672 bytes 57,665 bytes
Fp128 - nv32Onehot256 - Batched4 - D=64 - MultiGroupW8R2 21,868 bytes 7,168 bytes 28,672 bytes 57,708 bytes
Fp128 - nv32Onehot256 - D=64 - MultiChunkW2R2 21,875 bytes 7,168 bytes 28,672 bytes 57,715 bytes
Fp128 - nv32Onehot256 - D=64 - MultiChunkW4R2 21,850 bytes 7,168 bytes 28,672 bytes 57,690 bytes
Fp128 - nv32Onehot256 - D=64 - MultiChunkW8R2 21,831 bytes 7,168 bytes 28,672 bytes 57,671 bytes

The z column includes its per-segment length prefixes and Golomb payload; e and t are raw field bytes. These three columns sum exactly to the serialized terminal response.

Detailed schedule and proof-size breakdowns by fold level are available in the uploaded report.md benchmark artifact.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

no-spec PR has no spec file

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant