Skip to content
Open
Show file tree
Hide file tree
Changes from 3 commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
20 changes: 20 additions & 0 deletions CMakeLists.txt
Original file line number Diff line number Diff line change
Expand Up @@ -57,6 +57,7 @@ add_library(
src/cuda_runtime/timing_binding.cpp
src/cuda_runtime/vmm.cpp
src/hybrid/calibrator.cpp
src/prefetch/prefetch_model.cpp
src/profile/profile.cpp
src/host_service/request_dispatcher.cpp
src/host_service/backing_store.cpp
Expand Down Expand Up @@ -108,6 +109,19 @@ target_link_libraries(hbfsim_profile_tests PRIVATE hbfsim_core)
add_test(NAME profile COMMAND hbfsim_profile_tests)
set_tests_properties(profile PROPERTIES WORKING_DIRECTORY "${CMAKE_CURRENT_SOURCE_DIR}")

add_executable(hbf_prefetch_bench benchmarks/prefetch/hbf_prefetch_bench.cpp)
target_link_libraries(hbf_prefetch_bench PRIVATE hbfsim_core)

add_executable(
hbf_capacity_readahead_bench
benchmarks/prefetch/hbf_capacity_readahead_bench.cpp
)
target_link_libraries(hbf_capacity_readahead_bench PRIVATE hbfsim_core)

add_executable(prefetch_model_test tests/cpu/prefetch_model_test.cpp)
target_link_libraries(prefetch_model_test PRIVATE hbfsim_core)
add_test(NAME prefetch_model COMMAND prefetch_model_test)

add_executable(hbfsim_calibrator_tests tests/cpu/calibrator_test.cpp)
target_link_libraries(hbfsim_calibrator_tests PRIVATE hbfsim_core)
add_test(NAME calibrator COMMAND hbfsim_calibrator_tests)
Expand Down Expand Up @@ -167,6 +181,12 @@ add_test(
NAME capacity_page_service COMMAND hbfsim_capacity_page_service_tests
)

add_executable(
hbfsim_capacity_readahead_tests tests/cpu/capacity_readahead_test.cpp
)
target_link_libraries(hbfsim_capacity_readahead_tests PRIVATE hbfsim_core)
add_test(NAME capacity_readahead COMMAND hbfsim_capacity_readahead_tests)

add_executable(
hbfsim_capacity_handoff_tests tests/cpu/capacity_handoff_test.cpp
)
Expand Down
111 changes: 111 additions & 0 deletions benchmarks/prefetch/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,111 @@
# Prefetch experiment: what to run and where everything is

The reasoning behind the experiment, and what its numbers may and may not be
used to claim, is in `docs/46-预取实验设计.md`. This file is only the map and
the commands.

## Files

| Path | What it is |
|---|---|
| `include/hbfsim/prefetch_model.hpp` | The model's interface, with each knob's meaning |
| `src/prefetch/prefetch_model.cpp` | The model: a deterministic discrete-event simulation |
| `tests/cpu/prefetch_model_test.cpp` | 12 property assertions, written before the implementation |
| `benchmarks/prefetch/hbf_prefetch_bench.cpp` | Sweeps one access stream through the model and prints JSON |
| `src/host_service/capacity_page_service.cpp` | The real system-side readahead, in capacity mode |
| `tests/cpu/capacity_readahead_test.cpp` | Its contract: off by default, never forces a writeback, never drains on the demand's own path |
| `benchmarks/prefetch/hbf_capacity_readahead_bench.cpp` | Drives the real readahead and reports media reads avoided |
| `scripts/run_prefetch_accuracy_sweep.py` | Runs all three streams and writes the artifact and a CSV |
| `docs/proofs/artifacts/prefetch-accuracy-sweep.json` | The swept cells |
| `docs/proofs/artifacts/prefetch-accuracy-sweep.csv` | The same cells, flat, for plotting |

## Build

The model and its test are CPU-only. Neither needs a GPU or `nvcc`.

```
cmake -S . -B build -DHBFSIM_ENABLE_CUDA=OFF -DHBFSIM_ENABLE_MQSIM=ON \
-DCMAKE_BUILD_TYPE=Release
cmake --build build -j"$(nproc)" --target prefetch_model_test hbf_prefetch_bench
```

## Run the tests

```
./build/prefetch_model_test # exit code 0, prints nothing on success
```

On failure it prints the line number and the assertion that failed.

## Reproduce the sweep

```
python3 scripts/run_prefetch_accuracy_sweep.py --build-dir build
```

Writes `docs/proofs/artifacts/prefetch-accuracy-sweep.json` and the matching
`.csv`. Both carry a `disclaimer` field reading
`modeled, not measured on any device or GPU`.

## One stream at a time

```
./build/hbf_prefetch_bench --stream sequential --accesses 20000
./build/hbf_prefetch_bench --stream random --accesses 20000
./build/hbf_prefetch_bench --stream moe --accesses 20000 \
--pages-per-expert 2304
```

`--pages-per-expert 2304` is one Qwen3-30B-A3B expert: 3 x 2048 x 768
parameters in bf16 is 9,437,184 bytes, which is 2304 pages of 4 KiB.

Other options: `--compute-ns` (accelerator time between two accesses, the
interval a prefetch hides behind), `--lead` (how many accesses ahead a prefetch
is issued), `--buffer-pages`, `--max-in-flight`, `--seed`.

## The real readahead, as opposed to the model

`hbf_capacity_readahead_bench` exercises `CapacityPageService` directly. It
reports media reads avoided, not wall-clock time, and touches no GPU.

Does the implementation reproduce what the model predicts for a next-page
policy, which is (P-1)/P for P pages per expert:

```
for p in 4 8 16; do
./build/hbf_capacity_readahead_bench --stream moe --pages-per-expert $p \
--readahead $p --frames 256 --drain-per-demand 0
done
```

What the worker's drain rate is worth. `--drain-per-demand` is how many queued
pages the worker gets through between two demands, and 0 means it keeps up
completely:

```
for d in 0 4 2 1; do
./build/hbf_capacity_readahead_bench --stream moe --pages-per-expert 8 \
--readahead 8 --frames 256 --drain-per-demand $d
done
```

Draining one page per demand does not merely reduce the benefit. It turns an
86.79% reduction in media reads into a 14.17% increase, because queued pages go
stale before they are used while still costing bandwidth and frames.

## Reproducing the two results the paper leans on

The naive next-page policy's accuracy on a Mixture-of-Experts stream is
(P-1)/P, where P is how many pages one expert occupies:

```
for p in 1 2 4 8 16; do
./build/hbf_prefetch_bench --stream moe --accesses 20000 --pages-per-expert $p
done
```

One token of Qwen3-30B-A3B, which is 884,736 page accesses:

```
./build/hbf_prefetch_bench --stream moe --accesses 20000 --pages-per-expert 2304
```
Loading