ThreadBuildArena {
visited_gen : Vec<u32> // sized to max_doc_id + 1
gen_counter : u32 // bumped on each search; visited iff visited_gen[i] == gen_counter
to_visit : BinaryHeap<VisitorCandidate> // capacity ef_c, .clear() per call
found : BinaryHeap<Candidate> // capacity ef_c, .clear() per call
}
Round-3 perf push sub-issue (tracked under umbrella #535).
[M] Build-time visited-set / candidate-heap reuse with thread-local arenas
Where:
laurus/src/vector/index/hnsw/writer.rs(search_layer, currently around line 1243— the original
:842-938line reference had drifted; see investigation comment below)allocates
HashSet visited, twoBinaryHeaps, and aVecper call).Current behaviour: every node insertion calls
search_layeronceper layer (so
O(top_level)times per node); each call allocates afresh
HashSet<u64>+ twoBinaryHeaps. With rayon parallel buildthe allocator becomes a bottleneck.
Why it might be a bottleneck:
HashSet::new()is roughly freebut the first inserts trigger heap allocation; the heaps grow to >= ef
capacity. With M=16, ef_c=200, 10 M nodes, this is on the order of
hundreds of millions of allocator calls.
Reference precedent: hnswlib uses a thread-local "visitedList"
pool that resets a
u16generation counter per call instead ofclearing the set; FAISS does the same with
VisitedTable.Current data structure:
HashSet<u64>resized per call,BinaryHeap<Candidate>resized per call.Proposed layout (superseded — see investigation comment below for the adopted design, a
BitVec+ touched-list arena instead of thisVec<u32>generation-counter shape, which turnedout to retain hundreds of MB per thread for the process lifetime):
Reset is one increment + two
clear()s — no allocations after warm-up.Suggested direction: introduce
ThreadBuildArenaand pass itthrough
search_layerandprune_neighbors. Pair with the read-sidearena
HnswSearcheralready uses for the visited-set bitmap (perf(vector/hnsw): replace HashSet visited-set with bitmap #406).(Superseded — see investigation comment: the read side does not actually have a
thread_local!arena, and
prune_neighborsis out of scope for this PR; see task list below.)Risk / scope: medium; the arena needs sizing once
max_doc_idis known. Substantial CPU savings for large parallel builds.
Task list
SearchLayerArena(BitVec+touchedundo-list + twoBinaryHeaps) as athread_local!, wired throughsearch_layerpreserved, per-thread isolation) + mutation checks
cargo test -p laurus, existing HNSW recall/reachability/append tests unchanged)
ConcurrentHnswGraph::get_neighbors_view's per-candidateVec<u64>clone (found during investigation — likely a bigger allocation source than thisissue's original three containers, but out of scope here) — perf(vector/index): ConcurrentHnswGraph::get_neighbors_view clones a fresh Vec<u64> per call on the build-time hot path #1137
ID:
VI-14— see~/.claude/tasks/laurus/20260523_perf_round3_audit/task_list.mdfor the full Round-3 issue list.