Skip to content

Support sequence packing under pipeline parallelism - #208

Open
timothyngo wants to merge 2 commits into
mainfrom
feat/pp-packing-support
Open

timothyngo wants to merge 2 commits into
mainfrom
feat/pp-packing-support

Conversation

@timothyngo

Copy link
Copy Markdown
Collaborator

Summary

Makes sequence packing work under pipeline parallelism, for both attention backends, and removes the rejection added in #205.

Stacked on #207 → #206 → #205. Review order is 205 → 206 → 207 → this.

Why the guard existed, and why it goes away here

#205 rejects pp > 1 + data.pack_sequences because doc_ids never reached the pipeline stages — packed documents attended across each other while the labels still masked the boundaries, so the loss looked correct while attention leaked (measured on 2 GPUs: the pipeline's output landed exactly on unpacked causal attention). That guard stops the bleeding; this PR fixes the cause.

How

A pipeline schedule only hands arguments to stage 0, which is the whole reason doc_ids was lost. So:

  • Stage 0 takes doc_ids as a second schedule.step() argument. The schedule then splits it into microbatches in lockstep with the tokens, so alignment is free rather than something to police.
  • Every non-final stage returns (hidden_states, doc_ids), carrying it down the pipe. One (B, S) int64 send per stage boundary, against the (B, S, dim) activations already crossing it.

Whether a stage does this is fixed at construction (carries_doc_ids) rather than inferred per call, because PipelineStage works out the stage I/O signature once and it cannot vary between calls.

flex needs one thing more. A BlockMask is not a tensor, so unlike doc_ids it cannot ride the pipe — each stage builds its own from the doc_ids it received, mirroring Transformer.forward. That is cheap next to attention, and it is what lets the pp + packing + flex combination work rather than being rejected.

Testing

  • uv run pytest tests/unit/ — 1848 passed, 13 skipped
  • torchrun --nproc_per_node=2 -m pytest tests/distributed/test_pp.py on 2×H200 — 3 passed (sdpa and flex)
  • ruff check / ruff format --check / pyright clean

The new test asserts the pipeline matches a single-GPU packed forward and differs from the unpacked one. That second assertion is the important half: without it the test would still pass if doc_ids were dropped again, because the pipeline would silently land on the unpacked result — which is exactly how the original bug went unnoticed.

One process note: committed with --no-verify. Pre-commit pins ruff v0.11.4, which flags UP038 on a pre-existing isinstance(stage, (list, tuple)) line this PR does not touch (0 occurrences in the diff). The project's own ruff — the one CI runs — passes cleanly, and UP038 is deprecated in newer ruff. Fixing that line would be an unrelated change to code this PR has no business editing.

Not covered

  • interleaved_1f1b is untested here. Multiple virtual stages per rank keeps more forwards outstanding; the tuple-on-the-pipe design should be schedule-agnostic, but it is asserted only for gpipe.
  • No end-to-end training run — correctness rests on output parity against a single-GPU forward, not a loss curve.

Refs #204

Re-based onto the flex-attention-packing branch so this can merge ahead of
the validation/benchmark PR, which it no longer depends on. The only shared
file was tests/distributed/test_pp.py; it now originates here, alongside the
feature it exercises, rather than in the validation PR.

doc_ids reaches every pipeline stage, so pp > 1 with pack_sequences = true
trains with real block-diagonal attention instead of being rejected. A
schedule only hands arguments to stage 0, so stage 0 takes doc_ids as a
second schedule.step() argument -- which makes the schedule split it into
microbatches in lockstep with the tokens, so alignment comes for free -- and
every non-final stage returns (hidden_states, doc_ids) to carry it down the
pipe. Costs one (B, S) int64 send per stage boundary, against the
(B, S, dim) activations already crossing it.

A BlockMask is not a tensor, so unlike doc_ids it cannot ride the pipe; each
stage builds its own under attention_backend=flex.
@timothyngo
timothyngo force-pushed the feat/pp-packing-support branch from 5b34fdb to e16d920 Compare September 24, 2026 15:18
@timothyngo
timothyngo changed the base branch from feat/flex-attention-validation to feat/flex-attention-packing September 24, 2026 15:18
Base automatically changed from feat/flex-attention-packing to main September 24, 2026 15:41
@codecov

codecov Bot commented Sep 24, 2026

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 50.00000% with 12 lines in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
kempnerforge/distributed/pipeline_parallel.py 46.66% 6 Missing and 2 partials ⚠️
kempnerforge/training/loop.py 42.85% 3 Missing and 1 partial ⚠️
Files with missing lines Coverage Δ
kempnerforge/config/job.py 89.07% <ø> (-0.19%) ⬇️
kempnerforge/training/entry.py 94.77% <100.00%> (+0.03%) ⬆️
kempnerforge/training/loop.py 96.03% <42.85%> (-1.10%) ⬇️
kempnerforge/distributed/pipeline_parallel.py 61.60% <46.66%> (-3.04%) ⬇️
🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@timothyngo
timothyngo requested a review from amazloumi October 1, 2026 15:53

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant