Skip to content

Average sparse gradients by world size regardless of gradient_predivide_factor - #8772

Draft
vineethsaivs wants to merge 1 commit into
deepspeedai:masterfrom
vineethsaivs:fix/sparse-allreduce-predivide-factor
Draft

vineethsaivs wants to merge 1 commit into
deepspeedai:masterfrom
vineethsaivs:fix/sparse-allreduce-predivide-factor

Conversation

@vineethsaivs

Copy link
Copy Markdown
Contributor

With gradient_predivide_factor set to anything but 1, sparse gradients (for example nn.Embedding(sparse=True)) come out that many times the data-parallel average. Dense gradients are fine.

allreduce_bucket divides by the factor before the all-reduce and multiplies by factor / world_size after it. sparse_allreduce copies only the second half: it multiplies by factor / world_size and then all-gathers, with no predivide.

Fix: an all-gather just concatenates, so there is nothing to protect from overflow. Scale sparse values by 1 / world_size, the same as the prescale branch already does.

Test: test_averaging_sparse_gradients.py is now parametrized over factor 1.0 and 2.0. At 2.0 it fails on master and passes with this change. The whole sparse_tensor directory passes (3 tests, 2 CPU ranks). yapf and flake8 are clean.

sparse_allreduce scaled values by gradient_predivide_factor / dp_world_size
under postscale, but unlike allreduce_bucket it never predivided before the
all-gather, so with a factor other than 1 sparse gradients came out that many
times the average.

Signed-off-by: Vineeth Sai <vineethsai4444@gmail.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant