Is your feature request related to a problem? Please describe.
The legacy MoE gate does not use one combine-weight rule. On current master (991ccf0f) the three functions in deepspeed/moe/sharded_moe.py disagree, and MoE / TopKGate have no switch for it.
top1gating returns the chosen expert's softmax probability. The sparse return is (gates * mask1).sum(dim=1) with no division (around line 321). The dense path's comment says "Normalize gate probabilities" and then only multiplies by the mask (around lines 329-331).
top2gating divides the two kept probabilities by their sum before both the sparse and dense returns (around lines 402-411).
topkgating divides the kept probabilities by their sum before both returns (around line 501). MoE(k=1) does not call this function. TopKGate.forward sends k == 1 to top1gating, k == 2 to top2gating, and only k > 2 to topkgating (around lines 593-609).
TokenChoiceTopKRouter in deepspeed/moe/ep_router.py already takes route_norm. The config key expert_parallel.route_norm configures that AutoEP router. The legacy MoE constructor does not read it.
The auxiliary loss is a separate disagreement, and it is computed before those divisions. Top-1 is sum(me * ce) * E on the single assignment. Top-2 is mean(me * ce) * E * E on the first-expert mask only, which is the same value as top-1 when the first expert matches. Top-k is mean(me * ce) * E * E / k on the full top-k mask.
I measured this on CPU with torch 2.5.1, before changing the code, at 991ccf0f. drop_tokens=False, use_rts=False, and for top-2 top2_2nd_expert_sampling=False. Logits:
[[2.0, 0.0, -1.0, 0.5],
[0.2, 3.0, 0.1, -0.4],
[1.0, 1.1, 0.2, 0.0],
[-0.5, 0.3, 2.5, 0.1]]
| path |
combine weights |
per-token mass |
l_aux |
top1gating, sparse and dense |
0.710100, 0.870166, 0.378175, 0.799164 |
those values, not 1 |
1.261781 |
topkgating(k=1) |
1, 1, 1, 1 |
1 |
1.261781 |
top2gating |
the two kept probabilities, rescaled |
1 |
1.261781 |
topkgating(k=2) |
same weights as top-2 on this input |
1 |
1.144495 |
The k=1 load-balancing losses match. The k=2 losses do not (absolute difference 0.117286). On this input top-2 and top-k select the same experts, so the weight disagreement is top-1 versus top-k, and the loss disagreement is top-2 versus top-k. The gradient of the sum of combine weights with respect to the first token's logits is [0.205858, -0.068242, -0.025105, -0.112512] for top-1, and zeros for the renormalized gates, because those weights no longer depend on the softmax magnitude.
Unifying the weights, or unifying the auxiliary loss, changes numerics for existing users and checkpoints. Router parameter tensors keep their shapes. The MoE output and the expert gradients change, so a checkpoint trained with one rule does not reproduce its loss under the other.
I did not find an open issue for this. Search hits for route_norm in this repo are the AutoEP preset work (for example #8651), not the legacy gate.
Describe the solution you'd like
An opt-in route_norm on deepspeed.moe.layer.MoE, TopKGate, and the three gating functions. The default is None, which keeps today's rule inside each function:
top1gating: no renorm
top2gating and topkgating: renorm the kept weights so they sum to 1
True renorms for every k. False leaves the raw softmax probabilities for every k. The argument is trailing and defaulted, so existing positional callers stay valid, including the inference TopKGate(...) in deepspeed/ops/transformer/inference/moe_inference.py.
The auxiliary loss stays where it is, before the division, and this flag does not change it. A dropped token stays 0, because 0 / eps is 0.
A draft pull request implements only that. Tests require bitwise equality with the historical rule when the flag is omitted. I would rather not mark it ready for review until you say this shape is what you want.
Describe alternatives you've considered
- Always renorm, including k=1. That matches
topkgating(k=1) and the AutoEP presets that set route_norm=True, and it changes the MoE output of every existing top-1 checkpoint.
- Never renorm. That changes every existing top-2 and top-k checkpoint.
- A bool whose default is either the top-1 rule or the top-2 rule. One bool cannot mean both historical behaviors, which is why the default is
None.
- Also unifying
l_aux in this change. That is a second numeric change. If you want it, it should be its own opt-in and stay off by default.
Additional context
Questions for @tohtana:
- Is an opt-in
route_norm on the legacy gate welcome?
- Is
None meaning "keep the historical per-gate rule" the default you want, with explicit True / False overriding every k?
- Is the name
route_norm acceptable next to the existing AutoEP config key expert_parallel.route_norm, or do you want a different name on MoE?
- Should auxiliary-loss unification be a separate opt-in, and stay unchanged here?
- Two related legacy-gate mismatches are not in this change.
used_token is an argument of top1gating only; TopKGate.forward accepts it for every k and then does not pass it for k != 1. On the no-drop path, top-1 pads capacity to the tensor-parallel size and then clamps it back to the token count, which can undo that padding, while top-2 and top-k keep the padded capacity. Should those be separate bugfixes?
I will change or drop the draft PR if this is not the shape you want.
Is your feature request related to a problem? Please describe.
The legacy MoE gate does not use one combine-weight rule. On current master (
991ccf0f) the three functions indeepspeed/moe/sharded_moe.pydisagree, andMoE/TopKGatehave no switch for it.top1gatingreturns the chosen expert's softmax probability. The sparse return is(gates * mask1).sum(dim=1)with no division (around line 321). The dense path's comment says "Normalize gate probabilities" and then only multiplies by the mask (around lines 329-331).top2gatingdivides the two kept probabilities by their sum before both the sparse and dense returns (around lines 402-411).topkgatingdivides the kept probabilities by their sum before both returns (around line 501).MoE(k=1)does not call this function.TopKGate.forwardsendsk == 1totop1gating,k == 2totop2gating, and onlyk > 2totopkgating(around lines 593-609).TokenChoiceTopKRouterindeepspeed/moe/ep_router.pyalready takesroute_norm. The config keyexpert_parallel.route_normconfigures that AutoEP router. The legacyMoEconstructor does not read it.The auxiliary loss is a separate disagreement, and it is computed before those divisions. Top-1 is
sum(me * ce) * Eon the single assignment. Top-2 ismean(me * ce) * E * Eon the first-expert mask only, which is the same value as top-1 when the first expert matches. Top-k ismean(me * ce) * E * E / kon the full top-k mask.I measured this on CPU with torch 2.5.1, before changing the code, at
991ccf0f.drop_tokens=False,use_rts=False, and for top-2top2_2nd_expert_sampling=False. Logits:top1gating, sparse and densetopkgating(k=1)top2gatingtopkgating(k=2)The k=1 load-balancing losses match. The k=2 losses do not (absolute difference 0.117286). On this input top-2 and top-k select the same experts, so the weight disagreement is top-1 versus top-k, and the loss disagreement is top-2 versus top-k. The gradient of the sum of combine weights with respect to the first token's logits is
[0.205858, -0.068242, -0.025105, -0.112512]for top-1, and zeros for the renormalized gates, because those weights no longer depend on the softmax magnitude.Unifying the weights, or unifying the auxiliary loss, changes numerics for existing users and checkpoints. Router parameter tensors keep their shapes. The MoE output and the expert gradients change, so a checkpoint trained with one rule does not reproduce its loss under the other.
I did not find an open issue for this. Search hits for
route_normin this repo are the AutoEP preset work (for example #8651), not the legacy gate.Describe the solution you'd like
An opt-in
route_normondeepspeed.moe.layer.MoE,TopKGate, and the three gating functions. The default isNone, which keeps today's rule inside each function:top1gating: no renormtop2gatingandtopkgating: renorm the kept weights so they sum to 1Truerenorms for every k.Falseleaves the raw softmax probabilities for every k. The argument is trailing and defaulted, so existing positional callers stay valid, including the inferenceTopKGate(...)indeepspeed/ops/transformer/inference/moe_inference.py.The auxiliary loss stays where it is, before the division, and this flag does not change it. A dropped token stays 0, because
0 / epsis 0.A draft pull request implements only that. Tests require bitwise equality with the historical rule when the flag is omitted. I would rather not mark it ready for review until you say this shape is what you want.
Describe alternatives you've considered
topkgating(k=1)and the AutoEP presets that setroute_norm=True, and it changes the MoE output of every existing top-1 checkpoint.None.l_auxin this change. That is a second numeric change. If you want it, it should be its own opt-in and stay off by default.Additional context
Questions for @tohtana:
route_normon the legacy gate welcome?Nonemeaning "keep the historical per-gate rule" the default you want, with explicitTrue/Falseoverriding every k?route_normacceptable next to the existing AutoEP config keyexpert_parallel.route_norm, or do you want a different name onMoE?used_tokenis an argument oftop1gatingonly;TopKGate.forwardaccepts it for every k and then does not pass it fork != 1. On the no-drop path, top-1 pads capacity to the tensor-parallel size and then clamps it back to the token count, which can undo that padding, while top-2 and top-k keep the padded capacity. Should those be separate bugfixes?I will change or drop the draft PR if this is not the shape you want.