Skip to content

[REQUEST] Legacy MoE gate: opt-in route_norm, with historical combine weights left as the default #8768

Description

@Aravind-11

Is your feature request related to a problem? Please describe.

The legacy MoE gate does not use one combine-weight rule. On current master (991ccf0f) the three functions in deepspeed/moe/sharded_moe.py disagree, and MoE / TopKGate have no switch for it.

  • top1gating returns the chosen expert's softmax probability. The sparse return is (gates * mask1).sum(dim=1) with no division (around line 321). The dense path's comment says "Normalize gate probabilities" and then only multiplies by the mask (around lines 329-331).
  • top2gating divides the two kept probabilities by their sum before both the sparse and dense returns (around lines 402-411).
  • topkgating divides the kept probabilities by their sum before both returns (around line 501). MoE(k=1) does not call this function. TopKGate.forward sends k == 1 to top1gating, k == 2 to top2gating, and only k > 2 to topkgating (around lines 593-609).

TokenChoiceTopKRouter in deepspeed/moe/ep_router.py already takes route_norm. The config key expert_parallel.route_norm configures that AutoEP router. The legacy MoE constructor does not read it.

The auxiliary loss is a separate disagreement, and it is computed before those divisions. Top-1 is sum(me * ce) * E on the single assignment. Top-2 is mean(me * ce) * E * E on the first-expert mask only, which is the same value as top-1 when the first expert matches. Top-k is mean(me * ce) * E * E / k on the full top-k mask.

I measured this on CPU with torch 2.5.1, before changing the code, at 991ccf0f. drop_tokens=False, use_rts=False, and for top-2 top2_2nd_expert_sampling=False. Logits:

[[2.0, 0.0, -1.0, 0.5],
 [0.2, 3.0, 0.1, -0.4],
 [1.0, 1.1, 0.2, 0.0],
 [-0.5, 0.3, 2.5, 0.1]]
path combine weights per-token mass l_aux
top1gating, sparse and dense 0.710100, 0.870166, 0.378175, 0.799164 those values, not 1 1.261781
topkgating(k=1) 1, 1, 1, 1 1 1.261781
top2gating the two kept probabilities, rescaled 1 1.261781
topkgating(k=2) same weights as top-2 on this input 1 1.144495

The k=1 load-balancing losses match. The k=2 losses do not (absolute difference 0.117286). On this input top-2 and top-k select the same experts, so the weight disagreement is top-1 versus top-k, and the loss disagreement is top-2 versus top-k. The gradient of the sum of combine weights with respect to the first token's logits is [0.205858, -0.068242, -0.025105, -0.112512] for top-1, and zeros for the renormalized gates, because those weights no longer depend on the softmax magnitude.

Unifying the weights, or unifying the auxiliary loss, changes numerics for existing users and checkpoints. Router parameter tensors keep their shapes. The MoE output and the expert gradients change, so a checkpoint trained with one rule does not reproduce its loss under the other.

I did not find an open issue for this. Search hits for route_norm in this repo are the AutoEP preset work (for example #8651), not the legacy gate.

Describe the solution you'd like

An opt-in route_norm on deepspeed.moe.layer.MoE, TopKGate, and the three gating functions. The default is None, which keeps today's rule inside each function:

  • top1gating: no renorm
  • top2gating and topkgating: renorm the kept weights so they sum to 1

True renorms for every k. False leaves the raw softmax probabilities for every k. The argument is trailing and defaulted, so existing positional callers stay valid, including the inference TopKGate(...) in deepspeed/ops/transformer/inference/moe_inference.py.

The auxiliary loss stays where it is, before the division, and this flag does not change it. A dropped token stays 0, because 0 / eps is 0.

A draft pull request implements only that. Tests require bitwise equality with the historical rule when the flag is omitted. I would rather not mark it ready for review until you say this shape is what you want.

Describe alternatives you've considered

  • Always renorm, including k=1. That matches topkgating(k=1) and the AutoEP presets that set route_norm=True, and it changes the MoE output of every existing top-1 checkpoint.
  • Never renorm. That changes every existing top-2 and top-k checkpoint.
  • A bool whose default is either the top-1 rule or the top-2 rule. One bool cannot mean both historical behaviors, which is why the default is None.
  • Also unifying l_aux in this change. That is a second numeric change. If you want it, it should be its own opt-in and stay off by default.

Additional context

Questions for @tohtana:

  1. Is an opt-in route_norm on the legacy gate welcome?
  2. Is None meaning "keep the historical per-gate rule" the default you want, with explicit True / False overriding every k?
  3. Is the name route_norm acceptable next to the existing AutoEP config key expert_parallel.route_norm, or do you want a different name on MoE?
  4. Should auxiliary-loss unification be a separate opt-in, and stay unchanged here?
  5. Two related legacy-gate mismatches are not in this change. used_token is an argument of top1gating only; TopKGate.forward accepts it for every k and then does not pass it for k != 1. On the no-drop path, top-1 pads capacity to the tensor-parallel size and then clamps it back to the token count, which can undo that padding, while top-2 and top-k keep the padded capacity. Should those be separate bugfixes?

I will change or drop the draft PR if this is not the shape you want.

Activity

  1. Aravind-11 commented on Oct 6, 2026

    @Aravind-11
    Author

    Draft implementation, with the historical combine weights left as the default: #8769.

    Please treat that PR as a proposal for the questions above. I will change the flag or close it if this is not the shape you want.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions