dockyard_rl.models.dtensor.moe.sharding¶
Expert sharding placement policy for MoE blocks.
Distilled from torchtitan’s moe_sharding.py (pytorch/torchtitan, BSD-3) into
dockyard’s plain-DTensor idiom — WITHOUT torchtitan’s sharding framework
(ShardingConfig / NamedPlacement / MeshAxisName / decoder_sharding).
This module encodes only the placement DECISIONS: which DTensor placement each
MoE parameter category takes, ordered for a given device mesh’s dims.
The device-bound APPLICATION (distribute_tensor over the B.2 sparse mesh /
parallelize_module for the router + shared experts) lives with the
GroupedExperts module that B.4 vendors — it cannot be written or tested
against a phantom module, and is exercised only at bring-up (see
hardware-deferred-validation.md).
Param categories (torchtitan GroupedExperts / MoE module structure):
routed experts:
experts.{w1_EFD, w2_EDF, w3_EFD}— 3D(num_experts, *, *); EP on ->Shard(0)overep(sparse mesh); EP off -> dense TP shard on thetpaxis (colwise gate/up, rowwise down).router gate:
router.gate.weight—Replicateon every axis.shared experts: a normal dense FFN — use the standard Colwise/Rowwise
ParallelStylein the TP plan (no custom placement here; see notes).
Module Contents¶
Functions¶
Ordered DTensor placements for one routed-expert weight. |
|
Router gate weight is replicated on every mesh axis. |
|
Experts stored per EP rank ( |
Data¶
API¶
- dockyard_rl.models.dtensor.moe.sharding.ExpertRole¶
None
- dockyard_rl.models.dtensor.moe.sharding.routed_expert_placements(mesh_dim_names: Sequence[str], *, enable_ep: bool, role: dockyard_rl.models.dtensor.moe.sharding.ExpertRole) list[torch.distributed.tensor.placement_types.Placement]¶
Ordered DTensor placements for one routed-expert weight.
Args: mesh_dim_names: The target mesh’s dim names, in order. With EP this is the sparse mesh
("dp_replicate", "efsdp", "ep"); without EP it is the dense mesh("dp_replicate", "dp_shard", "cp", "tp"). enable_ep: Whether expert parallelism is active (ep > 1). role:gate_up(w1/w3) ordown(w2) — selects the TP shard dim when EP is off. Ignored when EP is on (experts alwaysShard(0)).Returns: A placement per mesh dim. EP on:
Shard(0)onep,Replicateelsewhere. EP off:Shard(1|2)ontp,Replicateelsewhere.
- dockyard_rl.models.dtensor.moe.sharding.router_gate_placements(mesh_dim_names: Sequence[str]) list[torch.distributed.tensor.placement_types.Placement]¶
Router gate weight is replicated on every mesh axis.
Every token needs the full set of routing logits, so the gate (tiny) is not sharded; it is replicated and computed locally.
- dockyard_rl.models.dtensor.moe.sharding.num_local_experts(num_experts: int, ep_size: int) int¶
Experts stored per EP rank (
num_expertsmust divideep_size).Mirrors
MoEParallelDims.num_local_experts; provided here so the sharding layer can size the local expert tensor without importing the mesh module.