DIET: Deletion-response Expert Trimming for Video Diffusion Transformers
Mixture-of-experts (MoE) video models reduce computation by activating only a small number of experts for each token, but they still need to store the full expert bank. This paper introduces DIET, a training-free method for physically removing redundant experts from a video diffusion transformer. Instead of ranking experts using static usage statistics, DIET measures how the layer changes when each expert is deleted and accounts for the re-routing that follows.
What the paper does
In an MoE layer, a router selects a subset of experts for each token. Removing one expert changes more than the model’s parameter count: tokens previously sent to that expert may be redirected, and the remaining experts’ gate weights may be renormalized. The authors argue that pruning methods based only on routing frequency or activation magnitude miss this post-deletion behavior.
DIET targets the storage problem directly by deleting entire experts. This can produce physical checkpoint reductions proportional to the number of experts removed, unlike approaches that merely reduce active computation. The method is designed for large video diffusion transformers, where the complete expert bank may otherwise require multiple GPUs simply to remain resident in memory.
The approach is evaluated on LingBot-Video 30B-A3B, which has 48 layers and 128 experts per layer. Each token activates eight experts through grouped routing, giving 6,144 experts in total. The main experiment retains half of them: 3,072 experts.
How deletion responses are measured
Naively testing every possible expert deletion would require thousands of additional model evaluations. DIET instead performs one instrumented calibration pass that evaluates all experts and records router logits, gate weights, routing selections, and expert outputs. The paper reports that this pass costs roughly 16 times the expert computation of a normal pass and produces about 375 GB of cached half-precision state, but it is performed only once.
Afterward, deleting an expert can be simulated using tensor arithmetic on the cached values. The method recomputes grouped top-k routing and gate redistribution at the captured states without running another model forward pass. For each expert, DIET records the change between the original layer output and the output under that expert’s deletion.
These changes are pooled across tokens and calibration cases into a deletion-response signature. The signature is intended to describe an expert’s functional role and its interaction with the other experts in the layer. DIET focuses on the direction of this response rather than its raw size, because the authors find that responses from multiple deletions do not combine linearly.
Selecting experts and distributing the budget
Within each layer, DIET uses an objective called Overall Diversity Loss, or ODL. For every deleted expert, the method finds the most similar retained expert in signature space and measures their cosine distance. The goal is therefore to retain a set that covers as many distinct response directions as possible. An expert whose deletion response is already represented by a surviving expert is considered less costly to remove.
The resulting selection problem is combinatorial. DIET uses greedy initialization with randomized restarts, single-swap refinement, and simulated annealing. It repeats this process for different retention levels, creating layer-specific trade-offs between the number of retained experts and the ODL objective.
The method also allows different layers to keep different numbers of experts. A regularized linear surrogate is fitted from end-to-end evaluations, estimating how layer-wise expert counts relate to benchmark dimensions. A constrained search then proposes allocations under a global budget, while enforcing feasibility and quality floors. The authors report that this inter-layer allocation improves results over using a uniform retention ratio, with a stated 1.2-point gain in their analysis.
This layer-aware storage allocation differs from methods such as KV-Kaizen’s context-adaptive cache compression choices, which address memory growth by changing how decoding caches are compressed rather than removing model parameters. DIET’s target is the persistent expert checkpoint itself.
Reported results and limitations
At 50% expert pruning, the LingBot-Video checkpoint shrinks from 57 GB to 30 GB. The authors say this enables deployment on a single 48 GB GPU without fine-tuning. On a fixed 284-case VBench evaluation using matched random seeds, the official VBench Total rises from 0.7941 for the unpruned model to 0.8115 after pruning. DIET also reportedly outperforms the tested pruning baselines adapted from large language models at the other retention budgets examined.
The paper does not treat the pruned model as a point-wise reconstruction of the original. Because deletion changes routing, generated videos may contain different motions even when they remain plausible. For that reason, the evaluation emphasizes perceptual and semantic video quality rather than metrics such as PSNR.
The reported evidence is centered on one large open-source video DiT-MoE model and a specific calibration and evaluation setup. The method also requires a substantial one-time calibration pass and very large cached tensors. Its central claim is not that every expert can be removed safely, but that measuring counterfactual deletion responses gives a more informative basis for deciding which experts are functionally redundant and how pruning should be distributed across layers.
Comments (0)
to join the discussion
No comments yet
Be the first to share your thoughts!