Beyond Retraining-Free MoE Compression:
A Cost-Normalized Study of Post-Compression Adjustment

EMNLP 2026 · Main Conference

1 Department of Electrical and Computer Engineering, Seoul National University
2 Interdisciplinary Program in Artificial Intelligence, Seoul National University

arXiv link will be updated after upload.

Post-Compression Adjustment (PCA)

Retraining-free compression is a strong start.
A small post-compression adjustment is an important recovery stage.

A compressed checkpoint is an initialization for further recovery. Design the compression and adjustment stages together.

Abstract

Retraining-free MoE compression reduces deployment memory by pruning or merging experts, but often treats the compressed checkpoint as the final artifact. We argue that this view is incomplete: compressed MoE checkpoints are better understood as compressed initializations that benefit from a small post-compression adjustment stage.

We study this stage under a small general-domain data budget, comparing causal language modeling fine-tuning and teacher-based knowledge distillation across different trainable parameter scopes. By considering performance recovery alongside measured GPU costs, we examine how the adjustment objective and scope affect the compressed model. The study motivates a practical two-stage view: retraining-free compression followed by a short recovery stage, with both stages considered when designing the final pipeline.

Post-Compression Adjustment

Expert pruning and merging change the available expert paths and, for merging, the functions implemented by the experts. The resulting mismatch can involve both router reweighting and expert/path errors. These perturbations can also propagate through subsequent attention, normalization, and shared components.

PCA applies short gradient updates to an already-compressed checkpoint using a small amount of general-domain text. Our study compares adjustment choices on the same compressed checkpoints under matched data budgets and measured GPU costs.

Which training objective?

Causal LM fine-tuning learns from observed next tokens. Token-level knowledge distillation matches the output distribution of the original, uncompressed teacher.

Which trainable parameters?

We compare router-only updates, router + selected experts, router + all experts, and full-parameter updates that also allow shared model components to change.

Matched adjustment protocol. The main experiments use 3,000 C4 examples, one epoch, and two NVIDIA H200 GPUs. Expert-selection and online-teacher passes are included in the measured adjustment cost when applicable. Experimental setup in the paper ↗

Experimental Results

Effect of Trainable Parameter Scope

How much of the compressed model should be allowed to change? We expand the trainable scope from router-only, to router + selected experts (top-8, top-16, or top-50 per layer), to router + all experts, and finally to all model parameters. Select a scope to see which components are updated and compare the mean recovery from causal LM fine-tuning and token-level KD.

Explore the trainable scope

Choose which parts of the compressed model can change during adjustment.

Trainable Frozen
Hidden states
Attention & shared transformationsFrozen
Compressed MoE layer
RouterTrainable

All retained experts frozen

Downstream & shared parametersFrozen

Conceptual view. Expert icons do not represent an exact count or proportion. Top-k experts are selected per layer before adjustment.

Router only

Only router parameters are updated. Retained experts and shared model components remain frozen.

Overall strategy recovery

Original paper Figure 1. Overall mean recovery gains of post-compression adjustment strategies; Full FT achieves the largest gain.
Figure 1. Overall recovery gain over the unadjusted compressed baseline. Values average the studied backbone–compressor–retention settings and downstream benchmarks. Original figure ↗

Expanding the trainable scope

Original paper Figure 4. Recovery gains for FT and KD increase from router-only through selected experts and all experts to full-parameter adjustment.
Figure 4. Recovery gain as the trainable scope expands. LM fine-tuning and KD both improve with broader parameter updates. Original figure ↗

Broader updates yield greater aggregate recovery. Figure 1 compares the adjustment strategies, while Figure 4 traces the effect of expanding the trainable scope. Full FT achieves the largest mean gain of +3.49 percentage points, followed by router + all experts FT (+2.31) and Full KD (+2.19). Under this small-budget protocol, causal LM fine-tuning outperforms token-level KD at each matched scope.

Router-only adjustment can reweight available experts, but cannot directly change their functions. Expert updates provide more ways to correct compression-induced errors. The further gain from router + all experts FT to Full FT suggests that allowing shared parameters to change also matters, consistent with errors propagating beyond the MoE modules. These are aggregate results under the tested adjustment budget; individual settings can differ.

Cost–Recovery Trade-off

Full FT provides the largest aggregate recovery while using less measured GPU time, memory occupancy, and energy than Full KD. Restricting adjustment to selected experts also incurs a separate expert-selection pass, so fewer trainable parameters do not necessarily imply lower adjustment cost.

Original paper Figure 2, top panel. Mean recovery gain versus effective GPU-hours.
Figure 2. Effective GPU-hours (top panel). Mean recovery gain versus measured GPU cost; upper-left indicates higher recovery at lower cost. Expert selection and online teacher passes are included when applicable. Original figure (PDF) ↗

Effective GPU-hours are utilization-weighted time; GiB-hours measure time-integrated memory occupancy. Compression, model loading, and benchmark evaluation are excluded. Full FT leads in aggregate under the measured protocol; it is not the cheapest strategy or the strongest in every individual setting.

Benchmark-Family Recovery

Recovery is not uniform across benchmark families. Full FT recovers most strongly on chain-of-thought reasoning and multiple-choice QA, followed by general reasoning and QA. Mathematics changes only modestly, and coding decreases slightly.

Original paper Figure 3. Full FT performance by family, with original and compressed score references. Gains are largest for CoT and MCQA; coding is slightly lower.
Figure 3. Benchmark-family recovery of Full FT. Bars show adjusted scores; dashed lines show the original model and the compressed baseline. Original figure (PDF) ↗

The result suggests that a short general-domain LM adjustment is especially useful for recovering general generation and reasoning behavior. Specialized domains may require task-specific data or longer adaptation.

Qualitative Examples: Before and After Adjustment

The family-level results show that Full FT helps chain-of-thought reasoning and multiple-choice QA most, followed by general reasoning and QA. The paired responses below add a generation-level view: can the compressed model sustain a natural-language answer without collapsing into repeated tokens? They complement the aggregate scores with concrete examples from both expert pruning and expert merging.

Across both pruning and merging, Full FT reduces repetition and restores more fluent continuations.

Example 1 · Expert Pruning

Gemma4 · AIMER · 50% expert retention

Prompt: what are forces (gravity, magnetism) made of?

Compressed model before Full FT

Imagine you are playing in a playground. You see a ball rolling down down down down down down down down down down down down down down down down down down down down down down down down down down down down down down down down down down down down down down down down

After Full FT

Imagine you are playing in a playground with a big, heavy ball. If you throw that ball down, it does not just go away and disappear. It stays in the ground, and if you throw it up, it stays in the sky. It does not just float on its own; it stays there because it is “stuck” to the ground or the sky. That “stickiness” is what we call gravity. … When we talk about forces like gravity or magnetism, we are talking about the rules of the universe. These rules are made of the way things are connected.

Example 2 · Expert Merging

Qwen3 · HC-SMoE · 62.5% expert retention

Prompt: How do overlays in NFL look so real? Like paint? How are they so accurate?

Compressed model before Full FT

Okay, let us talk about this like you are a kid who is just learning how to talk. So imagine you are watching a video of a football game, and you see this thing called an “overlay.” It looks like a flat layer of color that just sits on the ground, like paint. But here is the cool part: even though it looks like paint, it is not really paint at all. It is not like paint that you can touch or that you can walk on and it will fl fl fl fl fl fl fl fl fl fl fl fl fl fl fl

After Full FT

Okay, imagine you are playing football and you are on the field. You are running, jumping, dodging, and suddenly you see these bright lines and colors that look like paint on the grass. They are not just lines; they are little signs that show where the ball is going and where the players are. … Instead, they use a digital overlay technology. The computer knows exactly where every player is, where the ball is, and where the field lines are, so it draws the graphics in real time and makes them look as if they belong to the field.

Highlighted fragments mark repetition in the compressed model’s response. Ellipses indicate omitted text.

Robustness to Calibration Budget and Domain

We examine whether the aggregate strategy trend persists when the adjustment data change: a smaller C4 budget, and a shift from C4 to OpenR1-Math-220k. Both checks retain the one-epoch adjustment procedure and the same downstream evaluation suite.

A smaller calibration budget

3,000 → 1,024 C4 examples

Original paper Figure 5. Strategy recovery gains under a 1,024-example C4 budget. Full FT remains the leading aggregate strategy.
Figure 5. Mean recovery gain under a 1,024-example C4 budget. The one-epoch adjustment and downstream evaluation protocol are retained. Original figure (PDF) ↗

A different calibration domain

C4 → OpenR1-Math · 1,024 examples each

Original paper Figure 6. Strategy-level recovery under C4 versus OpenR1-Math calibration. The aggregate rankings remain strongly associated.
Figure 6. Recovery under C4 and OpenR1-Math calibration at the same 1,024-example budget. Upper-right indicates stronger recovery in both domains. Original figure (PDF) ↗

The leading aggregate pattern persists under both checks. With 1,024 C4 examples, Full FT gives the largest mean recovery (+2.78 percentage points), followed by router + all experts FT (+1.86) and Full KD (+1.77). Across calibration domains, the strategy rankings remain strongly associated (Spearman ρ = 0.973; Kendall τb = 0.897), and Full FT has the largest recovery averaged across the two domains.

The domain comparison uses six matched settings. This supports an aggregate trend, not an invariant ranking for every checkpoint: setting-level Spearman correlations range from 0.071 to 0.962.

Scope of the findings

The main conclusions concern small-budget, general-domain adjustment of the tested pruning and merging checkpoints under a unified implementation. Larger models, different training data or schedules, and system-level optimizations may change the trade-off. Expert Editing is reported separately as supplementary compatibility-mode evidence. PCA is an efficiency and capability-recovery stage; the study does not claim improved safety or factual reliability.

BibTeX

@inproceedings{hyeon2026beyond,
  title = {Beyond Retraining-Free {MoE} Compression:
           A Cost-Normalized Study of Post-Compression Adjustment},
  author = {Hyeon, Sieun and Do, Jaeyoung},
  booktitle = {Proceedings of the 2026 Conference on
               Empirical Methods in Natural Language Processing},
  year = {2026}
}

Paper figure

Open original figure ↗