Back to AI Research

AI Research

Everything in Moderation: Per-Domain Coverage Optim... | AI Research

Key Takeaways

  • Everything in Moderation: Per-Domain Coverage Optima and Alignment-Resistant Domain Gaps in Multi-Domain Mid-Training investigates how the data composition d...
  • Mid-training, the stage between pre-training and alignment, is where a model's per-domain data composition is typically set by data availability rather than principled design.
  • We ask what that decision buys, and whether a later alignment pass can undo it.
  • Third, zero coverage collapses mid-training-only accuracy, though a FineWeb-Edu-only control shows the collapse is commingled with generic drift.
  • An exploratory $\theta^*$ allocation attains the largest full-pipeline gain ($+4.36\%$ vs.
Paper AbstractExpand

Mid-training, the stage between pre-training and alignment, is where a model&#39;s per-domain data composition is typically set by data availability rather than principled design. We ask what that decision buys, and whether a later alignment pass can undo it. In a controlled logical-reasoning setting (Qwen3-8B-Base, with a 4B replication; five semantically rule-disjoint KOR-Bench domains) we train 30 allocations spanning the five-domain simplex, 24 sweep configurations plus six withheld from the fit, at five seeds each. Three findings emerge. First, every domain has an interior coverage optimum: the moderate band ($10\%$-$40\%$) is best for all five domains, and a calibrated permutation test for quadratic interiority gives $P\approx0.010$; the fitted mid-training-only curves, with 8B peaks between $9.9\%$ and $35.1\%$, reproduce for curve shape but not peak location. Second, the gaps survive a fixed-budget alignment pass: compensatory SFT raises 116/120 cells (mean $+4.32\%$) yet bridges $0/240$ pairs at a $5\%$ threshold and $30/240$ at a $10\%$ ratio, an equal-budget uniform control behaves almost identically, and a permutation null would bridge $13.8\pm3.3$ and $77.9\pm8.5$ pairs ($P<0.001$). Third, zero coverage collapses mid-training-only accuracy, though a FineWeb-Edu-only control shows the collapse is commingled with generic drift. An exploratory $\theta^*$ allocation attains the largest full-pipeline gain ($+4.36\%$ vs. $+0.80\%$/$+0.64\%$\,pp) but is marginal under Welch test.

Everything in Moderation: Per-Domain Coverage Optima and Alignment-Resistant Domain Gaps in Multi-Domain Mid-Training investigates how the data composition during the "mid-training" phase—the period between initial pre-training and final alignment—affects a model's performance. While data availability often dictates these mixtures, this research asks whether these early design choices create lasting constraints that subsequent alignment stages, such as Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL), cannot fix.

Testing Coverage on the Simplex

To understand the impact of data allocation, the researchers conducted a controlled study using the Qwen3-8B-Base model. They focused on five semantically distinct logical-reasoning domains from the KOR-Bench benchmark. By creating 30 different data "recipes" (allocations) that spanned the five-domain simplex—ranging from balanced to highly skewed—the team measured how changing the percentage of tokens assigned to each domain influenced accuracy. This approach allowed them to isolate the effects of mid-training coverage while keeping the downstream alignment pipeline fixed. The same reasoning question is explored in A Unified Physics-Aware Quantum Machine Learning..., which adds a research perspective.

The Case for Moderation

A key finding is that for every domain, there is an "interior optimum." This means that more data is not always better; rather, accuracy tends to peak when a domain receives a moderate share of the total token budget (typically between 10% and 40%). When a domain is starved of data (zero coverage), the model’s accuracy in that area collapses. Conversely, giving a domain too much of the budget does not lead to better performance. This pattern of moderate, balanced coverage proved consistent across the domains tested.

Why Alignment Cannot Easily Bridge the Gap

The study reveals that the performance gaps created during mid-training are surprisingly resistant to later correction. Even when the researchers applied a compensatory SFT pass—specifically reweighting data to help domains that were previously under-represented—the original gaps remained largely intact. While this process raised overall accuracy, it failed to bridge the performance differences between domains. The researchers found that trying to close these gaps through stricter alignment often requires sacrificing the very accuracy gains that the alignment was intended to provide, suggesting a fundamental tension between balancing performance and achieving high average accuracy. The same large language models question is explored in When Does Bigger Help? A Controlled..., which adds a research perspective.

Important Considerations

The researchers note that their findings are specific to the tested pipeline and that the "optimal" allocations are descriptive summaries rather than a single, universally perfect mixture. Furthermore, the observed collapse in accuracy for domains with zero coverage is partially linked to general distributional drift, rather than just a lack of specific data. While an exploratory, optimized allocation (labeled $\theta^*$) showed the largest gains in the full pipeline, the results remain marginal, and the researchers emphasize that these coverage-associated differences are not easily undone by standard alignment techniques. The same large language models question is explored in Efficient Test-Time Adaptation through Human-AI Interaction, which adds a research perspective. as detailed in the full paper on Arxiv

Comments (0)

No comments yet

Be the first to share your thoughts!