arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

一切适度:多领域中期训练中的每领域覆盖最优与对齐抵抗领域差距

Everything in Moderation: Per-Domain Coverage Optima and Alignment-Resistant Domain Gaps in Multi-Domain Mid-Training

Yunpeng Xu, Kun Zheng

arXiv 2609.09081首次发表:更新:

AI 中文总结

本研究通过受控实验发现多领域中期训练中每领域存在内部覆盖最优,且该最优在对齐后仍保持,零覆盖会导致性能崩溃。

AI 中文摘要

中期训练是预训练与对齐之间的阶段,在此阶段,模型每领域的数据组成通常由数据可用性而非原则性设计决定。我们探究这一决策的收益,以及后续的对齐过程能否抵消其影响。在受控的逻辑推理设置中(Qwen3-8B-Base,并辅以4B复现;五个语义规则不相交的KOR-Bench领域),我们训练了30种分配方案,覆盖五领域单纯形,其中24种为扫描配置,6种留出用于拟合,每种使用5个随机种子。我们得出三个发现。第一,每个领域都存在内部覆盖最优:中等区间(10%-40%)对所有五个领域均为最佳,且针对二次内部性的校准置换检验给出P≈0.010;仅中期训练的拟合曲线(8B峰值介于9.9%至35.1%之间)在曲线形状上可复现,但峰值位置不可复现。第二,这些差距在固定预算的对齐过程中依然存在:补偿性SFT提升了116/120个单元(平均+4.32%),但在5%阈值下仅弥合0/240对,在10%比率下弥合30/240对;等预算的均匀对照表现几乎相同,而置换零假设会弥合13.8±3.3和77.9±8.5对(P<0.001)。第三,零覆盖导致仅中期训练准确率崩溃,但FineWeb-Edu-only对照显示该崩溃与一般性漂移混杂。探索性的θ*分配实现了最大的全流程增益(+4.36%,对比+0.80%/+0.64%个百分点),但在Welch检验下不显著。

英文摘要

Mid-training, the stage between pre-training and alignment, is where a model's per-domain data composition is typically set by data availability rather than principled design. We ask what that decision buys, and whether a later alignment pass can undo it. In a controlled logical-reasoning setting (Qwen3-8B-Base, with a 4B replication; five semantically rule-disjoint KOR-Bench domains) we train 30 allocations spanning the five-domain simplex, 24 sweep configurations plus six withheld from the fit, at five seeds each. Three findings emerge. First, every domain has an interior coverage optimum: the moderate band ($10\%$-$40\%$) is best for all five domains, and a calibrated permutation test for quadratic interiority gives $P\approx0.010$; the fitted mid-training-only curves, with 8B peaks between $9.9\%$ and $35.1\%$, reproduce for curve shape but not peak location. Second, the gaps survive a fixed-budget alignment pass: compensatory SFT raises 116/120 cells (mean $+4.32\%$) yet bridges $0/240$ pairs at a $5\%$ threshold and $30/240$ at a $10\%$ ratio, an equal-budget uniform control behaves almost identically, and a permutation null would bridge $13.8\pm3.3$ and $77.9\pm8.5$ pairs ($P<0.001$). Third, zero coverage collapses mid-training-only accuracy, though a FineWeb-Edu-only control shows the collapse is commingled with generic drift. An exploratory $θ^*$ allocation attains the largest full-pipeline gain ($+4.36\%$ vs. $+0.80\%$/$+0.64\%$\,pp) but is marginal under Welch test.

CommentsAbout to commit to ICLR

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑