arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.11426cs.CL

收敛是不可避免的吗?将输出同质性追溯至基础模型

Is Convergence Inevitable? Tracing Output Homogeneity Back to Base Models

  • University of British Columbia(不列颠哥伦比亚大学)

机构由 AI 辅助整理,请以论文原文为准。

Alexandrine Fortier, Hazel Chen, Peter West

AI总结:

该研究指出大语言模型的语义收敛或源于预训练目标,仅靠对齐后干预难缓解,通过实验证实基础模型可被提示诱导类指令式坍缩。

AI中文摘要:

大语言模型(LM)内容多样性的缺失被广泛归因于对齐过程,但该“坍缩”究竟在流程中如何发生、从何处开始尚不明确。本文认为,输出同质性可能是在预训练阶段习得的,仅在对齐过程中被“揭示”或放大。具体而言,我们在首个对齐阶段——指令调优阶段(SFT)就观察到语义收敛,表明同质性可能已存在于对齐前的模型中。为探究这一点,我们开展受控SFT实验,研究训练数据如何影响特定输入/输出对的输出收敛,发现SFT数据可揭示并放大收敛,但不会引入收敛,支持其作为催化剂而非成因的作用。为进一步验证同质性是否起源于对齐之前,我们测量基础模型的收敛情况,发现仅通过提示(无需对齐)就能诱导类指令式坍缩。综上,我们的结果表明,语义收敛可能自然产生于LM训练的目标,仅靠对齐后干预难以缓解。

英文摘要:

The lack of diversity in LM content is widely attributed to the alignment process, but how and where exactly in the pipeline this collapse begins is unknown. We argue that output homogeneity is likely learned during the pretraining phase, and only revealed or magnified during the alignment process. Specifically, we find that semantic convergence is observed from the first alignment stage--the instruction-tuning phase (SFT)--suggesting that homogeneity might already exist in the pre-alignment model. To investigate this, we conduct controlled SFT experiments examining how training data influences output convergence on specific input/output pairs. We find that convergence can be revealed and amplified, but not introduced by the SFT data, supporting its role as a catalyst rather than a cause. To further test whether homogeneity originates before alignment, we measure convergence in base models. We find that instruct-like collapse can be induced through prompting alone, even without alignment. Taken together, our results suggest that semantic convergence may arise naturally from the objectives underlying LM training, making it difficult to mitigate through post-alignment interventions alone.

↑