arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

后训练中失去了什么?默认坍缩与跨多样视角的上下文内可引导性丧失

What Is Lost in Post-Training? Default Collapse and the Loss of In-Context Steerability Across Diverse Perspectives

Jessica Dierking, Itai Shapira, Niclas Boehmer

arXiv 2610.02614首次发表:更新:

发表机构

Hasso Plattner Institute; University of Potsdam; Harvard University(哈索·普拉特纳研究所; 波茨坦大学; 哈佛大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究揭示后训练会削弱语言模型在上下文中适应非偏好视角的能力,提出立场分布匹配作为兼顾多样性的替代目标。

AI 中文摘要

服务于异质人群的AI模型必须根据每个用户和情境采取适当的原则。虽然已有研究表明后训练会缩小大型语言模型所表达的观点范围,但先前的工作侧重于默认行为,而非适应上下文信息的能力。我们证明,后训练还会降低模型在上下文中被引导至其未被训练偏好的视角的能力。在受控实验中,我们将模型微调至文化价值分歧的一方,并评估整个训练过程中的检查点。被训练的一方在常规使用中变得越来越占主导地位,而识别并忠实执行对立观点的能力则下降。这些发现指出了在优先考虑单一价值观与保留服务多样化利益相关者所需的技术能力之间的张力。最后,我们提出并分析了一种替代目标,该目标在规定的表达视角分布下最大化奖励,并将立场分布匹配作为实际实施方案呈现。

英文摘要

AI models serving a heterogeneous population must act on the principles appropriate to each user and context. While post-training has been shown to narrow the views large language models express, prior work has focused on default behavior rather than the ability to adapt to in-context information. We show that post-training also degrades a model's ability to be steered in-context toward perspectives it was not trained to favor. In controlled experiments, we fine-tune models toward one side of cultural-value disagreements and evaluate checkpoints throughout training. The trained side becomes increasingly dominant in ordinary use, while the ability to recognize and faithfully enact the opposing view declines. These findings point to a tension between prioritizing a single set of values and preserving the technical capacity needed to serve diverse stakeholders. Finally, we propose and analyze an alternative objective that maximizes reward subject to a prescribed distribution over expressed perspectives, and present stance-distribution matching as a practical implementation.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑