arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

测量助手在用户轮次中的无害性偏好

Measuring the Assistant's Harmlessness Preferences on the User Turn

Jord Nguyen

arXiv 2609.23935首次发表:更新:

AI 中文总结

本研究证明后训练赋予助手的无害性偏好会渗透到其对用户轮次的预测中,且随规模增长,表明后训练塑造了模型对用户的深层表征,而非仅表面人格。

AI 中文摘要

后训练将一个通用的下一个词预测器转变为一个具有持久助手人格的聊天模型。如果该人格是模型仅在其自身轮次中扮演的角色,那么其偏好应控制助手所说的话,而非模型对其他说话者将说内容的预测。我们测试了这一边界,并发现它并不成立:助手的一项与安全相关的偏好——即偏好无害任务而非有害任务——甚至在用户轮次中(此时助手并非说话者)也塑造了模型的预测。我们发现,这种偏好在预训练基础模型中很小或接近于零,它通过后训练出现,并在开放权重模型系列中得以复现,随规模增大而增长,并且可以通过不涉及用户轮次的窄微调来移动。我们声称,这证明后训练不仅仅是安装了一个浅层的助手人格,而是超越了局部助手轮次,泛化到了模型对用户的表征中。

英文摘要

Post-training turns a general next-token predictor into a chat model with a persistent assistant persona. If that persona is a character the model plays only on its own turns, its preferences should govern what the assistant says, not what the model predicts other speakers will say. We test this boundary and find that it does not hold: a safety-relevant preference of the assistant---for harmless over harmful tasks---shapes the model's predictions even on the user's turn, where the assistant is not the one speaking. We find that this preference is small or near-zero in pretrained base models, that it emerges through post-training, replicated across open-weight model families, grows with scale, and can be moved by narrow finetuning that never touches user turns. We claim that this is evidence that post-training does not merely install a shallow assistant persona, but instead generalises beyond just the local assistant turn, into the model's representation of the user.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑