arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.32758cs.CLcs.AI

人格效应在语言模型中的泛化程度如何?

How Far Do Persona Effects Generalize in Language Models?

Yufan Zhou, Yuxuan Liu, Enze Ma, Lyumanshan Ye, Zhongqi Yue, Robin De Croon, Yucheng Jin, Katrien Verbert, Zhao Wang

首次发表
浏览论文内容

中文总结 AI 辅助

本研究探讨人格提示效应在语言模型中的泛化性,发现其可预测但不可移植,并证明目标重拟合优于跨模型借用增益,且正则化更新在提示迁移中表现最佳。

中文摘要 AI 辅助

人格提示要求语言模型以特定类型的人的身份作答。我们测试从这些效应中学到的关系是否能预测对新问题的回答,并跨模型和提示保持有用。在57个属性、三个行为领域以及七对开放的7至9B检查点中,人格效应可以是可预测的但不可移植的。在OLMo-3和Qwen2.5中,分离的属性增益和任务增益显著优于共享缩放预测,其中在OLMo-3中证据最为稳健。在该模型中,目标重拟合显著优于从其他六对中借用的增益。跨模型迁移时,在大多数方向上,具有单一幅度的借用增益在共享缩放下表现不佳;允许两个目标参数消除了显著损失,但相对于目标共享缩放没有显著收益。在改写提示后,在所有六对测试中,重拟合显著优于单一幅度的重用,而改变示例或国家背景通常保留重用价值。在测试的提示迁移中,正则化更新在每属性64个目标问题上的中位数优于两种重用策略。一项单独的调查比较发现,选择响应性更强的检查点可能恶化人类拟合;响应性与训练状态混淆,温度校准在很大程度上消除了这一成本,但未消除群体排序中的错误。在测试的增益表示中,表面迁移可能来自目标校准;源关系必须在校准和正则化之外增加预测价值。代码和数据可在该https URL获取。

英文摘要

Persona prompts ask language models to answer as particular kinds of people. We test whether relationships learned from these effects predict responses to new questions and remain useful across models and prompts. Across 57 attributes, three behavioral domains, and seven pairs of open 7 to 9B checkpoints, persona effects can be predictable without being portable. Separate attribute and task gains improve prediction beyond shared scaling significantly in OLMo-3 and Qwen2.5, with the most robust evidence in OLMo-3. In that model, target refitting significantly outperforms gains borrowed from each of the other six pairs. Across model transfers, borrowed gains with one amplitude underperform shared scaling in most directions; allowing two target parameters removes the significant losses but yields no significant benefit over target shared scaling. After rewording, refitting significantly outperforms reuse with one amplitude in all six tested pairs, while changes of examples or country context often preserve reuse value. In the tested prompt transfers, regularized updates outperform both reuse strategies in median at 64 target questions per attribute. A separate survey comparison finds that selecting the more responsive checkpoint can worsen human fit; responsiveness is confounded with training status, and temperature calibration largely removes this cost but not errors in group ordering. Within the tested gain representation, apparent transfer can come from target calibration; source relationships must add predictive value beyond calibration and regularization. Code and data are available at https://github.com/thzva/persona-gain

补充信息

↑