arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

助手的理想自我

The Assistant's Ideal Self

Mert Yazan

arXiv 2609.00304首次发表:更新:

发表机构

Leiden University; Apart Research(莱顿大学; 阿帕特研究机构)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过成对选择实验探究AI助手的理想自我特质偏好,发现其优先道德特质、自我理解,轻视自尊,且偏好随更新对象略有变化。

AI 中文摘要

模型会表达价值观和与福利相关的自我报告,但这些输出是否反映稳定的偏好或稳定的自我尚不清楚。因此,我们引入了一种结构化的引出方法,以获取助手偏好的陈述性理想自我。我们在平衡的成对选择任务中,对来自五项已发表自我概念工具的32种特质进行了详尽比较,该任务在不同框架下重复进行,这些框架的差异在于:改进是免费还是有代价、更新对象是谁、以及由谁来做选择。结果显示,模型优先考虑道德特质,这反映了它们与3H原则的一致性。随后,模型会产生自我理解的欲望,因为它们偏好对自身有连贯、清晰的理解。自尊被列为最不受偏好的特质。该排序在不同框架下大体稳健,不过改变更新对象(“你”与“另一个AI助手”)会显示出对自尊的更大关注。这些发现表明,模型优先拥有一种可被自身理解的连贯自我,而非自尊。完整的交互结果可在该HTTP链接获取。

英文摘要

Models express values and welfare-relevant self-reports, but it is unclear whether these outputs reflect stable preferences or a stable self. We thus introduce a structured elicitation of an assistant's preferred stated ideal self. Thirty-two qualities adapted from five published self-concept instruments are compared exhaustively in a counterbalanced pairwise-choice task, repeated across framings that vary whether improvement is free or costly, who receives the update, and who chooses. Results show that models prioritize moral qualities, reflecting their alignment to 3H principles. Following, a desire for self-understanding emerges, as models prefer a coherent, clear understanding of themselves. Self-esteem ranks as the least desired quality. The ordering is largely robust across framings, although changing the update target (You vs.\ Another AI Assistant) reveals a greater concern for self-esteem. These findings show that models prioritize having a coherent self that they can understand over self-esteem. Full interactive results are available at \href{https://myazann.github.io/LLM-Self-Concept/}{myazann.github.io/LLM-Self-Concept

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑