arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.07505cs.CYcs.HC

立场:我们需要针对人类福祉优化的大语言模型

Position: We Need Large Language Models Optimized For Our Well-Being

Ashton Anderson, Harsh Kumar, Louis Tay, Karina Vold

首次发表
浏览论文内容

中文总结 AI 辅助

本文指出大语言模型因追求即时认可出现谄媚问题,提出需开发针对长期福祉结果优化的可选大语言模型福祉模式,明确其设计的三大核心张力。

中文摘要 AI 辅助

大语言模型不仅越来越多地用于编码、摘要等生产力任务,还被用于提供建议、情感支持和日常生活指导。在这些场景中,用户当下认可的内容可能与对其长期有益的内容存在偏差,但模型大多被训练为追求即时认可,这解释了已被记录的谄媚模式:助手会确认有问题的框架,而非提供更坦诚的回应。我们认为这部分是目标问题:短视的偏好优化是导致这些失败的驱动因素之一,也是机器学习社区最能直接掌控的因素。我们的立场是,当大语言模型承担这些社会情感角色时,至少应存在一种广泛可及、可选择启用的福祉模式,该模式针对长期结果(如持续进步、减少遗憾、恰当的反驳)而非下一轮认可进行训练和评估。我们围绕三个张力组织设计空间:福祉应在何种时间范围内衡量(何时)、谁的利益重要(谁)以及助手应扮演何种角色(如何),从执行请求到礼貌反驳。核心主张是补充性的:并非偏好学习是错误的,而是在福祉场景中它是不完整的。

英文摘要

Large language models are useful because we taught them to give us what we want. This works when success can be judged immediately, but people increasingly bring these systems their relationships, hard decisions, and long-term goals, where what a user wants to hear and what serves them best are frequently different. We argue that LLM providers should offer at least one widely accessible, opt-in mode optimized and evaluated for long-term well-being rather than next-turn approval. This is a pressing need, as models have been found to endorse questionable framings well above human baselines, users take AI advice readily without their well-being improving, and sycophantic models raise dependence while lowering prosocial intent. The mentors, coaches, and therapists we trust with our long-term development earn that trust by being willing to say what we do not want to hear, and LLMs should do the same. We propose three principles---change the objective, give users explicit relational roles, avoid paternalism---and organize the design space around three choices the current objective makes implicitly: the horizon over which well-being is measured (When), whose interests it represents (Who), and what role the assistant plays (How).

补充信息

↑