arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.14250cs.CL

分离问题:大语言模型未意识到提示之外的人

The Severance Problem: LLMs are Unaware of the Person Beyond the Prompt

发表机构德克萨斯大学奥斯汀分校
查看机构详情
  • The University of Texas at Austin(德克萨斯大学奥斯汀分校)

机构由 AI 辅助整理,请以论文原文为准。

Dor Litvak, Liu Leqi

首次发表
浏览论文内容

中文总结 AI 辅助

研究指出个人AI助手存在不良行为,源于语言模型的“分离问题”。提出通过“分离模式”将结构化无知纳入语言模型上下文的解决方案,经实证,五个模型家族采用此模式后能减少不良行为,模型在信息缺失时会询问澄清问题。

中文摘要 AI 辅助

个人人工智能助手因能通过自动化日常任务、支持重要决策和协助处理日常个人事务来提升日常生活而备受关注。尽管近期技术进步迅速,但这些助手仍表现出不良行为,如谄媚、过度自信和幻觉。我们认为这些失败源于一个基本限制:语言模型缺乏对给定上下文之外的人的明确表示,即‘分离问题’。即便有丰富的个人上下文和强大的常识推理能力,当前人工智能助手仍无法表示关于用户未知的内容。我们提出一种简单解决方案:通过‘分离模式’将结构化无知纳入语言模型上下文,该模式明确勾勒出模型在用户物理性、时间性、后果、连续性、多样性和内在性等方面缺乏知识的维度。实证研究表明,在五个模型家族中,有了分离模式,助手能持续减少谄媚、有害建议和幻觉。值得注意的是,有该模式的模型在用户信息缺失时会询问澄清问题,而非从不完整的用户信息中自信地推断。

英文摘要

Personal AI assistants have attracted significant interest for their potential to enhance everyday life by automating routine tasks, supporting consequential decisions, and assisting with everyday personal matters. Yet despite rapid recent technical advances, these assistants continue to exhibit undesirable behaviors, such as sycophancy, overconfidence, and hallucination. We argue that these failures stem from a fundamental limitation: language models lack an explicit representation of the person beyond the context they are given, which we term as the \textbf{Severance Problem}. Even with rich personal context and strong commonsense reasoning capabilities from the backbone model, current AI assistants fail to represent what remains unknown about the user. We propose a simple solution: incorporating structured ignorance into the language model context via the \textbf{Severance Schema}, which explicitly outlines dimensions along which the model lacks knowledge about the user, including physicality, temporality, consequences, continuity, multiplicity, and interiority. Empirically, across five model families, with the Severance Schema, the assistant consistently reduces sycophancy, harmful advice, and hallucination. Notably, models with the schema ask clarifying questions when information about the user is missing, rather than confidently extrapolating from incomplete user information.

↑