arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.12086cs.HCcs.AIcs.CL

创建用于个性感知大语言模型交互的原子用户模型

Creating an Atomic User Model for Personality-Aware Large Language Model Interaction

  • Indian Institute of Science (IISc)(印度科学学院)

机构由 AI 辅助整理,请以论文原文为准。

B. Sankar, Deepthika S, Pawni Yadav, Amogh A S

AI总结:

针对大语言模型助手个性化,提出原子用户模型(AUM),以可解释的身份核心与外壳表示用户,通过检索而非提示前缀实现高效风格匹配,显著提升声音识别率,且对默认服务最差的用户收益最大。

AI中文摘要:

基于大语言模型构建的助手被期望以用户本人的风格进行写作,而目前的主流方法是单通道式的:从对话历史中总结出偏好,再将其重新插入到上下文中。这种方法颠倒了推理的顺序。偏好只是相对稳定的人格结构在任务层面的表面表现,因此,仅存储偏好的系统在任务变化时需要重新学习用户。首先,我们刻画了“人格渗漏”现象,即提示词的语言表面携带了人格指纹,助手在无法访问其背后人格的情况下会模仿这一指纹。其次,我们提出了原子用户模型(AUM),这是一种可读的表示形式,将一个人组织为稳定的身份核心,并带有四个可解释的外壳(心理、认知与经验、行为、社会),以及记录内部冲突和真实性的跨外壳条目。第三,我们将AUM视为针对个人的检索索引,而非提示词前缀,其流程包括一个任务分类器、组件选择函数和带预算的检索器,在生成时返回一小部分字段。第四,我们使用十六个由语言模型模拟的参与者、六个风格敏感任务和三个随机种子对其进行了评估,并进行了检索器的合成规模扩展研究。检索八个字段在23%的上下文(211个词元对比915个)上达到了完整用户模型的风格保真度,在五点量表上比扁平偏好笔记提高了0.24分(p < 0.001,dz = 0.50),并将参与者自身声音的强制选择识别率从14.9%提高到42.7%(随机水平为25%)。四个预注册的对照组返回了零结果,表明效果源于表示本身而非对其的搜索。对于非个性化助手复现效果最差的参与者,这种收益最大(rho = -0.61,p = 0.013):个性化对默认服务最差的人价值最大。

英文摘要:

Assistants built on large language models are expected to write in their users' own voice. Most systems summarise the user's preferences and include the summary in the prompt. This is the wrong way round. Preferences are only the surface of a person and change with the task, while the underlying personality stays the same, so storing preferences alone means relearning the user afresh whenever the task changes. This paper makes four contributions. First, we describe an effect we call personality seepage: the wording of a prompt carries traces of the writer's personality, which the assistant copies without knowing the writer. Second, we propose the Atomic User Model (AUM), a readable profile with a stable identity core surrounded by four layers covering psychological, cognitive, experiential, behavioral, and social details, plus notes on inner conflict and authenticity. Third, instead of inserting the entire profile, we use AUM as a searchable index, in which a task classifier, a selection step, and a budgeted retriever pass along only a few relevant fields. Fourth, we test the pipeline with 16 simulated users, 6 style-sensitive tasks, and 3 seeds. Eight retrieved fields matched the writing quality of the whole profile, while using only 23 percent of the context (211 tokens instead of 915). They scored 0.24 points higher than a plain preference note on a five-point scale. Accuracy in picking a user's own writing from four samples rose from 14.9 to 42.7 percent, where guessing gives 25 percent. Four pre-registered controls showed no effect, so the gain comes from the profile's structure rather than the search method. Personalization helps most for the users for whom a generic assistant imitates them the worst.

补充信息

↑