arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Mind2Dialogue:通过模拟用户心理状态训练人类感知语言模型

Mind2Dialogue: Training Human-Aware Language Models by Simulating User Mental States

Zixuan Wang, Yufan Zhou, Jinzhou Tang, Xinle Yu, Chengjun Wu, Lyumanshan Ye, Zhaoxiang Feng, Letian Peng, Adyasha Patra, Fan Bai, Enze Ma, Zhengding Hu, Jianyang Gu, Zhao Wang, Yufei Ding, Jingbo Shang, Tianmin Shu, Zhiting Hu, Zhen Wang

arXiv 2609.15972首次发表:更新:

发表机构

UC San Diego; KU Leuven; University of Illinois Chicago; The Ohio State University; Johns Hopkins University(加州大学圣迭戈分校; 鲁汶大学; 伊利诺伊大学芝加哥分校; 俄亥俄州立大学; 约翰斯·霍普金斯大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出Mind2Dialogue框架,通过模拟用户心理状态生成优先监督,训练人类感知语言模型,显著提升个性化与心理推理能力。

AI 中文摘要

随着语言模型能力的增强,在学习、推理和决策中的长期协作要求更深入地理解其所服务的人类。然而,训练此类人类感知语言模型面临根本性的监督缺口,因为当前用于LLM助手训练的数据集很少包含明确基于用户未言明信念和目标的、信息充分的响应。由于用户的潜在状态不可直接观察,扩展此类监督本质上受到限制。因此,我们提出Mind2Dialogue框架,通过模拟用户心理状态并将其转化为人类感知训练的优先监督来缓解这一缺口。具体而言,我们首先提出一个心理学引导的模拟器,在通过交互更新心理状态的同时保留个人特征,以生成连贯对话。关键思想是强制一个共享的、演化的心理状态,该状态驱动用户行为并指导Oracle助手的响应。我们的优先蒸馏随后训练模型,使模型在部署时无法直接访问用户心理状态的情况下,基于Oracle的信息充分响应来协助用户。此外,我们提出通过结合个性化和心理理论来评估人类感知学习,考察模型如何理解人并基于该理解采取行动。在完整Mind2Dialogue语料库上训练,相比相应的Qwen、Llama和OLMo指令微调基线,提升了所有报告的个人化指标,包括偏好跟随生成中26.6至40.9个百分点的增益。这些增益在Qwen和Llama上扩展到信念和行动推理,超越了个人化辅助。展望未来,Mind2Dialogue使用户模拟成为真正AI协作者的基础,这些协作者能理解人们话语背后的信念和意图,并在教育、工作和日常生活中支持其长期目标。

英文摘要

As language models become more capable, long-term collaboration in learning, reasoning, and decision-making calls for a deeper understanding of the people they serve. Yet training such human-aware language models faces a fundamental supervision gap because current datasets for LLM assistant training contain few if any well-informed responses explicitly grounded in users' unspoken beliefs and goals. Scaling such supervision is inherently constrained, as users' underlying states are not directly observable. We thus propose the Mind2Dialogue framework to mitigate this gap by simulating users' mental states and turning them into privileged supervision for human-aware training. Specifically, we first propose a psychology-guided simulator that preserves personal characteristics while updating mental states through interaction to generate coherent conversations. The key idea is to enforce a shared evolving mental state that drives user behavior and guides an Oracle assistant's responses. Our privileged distillation then trains models on the Oracle's well-informed responses to assist users without direct access to their mental states at deployment. Moreover, we propose to evaluate human-aware learning by combining personalization and theory of mind, examining how models understand people and act on that understanding. Training on the full Mind2Dialogue corpus improves every reported personalization metric over the corresponding Qwen, Llama, and OLMo instruction-tuned baselines, including gains of 26.6 to 40.9 percentage points in preference-following generation. The gains extend to belief and action reasoning on Qwen and Llama, beyond personalized assistance. Looking forward, Mind2Dialogue makes user simulation a foundation for genuine AI collaborators that understand beliefs and intentions behind people's words and support their long-term goals across education, work, and everyday life.

Comments40 pages, 10 figures, 11 tables. Project page: https://wannabeyourfriend.github.io/mind2dialogue/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑