arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.18085cs.CL

面向任务型对话的角色引导大语言模型智能体

Persona-Guided LLM Agents for Task-Oriented Dialogue

发表机构肯塔基大学
查看机构详情
  • University of Kentucky(肯塔基大学)

机构由 AI 辅助整理,请以论文原文为准。

Maryam Shoaeinaeini, Brent Harrison, A. B. Siddique

首次发表
浏览论文内容

中文总结 AI 辅助

该研究构建无需训练的框架,通过三种情境评估多模型在SGD数据集上的表现,发现适应用户人格存在权衡,Try情境的线索式适应可更可靠实现感知人格的任务型对话。

中文摘要 AI 辅助

已有研究表明,大语言模型(LLM)能在开放式文本生成中展现多样的人格特质,但尚不清楚它们在目标导向的对话中能否做到这一点且不损害任务完成度,以及适应用户人格是否会提升交互质量。我们在任务型对话(TOD)中研究这些问题,该场景下系统通过多轮交互帮助用户达成目标。我们构建了一个无需训练的框架,模拟两个LLM之间的TOD交互:一个展现目标人格的用户智能体,和一个在完成任务的同时适应用户的系统智能体。为隔离适应的效果,我们在三种情境下改变系统对用户人格的知晓程度:Neutral情境下系统未收到任何人格信息;Try情境下系统从对话线索中推断人格;Oracle情境下系统被明确告知人格。我们在Schema-Guided Dialogue(SGD)数据集的酒店和餐厅对话上,针对大五人格及其相反两极,评估GPT-4o、Qwen3-Next-80B和Gemini 2.0 Flash。我们发现,用户智能体可展现人格,同时系统维持较强的任务性能,尽管部分特质的实现可靠性远低于其他特质。适应用户人格会提升约束满足度、信息率和用户满意度,但会降低真实性,揭示了个性化与任务基础间的权衡。当目标特质被强烈表达时,Oracle的收益会增加,而Try的收益对实现强度基本不敏感。总体而言,Try中的基于线索的适应最能解决该权衡,为无需微调的感知人格的TOD提供了更可靠的路径。

英文摘要

Prior work has shown that large language models (LLMs) can express diverse personality traits in open-ended text generation. However, it remains unclear whether they can do so in a goal-directed dialogue without compromising task completion, and whether adapting to the user's personality improves the interaction quality. We study these questions in task-oriented dialogue (TOD), where a system helps a user accomplish a goal via multi-turn interaction. We build a training-free framework that simulates a TOD interaction between two LLMs: a user agent that exhibits a target personality and a system agent that adapts to the user while completing the task. To isolate the effect of adaptation, we vary how much the system knows about the user's personality across three conditions. In Neutral, the system receives no personality information. In Try, it infers the personality from dialogue cues. In Oracle, it is given the personality explicitly. We evaluate GPT-4o, Qwen3-Next-80B, and Gemini 2.0 Flash on Hotel and Restaurant dialogues from the Schema-Guided Dialogue (SGD) dataset, across the Big Five traits and their opposite poles. We find that the user agent can express personality while the system maintains strong task performance, although some traits are realized far less reliably than others. Adapting to the user's personality improves constraint satisfaction, inform rate, and user satisfaction, but lowers truthfulness, revealing a trade-off between personalization and task-grounding. Oracle's gains grow when the target trait is strongly expressed, whereas Try's gains are largely insensitive to realization strength. Overall, cue-based adaptation in Try best resolves this trade-off and offers a more reliable route to personality-aware TOD without fine-tuning.

补充信息

↑