发表机构
University of Science and Technology Beijing; University of South China(北京科技大学; 南华大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对人机对话中用户侧隐性冲突检测难题,构建基准UC-Bench并提出约束引导合成方法SynUC,生成训练集UC-Data,使轻量模型Qwen3.5-4B超越更大模型。
AI 中文摘要
在人机对话中,用户的后续话语可能与先前意图产生隐性冲突,导致大语言模型误解用户需求并生成不恰当的回复。一个可靠的对话系统应在生成回复之前主动检测用户侧冲突,并在必要时寻求澄清。然而,先前的工作主要聚焦于大语言模型侧的冲突,对用户侧冲突的探索不足。为填补这一空白,我们构建了UC-Bench,一个用于评估用户侧冲突检测的人工标注基准。初步实验表明,现有的大语言模型在此任务上表现不佳,尤其是当冲突源于基于对话历史的隐性不兼容时。为了在有限训练数据下改进轻量级大语言模型,我们研究了用于用户侧冲突检测的数据合成方法。现有的合成方法未显式建模历史与当前用户话语之间的隐性不兼容性,难以捕捉冲突的演变并生成可靠标注的隐性冲突样本。我们提出SynUC,一种约束引导的合成方法,在约束空间中表示用户侧冲突,并使用SPEAKING框架指导可追踪的约束变换。将SynUC应用于WildChat,我们构建了UC-Data,一个包含2,487个样本的用户侧冲突训练集。在UC-Bench上,基于UC-Data训练的Qwen3.5-4B优于更大的通用大语言模型(如Claude Opus 4.8),也优于使用现有方法合成数据训练的同一骨干模型。
英文摘要
In Human-LLM dialogue, follow-up user utterances may implicitly conflict with earlier intents, leading the LLM to misinterpret user needs and generate inappropriate responses. A reliable dialogue system should proactively detect user-side conflicts before generating a response and seek clarification when necessary. However, prior work has largely focused on LLM-side conflicts, leaving user-side conflicts underexplored. To fill this gap, we construct UC-Bench, a human-annotated benchmark for evaluating user-side conflict detection. Preliminary experiments show that existing LLMs struggle with this task, especially when conflicts arise from implicit incompatibilities grounded in dialogue history. To improve lightweight LLMs with limited training data, we investigate data synthesis for user-side conflict detection. Existing synthesis methods do not explicitly model the implicit incompatibilities between historical and current user utterances, making it difficult to capture the evolution of conflicts and to generate reliably labeled implicit conflict samples. We propose SynUC, a constraint-guided synthesis method that represents user-side conflicts in a constraint space and uses the SPEAKING framework to guide traceable constraint transformations. Applying SynUC to WildChat, we construct UC-Data, a user-side conflict training set containing 2,487 samples. On UC-Bench, Qwen3.5-4B trained on UC-Data outperforms larger general-purpose LLMs such as Claude Opus 4.8, as well as the same backbone trained on data synthesized by existing methods.
Comments24 pages, 13 figures