CLAIM:基于不确定性度量的大语言模型开放域主动澄清领先方案
CLAIM: Leading Open-domain Active Clarification of Large Language Models with Uncertainty Measurement
- Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学高瓴人工智能学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究提出基于不确定性度量的CLAIM框架,无需人工标注,结合监督学习与强化学习训练统一澄清决策模型,实现开放域人机交互中低成本鲁棒的主动澄清。
AI中文摘要:
在开放域人机交互场景中,大语言模型(LLMs)常遇到模糊或不完整的用户查询,直接生成答案往往会产生过度泛化、错误或低信息的响应,而提出澄清问题可大幅提升交互质量。然而现有方法仍严重依赖人工标注数据或偏好对齐来解决两个核心挑战:何时需要澄清,以及应澄清查询的哪个方面。这种依赖带来了高昂的标注成本并限制了泛化能力。为应对这些挑战,我们提出了CLAIM,一种用于开放域主动澄清学习的不确定性驱动框架。CLAIM通过量化多模型间答案分歧引发的熵来计算查询不确定性,无需显式的人类偏好标注。随后,该不确定性信号被用于构建高质量的合成数据,使我们能够结合监督学习和强化学习训练统一的澄清决策模型。具体而言,我们提出了一种熵驱动的合成数据生成流程,该流程将基于熵的不确定性估计与语义聚类、基于推理的判断相结合,实现了对澄清需求的可靠自动标注。为训练CLAIM,我们将澄清过程形式化为结构化决策生成问题,并采用结合监督微调(SFT)与组相对策略优化(GRPO)的训练范式。实验结果表明,CLAIM无需依赖人工标注数据即可学习稳定且可泛化的澄清策略,为现实世界中基于LLMs的开放域交互中的主动理解提供了低成本且鲁棒的解决方案。
英文摘要:
In open-domain human-computer interaction scenarios, large language models (LLMs) frequently encounter user queries that are ambiguous or incomplete. In such cases, directly producing an answer often leads to overgeneralized, erroneous, or low-information responses. In contrast, asking clarifying questions can substantially improve interaction quality. However, existing approaches still rely heavily on manually annotated data or preference alignment to address two fundamental challenges: when clarification is necessary, and which aspect of the query should be clarified. This reliance incurs high annotation costs and limits generalization. To address these challenges, we propose CLAIM, an uncertainty-driven framework for active clarification learning in open-domain settings. CLAIM eliminates the need for explicit human preference annotations by quantifying query uncertainty through the entropy induced by answer disagreements across multiple models. This uncertainty signal is then used to construct high-quality synthetic data, enabling the training of a unified clarification decision model through a combination of supervised learning and reinforcement learning. Specifically, we propose an entropy-driven synthetic data generation pipeline that integrates entropy-based uncertainty estimation with semantic clustering and reasoning-based judgments, enabling reliable automatic annotation of clarification requirements. To train CLAIM, we formulate the clarification process as a structured decision generation problem and adopt a training paradigm that combines supervised fine-tuning (SFT) with group-relative policy optimization (GRPO). Experimental results demonstrate that CLAIM can learn stable and generalizable clarification strategies without relying on manually labeled data, offering a low-cost and robust solution for proactive understanding in real-world open-domain interactions with LLMs.