arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.24290cs.AI

智能体应在何时以及如何澄清?CIGAsk:通过反事实信息增益教LLM进行澄清

When and How Should an Agent Clarify? CIGAsk: Teaching LLMs to Clarify via Counterfactual Information Gain

Yunxiang Li, Xixin Wu, Helen Meng

首次发表
浏览论文内容

中文总结 AI 辅助

针对LLM面对模糊查询时错误作答的问题,提出CIGAsk强化学习方案,通过反事实信息增益与非对称歧义奖励,同时教会模型何时提问及如何提问,在多个澄清基准上超越更强基线。

中文摘要 AI 辅助

指令微调的大语言模型(LLM)在面对表述不明确的查询时,往往倾向于对单一解释做出承诺,而不是请求澄清,从而产生自信但错误的答案。在我们的实验中,仅靠提示(prompting)是不够的:模型要么对每个查询都请求澄清,要么提出模糊的问题,无法恢复缺失的信息。解决这一失败需要学习两种相互关联的技能:何时提问而非回答,以及如何提出能够恢复歧义信息的问题。现有方法要么只解决其中一种技能,要么需要单独训练的批评模型。我们提出CIGAsk,一种强化学习(RL)方案,通过多轮GRPO循环中的两个互补奖励信号来同时教授这两种技能。反事实信息增益(CIG)在冻结的参考模型下,比较有和没有用户响应时黄金答案的对数似然,提供逐轮信用,指导如何提问。非对称歧义奖励(Asymmetric Ambiguity Bonus)根据黄金歧义标签在终止符处分配带符号的奖励,指导何时提问。在涵盖表格、段落和开放域问答的三个澄清基准测试中,CIGAsk-7B尽管使用较小的骨干网络,仍优于最强的外部基线。它还能跨数据集迁移,无需逐数据集调整,同时在分布外基准上保持单轮问答性能。

英文摘要

Instruction-tuned LLMs faced with underspecified queries often commit to a single interpretation rather than ask for clarification, producing confidently wrong answers. In our experiments, prompting alone is insufficient: models either ask for clarification on every query or ask vague questions that fail to recover the missing information. Addressing this failure requires learning two coupled skills: when to ask rather than answer and how to ask a question that recovers the disambiguating information. Existing recipes either address only one of these skills or require a separately trained critic. We propose CIGAsk, an RL recipe that teaches both skills through two complementary reward signals within a multi-turn GRPO loop. Counterfactual Information Gain (CIG) compares the gold-answer log-likelihood under a frozen reference model with and without the user response, providing per-turn credit that guides how to ask. The Asymmetric Ambiguity Bonus assigns a signed reward at the terminal token based on the gold ambiguity label, guiding when to ask. Across three clarification benchmarks spanning table, passage, and open-domain QA, CIGAsk-7B outperforms the strongest external baseline despite using a smaller backbone. It also transfers across datasets without per-dataset tuning while preserving single-turn QA performance on out-of-distribution benchmarks.

发表机构

  • The Chinese University of Hong Kong(香港中文大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑