发表机构
MBZUAI; BKAI Research Center, Hanoi University of Science and Technology; INSAIT, Sofia University "St. Kliment Ohridski"(穆罕默德·本·扎耶德人工智能大学; 河内科技大学BKAI研究中心; 索非亚大学“圣克莱门特·奥赫里德斯基”INSAIT)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究对话式大语言模型助手在模糊用户请求下的澄清决策问题,核心方法是引入RegretBench基准测试,贡献为揭示模型澄清的有效性和效率,表明有效澄清需在正确时间问对问题并适时停止。
AI 中文摘要
模糊的用户请求使澄清成为对话式大语言模型助手的序列决策问题:它们必须决定是否提问、问什么、何时停止以及何时回答。我们引入了RegretBench,这是一个多轮基准测试,将澄清评估为策略行为而非孤立的问题质量。RegretBench提供了模糊性的隐藏意图表述,支持基于语义状态跟踪的自由形式交互,并引入了基于遗憾的目标,以衡量模型相对于参考澄清策略损失的价值。在开放域问答和产品推荐场景上的实验表明,仅最终成功是不够的,因为具有相似准确率的模型在效率、对用户行为的鲁棒性和停止决策方面可能有很大差异。通过联合测量意图解析、交互成本、无效澄清和遗憾,RegretBench揭示了模型是否进行了有效和高效的澄清。我们的结果表明,有效的澄清需要的不仅仅是合理的问题:模型必须在正确的时间提出正确的问题,并在用户意图明确后停止。
英文摘要
Ambiguous user requests make clarification a sequential decision problem for conversational LLM assistants: they must decide whether to ask, what to ask, when to stop, and when to answer. We introduce RegretBench, a multi-turn benchmark that evaluates clarification as policy behavior rather than isolated question quality. RegretBench provides a hidden-intent formulation of ambiguity, supports free-form interaction grounded in semantic-state tracking, and introduces a regret-based objective that measures how much value a model loses relative to a reference clarification policy. Experiments on open-domain QA and product recommendation scenarios show that final success alone is insufficient, as models with similar accuracy can differ substantially in efficiency, robustness to user behaviors, and stopping decisions. By jointly measuring intent resolution, interaction cost, ineffective clarification, and regret, RegretBench reveals whether models clarify usefully and efficiently. Our results show that effective clarification requires more than plausible questions: models must ask the right question at the right time and stop once the user's intended meaning is clear.