arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.14808cs.AIcs.CLcs.LG

大型语言模型(LLMs)知道该问什么以及何时提问吗?评估多轮信息获取能力

Do LLMs Know What to Ask and When? Evaluating Multi-Turn Information Seeking

  • Harvard University(哈佛大学)
  • Google DeepMind(谷歌DeepMind)

机构由 AI 辅助整理,请以论文原文为准。

Yepeng Huang, Jiawen Zhang, Michelle Dai, Xiaorui Su, Shanghua Gao, Zi Wang, Marinka Zitnik

AI总结:

该研究构建MT-InfoSeek评估套件,从提问内容、时机及信息影响三方面评估LLMs的多轮信息获取能力,发现其存在低估信息需求等缺陷,且该能力未被现有评估覆盖。

AI中文摘要:

当用户的问题表述不明确时,一个有能力的模型应当意识到自身的上下文信息不足,识别出缺失的信息,针对该信息进行提问,并且仅在获取该信息后才给出唯一确定的答案。我们将多轮信息获取形式化为求解k个欠定约束满足问题,其中k是确定目标答案所需的联合变量数量,因此k可用于衡量信息缺失的程度。我们将该形式化方法实例化为MT-InfoSeek,这是一个受控评估套件,包含5251个问题和9006个任务实例,涵盖数学、逻辑、生物学、医学和通用知识领域。我们从三个维度评估模型:提问内容、提问时机以及获取的信息对最终答案的影响。随着欠定程度k的增加,模型在各领域的性能均出现下降。模型能够识别出需要额外信息,但会低估所需信息的数量;在k=2的逻辑问题中,模型低估信息缺失程度的频率约为高估的4倍。此外,模型无法识别出最小充分查询集合,当提供真实的k值时,性能仅略有提升,且常常在获取足够信息前就停止。在具有有序依赖关系的任务中,即使模型最终获取了所有必要信息,错误的查询顺序也会降低最终准确率。我们通过最终充分性直接衡量信息获取情况,最终充分性记录的是获取的信息是否能够独立确定目标答案,而非依赖答案生成过程。这种分离分析揭示了仅通过最终准确率无法捕捉到的模型差异,表明多轮信息获取能力与答案生成能力是不同的,且当前的LLM评估并未对该能力进行测量。

英文摘要:

When a user question is underspecified, a capable model should recognize that its context is insufficient, identify the missing information, ask for it, and respond only once that information determines a unique answer. We formalize multi-turn information seeking as solving a k-underspecified constraint satisfaction problem, where k is the number of variables jointly required to determine the target and therefore measures the degree of missing information. We instantiate the formulation in MT-InfoSeek, a controlled evaluation suite of 5,251 problems and 9,006 task instances spanning mathematics, logic, biology, medicine, and general knowledge. We evaluate models along three axes: what they ask, when they ask it, and how the acquired information affects the final answer. Performance degrades across models and domains as underspecification increases. Models recognize that additional information is needed but underestimate how much, and in logical problems at k = 2 they under-predict the degree of missing information about four times as often as they over-predict it. They also fail to identify a minimal sufficient set of queries, improve only marginally when given the true k, and often stop before acquiring sufficient information. In tasks with ordered dependencies, an incorrect query order reduces final accuracy even when the model eventually acquires all necessary information. We measure information seeking directly through final sufficiency, which records whether the acquired information determines the target independent of answer generation. This separation shows differences between models that final accuracy alone does not capture, and indicates that the ability to seek information over multiple turns is distinct from the ability to generate answers and is not measured by current LLM evaluations.

↑