arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

别让我主动索要:大语言模型在用于溯因推理的主动多轮信息获取中存在缺陷

Don't Let Me Ask for It: LLMs Show Deficiencies in Active Multi-Turn Information Acquisition for Abductive Inference

Shahrukh Mohiuddin, Chalamalasetti Kranti, Sherzod Hakimov, David Schlangen

arXiv 2608.03388首次发表:更新:

AI 中文总结

该研究通过Alien Abduction游戏探究LLMs的主动多轮信息获取能力,发现其在分轮提供证据、自行选择查询时表现欠佳,存在假设验证与停止决策的缺陷。

AI 中文摘要

溯因推理需要构建能够解释观测证据的假设,并在新证据出现时对假设进行修正。虽然大语言模型(LLMs)常被评估是否能正确解决溯因推理任务,但人们对其获取证据、更新假设以及决定何时停止的过程知之甚少。我们引入了Alien Abduction游戏,这是一种交互式探测工具,用于研究不同交互模式下的上述行为。交互模式的差异在于证据是预先提供还是分轮提供,以及查询是由模型选择还是由“神谕”提供示例。在所有模型中,预先提供证据的成功率高于分轮提供;在多轮设置中,部分模型在使用完可用证据前就做出了决策,而另一些模型则在耗尽轮次预算后仍未收敛。当“神谕”提供示例时,模型的成功率高于模型自行选择查询时,尽管模型的最终假设与自行选择的证据更一致。这些发现表明,模型可能会构建出符合自行选择证据却未充分区分于其他替代假设的假设,且可能难以验证和修正自身假设,或难以确定何时停止。

英文摘要

Abductive reasoning requires forming hypotheses that explain observed evidence and revising them as new evidence becomes available. While large language models (LLMs) are often evaluated on whether they solve abductive reasoning tasks correctly, less is known about how they acquire evidence, update their hypotheses, and decide when to stop. We introduce Alien Abduction game, an interactive probe for studying these behaviours under different interaction modes. The modes vary in whether evidence is provided upfront or across turns, and whether queries are selected by the model or examples are provided by the oracle. Across models, providing evidence upfront leads to higher success rates than distributing it across turns. In multi-turn settings, some models commit before using the available evidence, while others exhaust the turn budget without converging. Models also achieve higher success rates when examples are provided by the oracle than when they select their own queries, although their final hypotheses are more consistent with the evidence they selected. These findings suggest that models may form hypotheses that fit self-selected evidence without sufficiently distinguishing them from alternatives, and may struggle to validate and refine their hypotheses or determine when to stop.

CommentsPreprint

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑