arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.10444cs.CLcs.AI

从推理深度到推理广度:评估大型语言模型中的多点关联推理

From Reasoning Depth to Reasoning Breadth: Evaluating Multi-Point Associative Reasoning in Large Language Models

  • Beijing University of Posts and Telecommunications(北京邮电大学)
  • Peking University(北京大学)
  • Kuaishou Technology(快手科技)

机构由 AI 辅助整理,请以论文原文为准。

Si'an Xie, Jiaxun Liu, Biao Yang, Wei Yuan, Fan Yang, Tingting Gao, Ming Wu

AI总结:

该研究构建了中英双语多点关联推理基准MPAR-Bench,发现大型语言模型的推理深度未自动带来稳健的推理广度,当前基准未充分覆盖推理广度。

AI中文摘要:

大型语言模型(LLMs)在需要越来越长且复杂推理链的推理任务上已取得显著进展,这一进展主要体现为推理深度。与之互补且相对未被研究的能力是推理广度:并行探索多个语义方向并将所得线索整合为一个连贯答案。我们推出MPAR-Bench,这是一个通过多点关联推理分离推理广度的中英双语基准。受合作游戏《Just One》启发,每个条目要求模型从若干独立生成的、语义多样的线索中恢复隐藏目标。我们使用多智能体线索生成流水线、基于嵌入的多样性过滤以及人工验证构建了1000个条目;仅答案空间来自公共词表,而每个线索集均为从头生成。除精确匹配准确率外,我们还使用准确率、ANLS、嵌入相似度、推理轨迹验证以及四种扰动(线索掩码、顺序打乱、干扰项注入和多步线索)评估模型。在被评估模型中,扰动使英文准确率降低9-18个百分点,中文准确率降低5-12个百分点。思维模式提升了基准设定的准确率,尤其在英文场景中,但并未始终如一地降低对扰动的敏感性。案例级分析还显示,扩展推理可推翻最初正确的假设。这些结果表明,更强的推理深度并不自动带来稳健的推理广度,且当前基准在很大程度上未覆盖推理广度。

英文摘要:

Large language models (LLMs) have made substantial progress on reasoning tasks that require increasingly long and complex inferential chains. This progress primarily reflects reasoning depth. A complementary and comparatively unexamined capability is reasoning breadth: exploring multiple semantic directions in parallel and integrating the resulting clues into one coherent answer. We introduce MPAR-Bench, a bilingual English-Chinese benchmark that isolates reasoning breadth through multi-point associative reasoning. Inspired by the cooperative game Just One, each item asks a model to recover a hidden target from several independently generated, semantically diverse clues. We construct 1,000 items using a multi-agent clue-generation pipeline, embedding-based diversity filtering, and human verification. Only the answer space is drawn from public word lists, whereas every clue set is generated from scratch. Beyond exact-match accuracy, we evaluate models using accuracy, ANLS, embedding similarity, reasoning-trace verification, and four perturbations: clue masking, order shuffling, distractor injection, and multi-step clues. Across evaluated models, perturbations reduce accuracy by 9-18 percentage points in English and 5-12 percentage points in Chinese. Thinking mode improves standard-setting accuracy, especially in English, but does not consistently reduce sensitivity to perturbations. Case-level analysis also shows that extended reasoning can overturn an initially correct hypothesis. These results indicate that greater reasoning depth does not automatically confer robust reasoning breadth, and that reasoning breadth remains largely uncovered by current benchmarks.

↑