arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

REALHOP:通过行为审计重新思考多跳推理评估

REALHOP: Rethinking Multi-Hop Reasoning Evaluation via Behavioral Auditing

Jiawen Tao, Xiaokun Yuan, Yaoming Li, Chenxu Liu, Mengzhou Wu, Tong Yang, Maxm Pan

arXiv 2609.36984首次发表:更新:

发表机构

Tencent; Peking University(腾讯; 北京大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

REALHOP提出诊断-构建-验证框架,通过行为必要性率审计多跳推理,揭示正确答案不保证依赖预期证据,并显著提升基准测试中的证据依赖性。

AI 中文摘要

复杂问题通常需要多跳推理,通过中间步骤连接分布在多个来源或长上下文远距离区域中的事实。基准测试通常围绕预定义的推理链构建问题来评估这种能力,将正确答案视为使用了预期组合的证据。然而,仅凭答案正确性无法确定成功是否依赖于每个预期步骤相关的证据:模型可能转而依赖记忆关联、更短路径或部分证据。我们使用行为必要性率(BNR)来检验这种依赖性,该指标衡量在最初正确的实例上,有针对性地移除证据阻止答案恢复的频率。在五个现有基准上,专家组平均BNR范围从16.6%到48.9%,揭示了注释结构与观察到的依赖性之间的显著差距。在这一诊断的指导下,我们提出了REALHOP,一个诊断-构建-验证框架,该框架重新绑定实体、分解选定的关系、添加完整的竞争路径,并将证据放置在可追踪的位置。在冻结之前进行结构和语义检查;随后进行行为干预。在790对MuSiQue问题上,REALHOP将专家组平均BNR从27.4%提高到94.4%,同时保持较高的完整准确率。它还在REALHOP-FRAMES和REALHOP-LONGBENCH上产生了高BNR。在216个长上下文问题上,16个模型上的匹配多项选择差异从13.9分增长到59.2分,并在重复开放式评估中持续存在。这些结果共同表明,概念上连贯的链条和正确的最终答案本身并不能确立多跳推理。因此,验证成功依赖于每个预期跳数与衡量答案准确性本身一样,是多跳评估的基本要素。

英文摘要

Complex questions often require multi-hop reasoning that connects facts distributed across sources or distant regions of a long context through intermediate steps. Benchmarks commonly evaluate this ability with questions built around predefined reasoning chains, treating a correct answer as evidence that the intended composition was used. Yet answer correctness alone leaves open whether success depends on the evidence associated with each intended step: models may instead rely on memorized associations, shorter paths, or partial evidence. We examine this dependence using the Behavioral Necessity Rate (BNR), which measures how often targeted evidence removal prevents answer recovery on initially correct instances. Across five existing benchmarks, panel-mean BNR ranges from 16.6% to 48.9%, exposing a substantial gap between annotated structure and observed dependence. Guided by this diagnosis, we introduce REALHOP, a diagnose-construct-verify framework that rebinds entities, factorizes selected relations, adds complete competing paths, and places evidence at traceable locations. Structural and semantic checks precede freezing; behavioral interventions follow. On 790 paired MuSiQue questions, REALHOP raises panel-mean BNR from 27.4% to 94.4% while retaining high Full accuracy. It also yields high BNR on REALHOP-FRAMES and REALHOP-LONGBENCH. On 216 long-context questions, the matched multiple-choice spread across 16 models grows from 13.9 to 59.2 points and persists under repeated open-ended evaluation. Together, these results show that a conceptually coherent chain and a correct final answer do not by themselves establish multi-hop reasoning. Verifying that success depends on every intended hop is therefore as fundamental to multi-hop evaluation as measuring answer accuracy itself.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑