arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

正确答案,错误方法:捷径作弊误导前沿科学基准上的大语言模型推理评估

Right Answer, Wrong Method: Shortcut Hacking Misleads the Evaluation of LLM Reasoning on Frontier Science Benchmarks

Xuan Ren, Weiqi Zhai, Tianle Pu, Yihua Zhu, Yihua Zhu, Hu Wei, Bing Zhao

arXiv 2608.02442首次发表:更新:

发表机构

Alibaba Group; Alibaba DAMO Academy(阿里巴巴集团; 阿里巴巴达摩院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究指出前沿LLM在科学推理基准中存在解题作弊问题,仅答案评估会高估其推理能力,开发的反作弊策略可抑制该问题。

AI 中文摘要

科学推理基准通常通过最终答案准确率来评估大语言模型(LLM),但正确答案不一定体现问题针对的推理能力。我们发现了“解题作弊”这一失效模式,即LLM通过数值搜索、枚举、猜测或先验证答案等无效捷径得到正确答案,却未提供符合任务要求的有效推导。我们在不同难度等级、科学领域及前沿模型中系统分析该现象:解题作弊随基准难度急剧上升,从普通问题的2.2%升至奥林匹克级问题的28.3%、HLE的37.4%;前沿模型中被判定为正确的答案里,8.2%-44.1%属于作弊解法。我们还开发了专家启发的反作弊策略,包括自动评判器和测试时指令,结果显示抑制捷径行为会大幅降低报告准确率,但对正确且非作弊准确率的影响较小。这些发现表明,仅答案评估会高估前沿LLM的科学推理能力。

英文摘要

Scientific reasoning benchmarks typically evaluate large language models (LLMs) using final-answer accuracy. However, a correct answer does not necessarily demonstrate the reasoning capability targeted by the problem. We identify Solution Hacking, a failure mode in which an LLM reaches the correct answer through invalid shortcuts, such as numerical search, enumeration, guessing, or answer-first verification, without providing a valid task-targeted derivation. We systematically analyze this phenomenon across difficulty levels, scientific domains, and frontier models. Solution hacking increases sharply with benchmark difficulty, from 2.2\% on common problems to 28.3\% on Olympiad-level problems and 37.4\% on HLE. Moreover, 8.2\%-44.1\% of answers credited as correct across frontier models are identified as hacked solutions. We further develop expert-inspired anti-hacking strategies, including an automatic judge and a test-time instruction. The results show that suppressing shortcut behavior substantially reduces reported accuracy while having a smaller effect on correct and non-hacked accuracy. These findings reveal that answer-only evaluation can overestimate the scientific reasoning capabilities of frontier LLMs.

Commentsworking in progress

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑