arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

可达性并非可实现性:追溯大语言模型基准测试增益的来源

Reachability Is Not Realization: Tracing the Sources of LLM Benchmark Gains

Yanchao Li, Wanhao Liu, Jiaqing Xie, Ben Gao, Yanbo Wang, Tianfan Fu, Yuqiang Li

arXiv 2608.03219首次发表:更新:

AI 中文总结

该研究针对LLM基准测试增益,建立问题级审计方法,发现可实现性与可达性常不同步,提出能力扩展声明需报告匹配条件下的可实现性能与可达性。

AI 中文摘要

基准测试增益常被视为大语言模型(LLM)能力更强的证据,但相同的增益可能反映模型行为的不同变化:模型可能得出新答案,也可能得出已在可达范围内的答案,而 aggregate 分数无法逐问题区分这些变化。我们在固定预算、温度和答案格式下建立了问题级审计方法:当默认部署流程产生正确答案时,该问题被视为可实现(realized);当指定探测方法在固定预算内找到该答案时,该问题被视为可达(reachable)。我们首先测试推理时的层路由是否能扩大可达性:在匹配的预算下,随机路由在全部 43 种模型与任务设置中均达到或超过结构化搜索的效果;无答案的程序几乎无法保留该增益,而这需要访问正确答案才能实现。接着我们探究可达答案有时为何未出现:在涵盖 0.5B 至 31B 的 6 种案例中,抑制一个已识别的 MLP 模块可修复预定义失败集的 68%至 92%。随后我们测试训练是否能通过扩大可达性缩小差距:在 6 项匹配评估中的 5 项里,部署性能上升,而可达上限保持平稳或下降;对于 DAPO,部署分数上升 14.7 个百分点,而可达上限下降 13.3 个百分点。因此,在我们审计的设置中,可实现性与可达性并不总是同步变化,能力扩展的声明应报告匹配评估条件下的可实现性能与可达性。代码可在该 https URL 获取

英文摘要

Benchmark gains are often treated as evidence of greater LLM capability. Yet the same gain can reflect different changes in model behavior. A model may reach new answers, or produce answers that were already within reach. Aggregate scores do not distinguish these changes question by question. We establish a question-level audit under fixed budgets, temperatures, and answer formats. A question is realized when the default deployment procedure produces the correct answer. A question is reachable when a specified probe finds that answer within a fixed budget. We first test whether inference-time layer routing can expand reachability. Under a matched budget, random routes match or exceed structured search in all 43 model and task settings. Answer-blind procedures retain almost none of this gain, which instead requires access to the correct answer. We then ask why reachable answers sometimes fail to appear. Across six cases spanning 0.5B to 31B, silencing one identified MLP block repairs 68 to 92 percent of a predefined failure set. We next test whether training closes the gap by expanding reachability. In five of six matched evaluations, deployed performance rises while the reachable ceiling remains flat or falls. For DAPO, the deployed score rises by 14.7 points while the reachable ceiling falls by 13.3 points. Across the settings we audit, realization and reachability therefore do not always change together. Claims of capability expansion should report both realized performance and reachability under matched evaluation conditions. Code is available at https://github.com/LiZaiyuan0619/reachability-not-realization

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑