arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.24268cs.CL

准确性掩盖了语言模型的失败方式:在匹配输出预算下测量失败状态

Accuracy Hides How Language Models Fail: Measuring Failure States Under Matched Output Budgets

Zongyou Yang, Yinghan Hou

AI总结:

研究指出语言模型基准测试中准确性分数掩盖问题,引入两层评估框架分离执行证据与正确性。通过实验表明不同模型在匹配输出限制下执行组合不同,验证策略会影响准确性估计,强调评估应同时报告执行状态、覆盖范围和评分者来源。

AI中文摘要:

语言模型基准测试将两个不同的测量问题合并为一个单一的准确性分数:一个响应是否达到可评估状态,以及其答案是否被判定正确。我们引入了一个两层评估框架,将与评分者无关的执行证据(包括终止、答案暴露、可解析性和完成长度)与依赖评分者的正确性分开。在MATH和ARC-Challenge上对五个固定的Qwen和DeepSeek配置的2550个输出进行测试,匹配2048-token限制会产生截然不同的执行组合。在相同的300个DeepSeek MATH问题-模型对中,在8192个token时未观察到缺少最终长度终止的情况。一项覆盖审核的目标验证研究进一步表明,候选选择和聚合策略会显著改变比较准确性估计。这些结果表明,准确性将执行案例组合与验证策略混为一谈。因此,对测试时方法的评估应在报告准确性的同时,报告干预前的执行状态、验证覆盖范围和评分者来源。

英文摘要:

Language-model benchmarks collapse two distinct measurement questions into a single accuracy score: whether a response reached an evaluable state, and whether its answer was judged correct. We introduce a two-layer evaluation framework that separates scorer-independent execution evidence, including termination, answer exposure, parseability, and completion length, from scorer-dependent correctness. Across 2,550 outputs from five fixed Qwen and DeepSeek configurations on MATH and ARC-Challenge, matched 2,048-token limits produce sharply different execution mixtures: 49 of 450 Qwen MATH outputs terminate without a final answer, compared with 5 of 300 DeepSeek MATH outputs and none of the 750 ARC outputs. Among the same 300 DeepSeek MATH question-model pairs, no missing-final length termination is observed at 8,192 tokens. A coverage-audited targeted verification study further shows that candidate-selection and aggregation policies can substantially alter comparative accuracy estimates. These results demonstrate that accuracy conflates execution case mix with verification policy. Evaluations of test-time methods should therefore report pre-intervention execution states, verification coverage, and scorer provenance alongside accuracy.

补充信息

↑