arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AERA:用于高效测试时推理的自适应证据残差分配

AERA: Adaptive Evidence Residual Allocation for Efficient Test-Time Reasoning

Ziming Wang, Ivor Tsang, Hangwei Qian

arXiv 2608.27964首次发表:更新:

发表机构

National University of Singapore; Agency for Science, Technology and Research (A*STAR); Nanyang Technological University(新加坡国立大学; 新加坡科技研究局; 南洋理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对测试时推理的计算资源浪费问题,提出自适应证据残差分配(AERA)控制器,在GSM8K等数据集上减少95.99%完成令牌数的同时保持较高准确率,证明需估计计算未来价值而非依赖当前置信度。

AI 中文摘要

测试时缩放通过生成额外候选解提升语言模型推理能力,但为每个问题分配相同推理预算会造成计算资源浪费。现有自适应停止方法通常依赖置信度、一致性或答案稳定性,隐含假设当前证据越强则无需进一步计算,本文证明该假设可能失效:检查点级别的正确性呈非单调变化,可观测证据可能在答案崩溃前增强或恢复前减弱。针对该不匹配问题,本文提出自适应证据残差分配(Adaptive Evidence Residual Allocation,AERA),这是一种序列控制器,可从检查点可观测证据中学习额外计算是否可能恢复更优答案。AERA通过答案分布、时间、重求解、语义及计算特征刻画累积响应前缀,并反复决定是停止还是分配下一个响应块。未来检查点的正确性仅用于构建离线监督,推理时控制器无法获取。在GSM8K和GPQA Diamond数据集上,AERA可识别问题特定的残差机会,同时大幅减少推理计算量。在对300个未接触过的GSM8K问题进行的冻结阈值增量生成评估中,AERA达到92.61%的准确率,而使用128个响应的基准方法准确率为93.01%,同时将完成令牌数减少95.99%。这些结果表明,自适应推理应估计计算的未来价值,而非将当前置信度等同于正确性。

英文摘要

Test-time scaling improves language-model reasoning by generating additional candidate solutions, but allocating the same inference budget to every problem is computationally wasteful. Existing adaptive stopping methods commonly rely on confidence, agreement, or answer stability, implicitly assuming that stronger current evidence indicates that further computation is unnecessary. We show that this assumption can fail: checkpoint-level correctness evolves non-monotonically, and observable evidence may strengthen before an answer collapses or weaken before it recovers. Motivated by this mismatch, we introduce Adaptive Evidence Residual Allocation (AERA), a sequential controller that learns whether additional computation is likely to recover a better answer from checkpoint-observable evidence. AERA characterizes cumulative response prefixes using answer-distribution, temporal, re-solving, semantic, and compute features, and repeatedly decides whether to stop or allocate the next response block. Future checkpoint correctness is used only to construct offline supervision and is never available to the controller at inference time. Across GSM8K and GPQA Diamond, AERA identifies question-specific residual opportunities while substantially reducing inference computation. In a frozen-threshold incremental-generation evaluation on 300 untouched GSM8K questions, AERA achieves 92.61% accuracy versus 93.01% with 128 responses while reducing completion tokens by 95.99%. These results suggest that adaptive reasoning should estimate the future value of computation rather than equating present confidence with correctness.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑