发表机构
AWS AI Labs(亚马逊云科技人工智能实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对现有基于置信度的无验证器测试时扩展(VF-TTS)方法在复杂任务上失效的问题,提出一致性(consilience)框架,通过评估置信度时间不对称性提升LLM推理性能,在数学和代码生成任务上优于基线。
AI 中文摘要
测试时扩展(Test-time scaling)通常会使用外部验证器,例如编码场景中的编译器和测试用例、机器人应用中的训练值函数,来获得高质量的 rollout。无验证器的测试时扩展(Verifier-free test-time scaling,简称 VF-TTS)作为增强大型语言模型(Large Language Model,简称 LLM)推理能力的机制,正受到广泛关注,主要原因是在许多实际应用中无法获取此类高质量验证器。在现有的 VF-TTS 方法中,仅通过置信度计算和排序 rollout 的基于置信度的 VF-TTS 方法尤为有前景。这类方法的样本评估开销几乎为零,且对内部模型状态的访问需求极小,因此在模型和任务间具有极高的灵活性。本文中,我们证明了现有基于置信度的 VF-TTS 方法存在一个关键局限:这类方法在复杂任务上会灾难性失效。我们观察到一个非常有趣的现象:均匀高置信度往往意味着探索不足,偏向于置信度高的错误答案。为解决这一问题,我们的核心见解是:稳健的认知搜索需要特定的置信度轨迹模式——这类方法在初始阶段执行探索性分支,表现为初始置信度较低,最终收敛到高置信度的解决方案。为实现这一见解,我们引入了一种名为一致性(consilience)的新型选择框架,该框架明确评估推理中置信度的时间不对称性。我们通过一种组合指标来实现这一点,该指标主动惩罚高初始置信度,同时严格要求最终置信度确定性。涵盖研究生数学问题和自由形式代码生成的大量实验表明,一致性方法显著优于现有基线,验证了我们对完成置信度的新视角。
英文摘要
Test-time scaling often uses an external verifier, such as compilers and test cases in coding or trained value functions in robotics applications, to obtain high-quality rollouts. Verifier-free test-time scaling (or VF-TTS) is gaining extensive attention as a mechanism to enhance Large Language Model (LLM) reasoning, primarily because we do not have access to such high-quality verifiers in many real-world applications. Among existing VF-TTS methods, confidence-based VF-TTS methods, which compute and rank rollouts solely by confidence, are particularly promising. Such methods introduce near-zero overhead for sample evaluation and require minimal access to internal model states, making the methods highly flexible across models and tasks. In this paper, we demonstrate a critical limitation of existing confidence-based VF-TTS methods by showing that such methods catastrophically break down on complex tasks. We observe a very interesting phenomenon: uniformly high confidence frequently indicates a failure to explore, favoring confidently wrong answers. To address this, our core insight is that robust cognitive search requires a specific confidence trajectory pattern: such methods perform exploratory branching at the beginning, as manifested by low initial confidence, and converge to a high final confidence solution. To implement this insight, we introduce consilience, a novel selection framework that explicitly evaluates the temporal asymmetry of confidence in reasoning. We operationalize this via a combinatorial metric that actively penalizes high initial confidence while strictly demanding final certainty. Extensive experiments covering both graduate-level mathematics problems and free-form code generation demonstrate that consilience effectively outperforms existing baselines, validating our novel perspective on completion confidence.