Chronos: 学习推理链的时序动态以实现测试时扩展
Chronos: Learning Temporal Dynamics of Reasoning Chains for Test-Time Scaling
- University of Science and Technology of China(中国科学技术大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
Chronos通过学习推理链的时序动态,提升大语言模型在测试时的推理性能,实现显著的性能提升。
AI中文摘要:
测试时扩展(TTS)已作为一种有效范式,用于提升大语言模型(LLMs)的推理性能。然而,现有方法——尤其是多数投票和启发式token级评分——将推理轨迹或token视为平等,因此容易受到轨迹质量的显著变化和局部逻辑故障的影响。在本工作中,我们引入Chronos,一种轻量级且即插即用的时序推理评分器,将每个轨迹视为时间序列。具体而言,Chronos学习捕捉token概率的轨迹特征,并据此分配质量分数,同时采用加权投票机制。在域内和域外基准上的广泛评估表明,Chronos在各种模型上均能带来显著的提升,且计算开销极低。值得注意的是,Chronos@128在使用Qwen3-4B-Thinking-2507在HMMT25上相对于Pass@1和Maj@128的相对改进分别为34.21%和22.70%,凸显了其有效性。
英文摘要:
Test-Time Scaling (TTS) has emerged as an effective paradigm for improving the reasoning performance of large language models (LLMs). However, existing methods -- most notably majority voting and heuristic token-level scoring -- treat reasoning traces or tokens equally, thereby being susceptible to substantial variations in trajectory quality and localized logical failures. In this work, we introduce \textbf{Chronos}, a lightweight and plug-and-play chronological reasoning scorer that models each trajectory as a time series. Specifically, Chronos learns to capture trajectory features of token probabilities, assigns quality scores accordingly, and employs a weighted voting mechanism. Extensive evaluations on both in-domain and out-of-domain benchmarks demonstrate that Chronos consistently delivers substantial gains across a variety of models, with negligible computational overhead. Notably, Chronos@128 achieves relative improvements of 34.21\% over Pass@1 and 22.70\% over Maj@128 on HMMT25 using Qwen3-4B-Thinking-2507, highlighting its effectiveness.