arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2602.01208cs.CL

Chronos: 学习推理链的时序动态以实现测试时扩展

Chronos: Learning Temporal Dynamics of Reasoning Chains for Test-Time Scaling

  • University of Science and Technology of China(中国科学技术大学)

机构由 AI 辅助整理,请以论文原文为准。

Kai Zhang, Jiayi Liao, Chengpeng Li, Ziyuan Xie, Sihang Li, Xiang Wang

更新

AI总结:

Chronos通过学习推理链的时序动态,提升大语言模型在测试时的推理性能,实现显著的性能提升。

AI中文摘要:

测试时扩展(TTS)已作为一种有效范式,用于提升大语言模型(LLMs)的推理性能。然而,现有方法——尤其是多数投票和启发式token级评分——将推理轨迹或token视为平等,因此容易受到轨迹质量的显著变化和局部逻辑故障的影响。在本工作中,我们引入Chronos,一种轻量级且即插即用的时序推理评分器,将每个轨迹视为时间序列。具体而言,Chronos学习捕捉token概率的轨迹特征,并据此分配质量分数,同时采用加权投票机制。在域内和域外基准上的广泛评估表明,Chronos在各种模型上均能带来显著的提升,且计算开销极低。值得注意的是,Chronos@128在使用Qwen3-4B-Thinking-2507在HMMT25上相对于Pass@1和Maj@128的相对改进分别为34.21%和22.70%,凸显了其有效性。

英文摘要:

Test-Time Scaling (TTS) has emerged as an effective paradigm for improving the reasoning performance of large language models (LLMs). However, existing methods -- most notably majority voting and heuristic token-level scoring -- treat reasoning traces or tokens equally, thereby being susceptible to substantial variations in trajectory quality and localized logical failures. In this work, we introduce \textbf{Chronos}, a lightweight and plug-and-play chronological reasoning scorer that models each trajectory as a time series. Specifically, Chronos learns to capture trajectory features of token probabilities, assigns quality scores accordingly, and employs a weighted voting mechanism. Extensive evaluations on both in-domain and out-of-domain benchmarks demonstrate that Chronos consistently delivers substantial gains across a variety of models, with negligible computational overhead. Notably, Chronos@128 achieves relative improvements of 34.21\% over Pass@1 and 22.70\% over Maj@128 on HMMT25 using Qwen3-4B-Thinking-2507, highlighting its effectiveness.

↑