arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过难度控制和动态测试生成对语言推理模型进行时间推理基准测试框架

A Temporal Reasoning Benchmarking Framework for LRMs via Difficulty-controlled and Dynamic Test Generation

Shide Zhou, Kailong Wang, Ling Shi, Haoyu Wang

arXiv 2607.04784首次发表:更新:

发表机构

Huazhong University of Science and Technology; National University of Singapore; Nanyang Technological University(华中科技大学; 新加坡国立大学; 南洋理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对语言推理模型定义推理边界和确保可靠性的挑战,引入TRACE框架,将时间推理建模为约束满足问题,构建TRACEBench基准,评估模型,揭示推理问题及规模相关失败模式。

AI 中文摘要

定义大型推理模型(LRMs)的推理边界并确保其可靠性仍然是一项关键挑战。当前基准主要依赖易受数据污染的静态数据集或缺乏细粒度难度控制的合成任务。此外,基于标准结果的评估往往忽略推理过程而掩盖推理缺陷。为解决这些限制,我们引入TRACE,一个通过艾伦区间代数将时间推理建模为约束满足问题的测试框架。这种方法能够精确调节逻辑复杂性,并纳入基于跟踪的验证预言机以验证推理忠实性。使用这个框架,我们构建了TRACEBench,一个包含1200个跨难度级别的合成测试实例的广泛基准。我们使用TRACE在TRACEBench上评估八个广泛使用的LRMs。结果证实模型性能与我们的难度指标之间存在很强的负相关(皮尔逊r约为-0.96),验证了我们难度控制机制的有效性。此外,我们基于跟踪的分析揭示了推理有效性与最终答案之间的显著差异,揭示了中型模型中约28%的高虚假猜测率。此外,我们诊断了与规模相关的失败模式,从小型模型中的退化循环到高级架构中的推理爆炸。TRACE因此为基准测试LRMs的真正时间推理能力提供了一个强大的自动化平台。

英文摘要

Defining the reasoning boundaries and ensuring the reliability of Large Reasoning Models (LRMs) remains a critical challenge. Current benchmarks primarily rely on static datasets susceptible to data contamination or synthetic tasks lacking fine-grained difficulty control. Furthermore, standard outcome-based evaluations often conceal reasoning flaws by neglecting the reasoning process. To address these limitations, we introduce TRACE, a testing framework that models temporal reasoning as constraint satisfaction problems via Allen's Interval Algebra. This approach enables precise regulation of logical complexity and incorporates a Trace-Based Verification Oracle to validate reasoning faithfulness. Using this framework, we construct TRACEBench, an extensive benchmark comprising 1,200 synthesized test instances across graded difficulty levels. We employ TRACE to evaluate eight widely used LRMs on TRACEBench. The results confirm a strong negative correlation between model performance and our difficulty metric (Pearson's r approximately -0.96), validating the effectiveness of our difficulty control mechanism. Moreover, our trace-based analysis exposes significant discrepancies between reasoning validity and final answers, revealing a high spurious guessing rate of approximately 28% in mid-sized models. In addition, we diagnose scale-dependent failure modes, ranging from Degenerative Loops in small models to Reasoning Explosion in advanced architectures. TRACE thus provides a robust, automated platform for benchmarking the true temporal reasoning capabilities of LRMs.

CommentsAccepted to ISSTA 2026. 23 pages including references, 3 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑