arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.13457cs.AI

TimeThink:激发时间序列大语言模型中的组合推理能力

TimeThink: Eliciting Compositional Reasoning in Timeseries Large Language Models

Sudarshan Regmi, Arvind Pillai, Yu Yvonne Wu, Yuliang Chen, Bibek Panthi, Tess Z. Griffin, Michael V. Heinz, Lisa Marsch, Nicholas C. Jacobson, Andrew Campbell

首次发表
浏览论文内容

中文总结 AI 辅助

TimeThink提出合成数据生成与可验证奖励强化学习框架,激发时间序列大语言模型的显式组合推理,仅用合成数据即在合成及真实基准上显著超越强基线。

中文摘要 AI 辅助

时间序列多模态大语言模型(TS-MLLMs)近期开始利用大语言模型(LLMs)的推理能力来执行问答任务。然而,这些模型往往无法捕捉动态的时间模式,仅提供隐式推理,缺乏在高风险应用(如医疗保健)中至关重要的底层解释。虽然基于强化学习(RL)的时间序列语言模型旨在解决这一问题,但它们通常因在狭窄的分布内数据上训练而表现不佳,并且在处理分布外的组合性问题时存在困难。为应对这些挑战,我们提出了TimeThink,一个用于激发组合时间序列推理的合成框架。核心时间序列原语(如趋势、季节性)是领域无关的,并且可以被确定性地生成。基于这一前提,TimeThink首先设计了一个合成数据生成器,产生原子性和组合性的问答对,提供带有推理轨迹的客观真实值。在此框架基础上,TimeThink采用了一种带可验证奖励的强化学习(RLVR)训练策略,鼓励显式推理。与依赖模板的方法不同,该方法使模型能够学习组合的底层逻辑,而非简单模仿轨迹。大量实验表明,仅使用合成数据训练的TimeThink在合成和真实世界基准测试中均显著优于强基线模型。

英文摘要

Timeseries multimodal large language models (TS-MLLMs) have recently begun leveraging the reasoning capabilities of large language models (LLMs) for question-answering tasks. However, these models often fail to capture dynamic temporal patterns, providing only implicit reasoning that lacks the underlying explanations critical for high-stakes applications like healthcare. While reinforcement learning (RL)-based timeseries language models aim to address this, they often fall short because they are trained on narrow, in-distribution data and struggle with out-of-distribution compositional questions. To address these challenges, we present TimeThink, a synthetic framework for eliciting compositional timeseries reasoning. Core timeseries primitives (e.g., trend, seasonality) are domain-independent and can be deterministically generated. Guided by this premise, TimeThink first designs a synthetic data generator that produces atomic and composite question-answer pairs, providing objective ground truth with reasoning traces. Building on this framework, TimeThink employs a reinforcement learning with verifiable rewards (RLVR) training strategy that encourages explicit reasoning. Unlike template-reliant methods, this approach enables the model to learn the underlying logic of composition rather than simply imitating traces. Extensive experiments show that TimeThink, trained only on synthetic data, significantly outperforms strong baselines on both synthetic and real-world benchmarks.

发表机构

  • Dartmouth College(达特茅斯学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑