AI 中文总结
本文提出任务融合模型ReasonCast,通过微调LLM实现时间序列预测与可解释文本推理的联合生成,在预测准确性上优于LLM和TS模型,还构建了联合评估基准ReasonTS-Bench。
AI 中文摘要
大多数时间序列(TS)模型专门针对单一任务,要么用于理解(即返回关于时间序列的文本答案),要么用于生成(即返回数值预测)。直到最近,才出现了统一模型,开始在单一架构中处理这两项任务。然而,即使是这些模型,也会将两项输出作为任务分离的路径生成,无法在单一连贯响应中同时预测序列并解释预测的由来。本文提出一种任务融合模型,可联合生成1)预测(生成)和2)自解释(理解),从而在单一响应中整合1)数值时间序列预测和2)可解释文本推理。为实现对该能力的系统研究,本文同时提供基准和方法,共同解决两项任务。基准ReasonTS-Bench识别时间序列的五种基本模式,可对两项任务进行联合评估。本文的微调方法ReasonCast可对任意大语言模型(LLM)进行微调,使其联合执行两项任务,生成的模型可在单次自回归过程中同时生成推理链和预测。大量实验表明,ReasonCast在预测准确性上优于LLM和TS模型,同时生成可验证的因果推理。代码可从该https URL获取。
英文摘要
Most time series (TS) models are specialized for a single task, either understanding (i.e., returning text answers about a TS) or generation (i.e., returning a numeric forecast). Only recently have unified models begun to handle the two within a single architecture. Even these models, however, produce the two outputs as task-separated paths and cannot predict a series and explain why that prediction arises within a single coherent response. In this paper, we argue for a task-fused model that jointly produces 1) prediction (generation) and 2) selfexplanation (understanding), thereby integrating 1) numerical TS forecasting and 2) interpretable text reasoning within a single response. To enable the systematic study of this capability, we present both a benchmark and a recipe that jointly address the two tasks. The benchmark, ReasonTS-Bench, identifies five fundamental patterns underlying TS and enables the joint evaluation of both tasks. ReasonCast, our recipe for finetuning any LLM to perform both tasks jointly, yields a model that generates a reasoning chain and a forecast together in a single autoregressive pass. Extensive experiments show that ReasonCast outperforms both LLMs and TS models on prediction accuracy while producing verifiable, causal reasoning. Code is available at: https://github.com/seunghan96/reasoncast.