arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LineupRL:通过描述到序列识别实现时间序列描述的可验证强化学习

LineupRL: Verifiable Reinforcement Learning for Time Series Captioning via Caption-to-Series Identification

Haochen Zhang, Laura Yao, Zachary Plotkin, Gengwei Zhang, Tianlong Chen

arXiv 2610.01800首次发表:更新:

AI 中文总结

提出LineupRL,一种基于描述到序列识别的可验证奖励强化学习方法,用于时间序列描述,在多个基准上超越SFT和RL基线,并以1/24参数超越72B VLM。

AI 中文摘要

时间序列描述是时间序列理解的基础步骤,也可作为信号与自然语言之间的桥梁。监督微调(SFT)依赖于更大模型的描述,无法超越其质量。强化学习(RL)可以超越,但其奖励是为其他模态和其他任务设计的,难以迁移到时间序列领域的开放式生成任务。我们通过提出LineupRL来解决这一问题,这是一种具有可验证奖励的强化学习(RLVR)流程,其奖励为描述到序列的识别。奖励模型是一个冻结的大型语言模型(LLM)验证器,它读取生成的描述和候选时间序列的原始值(而非图表),并必须从多个干扰项中选出所描述的时间序列。匹配对验证器的要求远低于编写问题或评判描述,因此现成的LLM可以提供奖励。在两个描述基准测试中,以及在预测和重建任务(预测器仅看到描述)中,LineupRL在每项指标上都优于SFT和RL基线。由LineupRL训练的3B视觉语言模型(VLM)在参数仅为1/24的情况下,也优于SFT基线所蒸馏的72B VLM。我们的案例研究表明,LineupRL能够抵抗奖励黑客攻击,并且其训练的描述器既能追踪趋势,也能在关键点标注数值。

英文摘要

Time series captioning is a fundamental step in time series understanding and can also serve as the bridge between signal and natural language. Supervised fine-tuning (SFT) relies on a larger model's captions and cannot exceed their quality. Reinforcement learning (RL) can, but its rewards were designed for other modalities and other tasks, and they transfer poorly to open-ended generation in the time series domain. We address this by proposing LineupRL, a reinforcement learning with verifiable rewards (RLVR) pipeline whose reward is caption-to-series identification. The reward model is a frozen large language model (LLM) verifier that reads the generated caption and the candidate time series as raw values, never the chart, and must pick the described time series from multiple distractors. Matching is a far lighter demand on the verifier than writing questions or judging a caption, so an off-the-shelf LLM can supply the reward. Across two captioning benchmarks, and on forecasting and reconstruction where the predictor sees only the caption, LineupRL outperforms SFT and RL baselines on every metric. The 3B vision language model (VLM) trained by LineupRL also outperforms, at 1/24 of the parameters, the 72B VLM whose captions the SFT baseline is distilled from. Our case study shows that LineupRL resists reward hacking, and that the captioner it trains both traces the trend and names the values at key points.

Comments28 pages, 4 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑