arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TimeRLM:递归语言模型可实现长时序数据中的精准异常定位

TimeRLM: Recursive Language Models Enable Precise Anomaly Localization in Long-Context Time-Series

Nicolas Zumarraga, Lorenzo Steno, Ning Wang, Max Rosenblattl, Thomas Kaar, Maxwell A. Xu, Kevin O'Sullivan, Markus Kreft, Elgar Fleisch, Paul Schmiedmayer, Patrick Langer, Robert Jakob

arXiv 2608.03391首次发表:更新:

AI 中文总结

提出TimeRLM(基于递归语言模型的时序异常定位框架),结合强化学习后训练后,在合成基准AnomalyXL及真实世界数据上,均优于现有时序语言模型,可实现长时序数据的精准异常定位。

AI 中文摘要

在临床护理、工业运营、金融服务及物流等领域的监控应用中,长时序数据的精准异常定位是一项关键任务,简短的异常证据可能隐藏在长跨度的高频数据中。时序语言模型(TSLMs)能够输入时序数据并以自然语言表述异常发现;然而,近期基准测试显示,长上下文下的检索性能出现下降,这与文本、视觉及音频领域的失效模式类似。在文本领域,递归语言模型(RLMs)通过将上下文保留在大型语言模型(LLM)外部,允许模型通过代码查询上下文,可恢复大部分丢失的性能。我们提出TimeRLM,一种针对时序数据的RLM框架,通过代码和视觉能力对信号进行顺序处理。我们进一步引入AnomalyXL,这是一个合成长时序异常定位基准,包含通过编程注入的、需要精准检索的异常。我们设置了五个不同的任务类别和两个变体:AnomalyXL-MCQ与AnomalyXL-Localize。TimeRLM在AnomalyXL-Localize的5项任务中的4项上,性能优于所有被评估的TSLMs及单遍基线模型,定位任务的交并比(IoU)达到0.682,带证据分类任务达到0.745,而所有基线模型的对应数值最高仅为0.329和0.072。我们使用强化学习对TimeRLM进行后训练,得到的模型进一步提升了性能,且产生最终答案所需的智能体交互轮次约为其未训练基础模型的三分之一。在未见过的真实世界心电图(ECG)、睡眠数据及软件可观测性记录上,经后训练的TimeRLM保持或提升了性能,尽管仅在合成数据上训练,仍超越了TSLMs。我们的研究表明,与时序数据进行递归交互是长距离检索的有效方法。

英文摘要

Precise anomaly localization over long-context time series is a crucial task in monitoring applications across clinical care, industrial operations, financial services, and logistics, where brief evidence may hide inside long spans of high-frequency data. Time-Series Language Models (TSLMs) are able to ingest time series data and verbalize findings on anomalies in natural language; however, recent benchmarks report a decrease in retrieval performance at long contexts, mirroring failure modes in text, vision, and audio. In the text domain, Recursive Language Models (RLMs) can recover much of this lost performance by keeping context external to the large language model (LLM), allowing the model to query it through code. We present TimeRLM, an RLM formulation for time-series that sequentially manipulates the signal using code and vision capabilities. We further introduce AnomalyXL, a synthetic long-context anomaly localization benchmark with programmatically injected anomalies that require precise retrieval. We implement five different task categories and two variants: AnomalyXL-MCQ and AnomalyXL-Localize. TimeRLM outperforms every evaluated TSLM and single-pass baseline on four of the five AnomalyXL-Localize tasks, reaching 0.682 IoU on localization and 0.745 on classify-with-evidence, versus at most 0.329 and 0.072 across all baselines. We post-train TimeRLM using reinforcement learning. The resulting model further improves performance and requires approximately one-third as many agent interaction turns as its untrained base model to produce a final answer. On unseen real-world ECG, sleep and software observability recordings, the post-trained TimeRLM retains or improves performance, surpassing TSLMs despite being trained exclusively on synthetic data. Our findings suggest recursive interaction with time-series is an effective approach for long-horizon retrieval.

CommentsOpen source code and datasets: https://github.com/OpenTSLM/TimeRLM

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑