arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.14270cs.AI

TimeSage-EV:面向动态环境下智能体时间序列分析的实时基准测试集

TimeSage-EV: A Live Benchmark for Agentic Time Series Analysis in Evolving Environments

  • Eindhoven University of Technology(埃因霍温理工大学)
  • University of Oxford(牛津大学)
  • VulpiVox Intelligence(VulpiVox智能公司)
  • Squirrel Ai Learning(Squirrel AI学习公司)
  • Griffith University(格里菲斯大学)

机构由 AI 辅助整理,请以论文原文为准。

Qingren Yao, Yaxuan Kong, Yuqi Nie, Yichen Li, Stefan Zohren, Anna Vettoruzzo, Qingsong Wen, Ming Jin, Joaquin Vanschoren

AI总结:

针对现有时间序列基准未评估时间有效性等问题,推出动态环境智能体时间序列分析基准TimeSage-EV,含多领域真实场景,实验揭示模型性能差距及失败模式,提供研究资源。

AI中文摘要:

高风险领域的时间序列分析依赖于周期性数据发布,新观测结果可能会改变证据基础及后续结论的有效性。现有时间序列问答基准大多基于固定快照,未对时间有效性和感知截止日期的证据使用进行评估。我们推出TimeSage-EV,这是一个面向动态环境下智能体时间序列分析的实时基准测试集。它涵盖6个领域的60个真实机构场景,包含2023年2月至2026年5月期间的1485个场景-时间段问答对,覆盖月度、每周、每日及不规则发布节奏。每个时间段,大语言模型(LLM)智能体将接收时间序列数据和源报告,而 withheld 目标发布提供真实值。TimeSage-EV评估状态识别、数据汇总和展望推理。对前沿LLM智能体及TimeSage-1.0(一种具有可复用分析技能库的新型自进化智能体)的实验显示,不同模型层级间存在显著性能差距,且在时间有效性、外生上下文使用和适应性方面存在反复出现的失败。我们将TimeSage-EV作为研究资源发布,包含月度更新、代码、排行榜及失败模式分析。

英文摘要:

Time series analysis in high-stakes domains relies on recurring data releases, where new observations can alter the evidence base and the validity of later conclusions. Existing time series QA benchmarks mostly rely on fixed snapshots, leaving temporal validity and cutoff-aware evidence use unevaluated. We introduce TimeSage-EV, a live benchmark for agentic time series analysis in evolving environments. It tracks 60 real institutional scenarios across 6 domains, comprising 1,485 scenario-period QA pairs from Feb 2023 to May 2026 and spanning monthly, weekly, daily, and irregular release cadences. At each period, large language model (LLM) agents receive time series data and source reports, while the withheld target release provides ground truth. TimeSage-EV evaluates state identification, data summarization, and outlook reasoning. Experiments with frontier LLM agents and TimeSage-1.0, a novel self-evolving agent with a reusable analytical skill library, reveal significant performance gaps across model tiers and recurring failures in temporal validity, exogenous context use, and adaptation. We release TimeSage-EV as a research resource with monthly updates, code, a leaderboard, and failure-mode analyses.

↑