arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.27182cs.LG

TraceBench:用于时间序列根因归因的LLM智能体受控评估框架

TraceBench: Controlled Evaluation of LLM Agents for Time-Series Root-Cause Attribution

Tommaso Bendinelli, Artur Dox, Christian Holz

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对LLM智能体时间序列根因归因性能缺乏受控评估的问题,推出TraceBench框架,评估了四个LLM智能体的表现并揭示其分析规律,相关资源已公开。

中文摘要 AI 辅助

大型语言模型(LLM)智能体正越来越多地应用于从现实系统收集的时间序列观测数据的异常检测和根因分析,但它们在这些任务上的性能尚未在受控条件下得到系统评估。我们推出TraceBench,这是一个基于模拟的框架,用于生成受控的根因归因任务。在每个生成的任务中,智能体接收模拟物理动力系统产生的时间序列观测数据,必须确定模拟过程中是否改变了某个系统参数,若改变了则确定是哪一个参数。我们使用TraceBench从三个可解释的机械系统生成任务,并在受控实验条件下系统评估四个LLM智能体,获得了这些智能体如何分析动力系统时间序列观测数据的新见解。我们的结果表明,智能体从领域上下文获益显著,且主要通过数值控制台输出而非可视化来探索数据;我们还发现,当智能体需要生成Python脚本以将每个时间序列样本映射到预测的根因标签时,其表现通常比直接提交预测时更差。我们在网站(this http URL)上发布了我们的数据集、智能体轨迹、实验结果和排行榜。

英文摘要

LLM agents are increasingly applied to anomaly detection and root-cause analysis in time-series observations collected from real-world systems; however, their performance on these tasks has not been systematically evaluated under controlled conditions. We introduce TraceBench, a simulation-based framework for generating controlled root-cause attribution tasks. In each generated task, an agent receives time-series observations produced by simulating a physical dynamical system and must determine whether a system parameter was altered during the simulation and, if so, which one. Using TraceBench, we generate tasks from three interpretable mechanical systems and systematically evaluate four LLM agents across controlled experimental conditions, yielding new insights into how these agents analyze time-series observations from dynamical systems. Our results show that agents benefit substantially from domain context and explore data primarily through numerical console output rather than visualizations. We also find that agents generally perform worse when required to produce a Python script that maps each time-series sample to a predicted root-cause label than when they submit predictions directly. We release our datasets, agent trajectories, experimental results, and a leaderboard on our website, tracebench.github.io.

发表机构

  • ETH Zürich(苏黎世联邦理工学院)
  • CSEM SA(CSEM公司)

机构由 AI 辅助整理,请以论文原文为准。

↑