发表机构
University of Illinois at Chicago; Visa Research; University of Florida; The Ohio State University; Case Western Reserve University(伊利诺伊大学芝加哥分校; Visa研究院; 佛罗里达大学; 俄亥俄州立大学; 凯斯西储大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
TimeEvo通过聚类智能体失败、规划测量并合成证据工具,经配对准入门实现自我进化,在十个时间序列问答任务和三个骨干模型上从空库开始提升准确率。
AI 中文摘要
时间序列智能体通过调用外部工具来回答分析性问题,而它们携带哪些工具由人们在智能体运行前决定。然而,我们发现了这种设置中的两个失败。人-智能体工具错配:一个包含21个专家精选工具的工具库在某些任务上有帮助,但在其他任务上有害,在我们测试的每个骨干模型下都会降低异常检测的准确率。无声伤害:一轮通用的自我修订改变了147个答案,并破坏了其中56个,而最终得分变动不到一分。两者都源于同一个差距:工具是否有帮助是在运行时逐问题决定的,而工具是预先提供并仅通过单一平均值来评判的。为解决这一问题,我们提出TimeEvo,它将智能体诊断出的失败聚类为能力差距,为每个差距规划测量方法,合成仅基于证据的工具来填补这些差距,并通过一个配对准入门来接纳候选工具库。在十个时间序列问答任务和三个骨干模型上的实验表明,TimeEvo从空库开始,在每个任务和每个骨干模型上都提高了准确率,并且在一个廉价模型上成长起来的工具库在安装到更强模型上时仍然有效。代码可在https://github.com/Muyiiiii/TimeEvo获取。
英文摘要
Time series agents answer analytical questions by calling external tools, and which tools they carry is decided by people before the agent runs. However, we identify two failures in this setup. Human-Agent Tool Misalignment: a library of 21 expert-curated tools helps on some tasks and hurts on others, dropping anomaly accuracy under every backbone we test. Silent Harm: one round of generic self-revision changes 147 answers and breaks 56 of them, while the final score moves by less than a point. Both follow from the same gap: whether a tool helps is decided question by question at runtime, while tools are supplied in advance and judged by a single average. To address this, we propose TimeEvo, which clusters an agent's diagnosed failures into capability gaps, plans a measurement for each, synthesizes evidence-only tools that fill them, and admits the candidate library only through a paired admission gate. Experiments on ten time series QA tasks and three backbones show that TimeEvo, starting from an empty library, improves accuracy on every task and every backbone, and that a library grown on a cheap model still gains when it is installed into stronger ones. Code is available at https://github.com/Muyiiiii/TimeEvo.