arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

智能体错误数据集:用于失败分析与错误感知后训练的50,000对错误诊断数据

Agent Error Dataset: Scaling 50,000 Error--Diagnosis Pairs for Failure Analysis and Error-Aware Post-Training

Kunlun Zhu, Xuyan Ye, Yibo Li, Cheng Qian, Beibin Li, Heng Ji

arXiv 2609.40111首次发表:更新:

发表机构

Apodex(Apodex)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出智能体错误数据集(AED),含50,228对错误诊断数据,并设计五阶段AET流程,通过错误诊断与修正显著提升验证器通过率及模型微调一致性。

AI 中文摘要

一次不成功的LLM智能体运行所包含的信息远多于其最终奖励:智能体可获得的观测、它选择的动作以及环境的响应。要复用这些经验进行学习,需要识别出需要修正的决策,并测试一个具体的替代方案。我们引入了智能体错误数据集(AED),该数据集包含来自33个环境、19个工具框架族和23个策略模型(均为文本型智能体系统)中9,961个源任务的50,228对错误诊断数据。我们保留了源轨迹和执行元数据,以支持跨设置的失败分析和重新诊断,而无需重复原始运行。我们的五阶段智能体错误到训练(AET)流程收集自然失败,生成诊断和提议修正,并对照记录的证据进行核查。在支持重放的情况下,我们在匹配的执行设置下,将修正与来自同一检查点的原始动作重试进行比较。然后,我们为诊断和演员恢复构建独立的训练视图。在3,062对匹配的重放对中,首次提议的修正将验证器通过率从18.4%提升至51.1%,提升了32.7个百分点。使用单独冻结的诊断发布,在1,656个源任务上进行全诊断微调,将Qwen3-8B与内部教师标签的精确步骤一致性(在943例保留集上,三个种子的平均值)从47.2%提升至63.6%。该比较中最强的提示参考得分54.7%,并且在四个递增的训练集规模下,平均一致性均有所提升。在单种子比较的演员训练方案中,仅动作修复训练在WebShop-lite上比仅成功训练高出6.67个百分点。

英文摘要

An unsuccessful LLM agent rollout contains more information than its final reward: the observations available to the agent, the actions it chose, and the environment's responses. Reusing this experience for learning requires identifying a decision to revise and testing a concrete alternative. We introduce the Agent Error Dataset (AED), comprising 50,228 error-diagnosis pairs from 9,961 source tasks across 33 environments, 19 harness families, and 23 policy models in text-based agent systems. We retain source traces and execution metadata to support cross-setting failure analysis and re-diagnosis without repeating the original rollout. Our five-stage Agentic Error-to-Training (AET) pipeline collects natural failures, generates diagnoses and proposed corrections, and checks them against recorded evidence. Where replay is supported, we compare corrections with original-action retries from the same checkpoint under matched execution settings. We then construct separate training views for diagnosis and actor recovery. Across 3,062 matched replay pairs, first-proposal corrections raise verifier pass rates from 18.4% to 51.1%, a gain of 32.7 percentage points. Using a separately frozen diagnosis release, full-diagnosis fine-tuning on 1,656 source tasks raises Qwen3-8B's exact-step agreement with internal teacher labels from 47.2% to 63.6%, averaged over three seeds on a 943-case holdout. The strongest prompted reference in this comparison scores 54.7%, and mean agreement improves at each of four increasing training-set sizes. In a single-seed comparison of actor-training recipes, action-only repair training scores 6.67 percentage points higher on WebShop-lite than success-only training.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑