奇妙的自适应分类法及其使用方法
Fantastic Adaptive Taxonomies and How to Use Them
浏览论文内容
中文总结 AI 辅助
研究智能体系统失败反馈问题,提出AdaMAST方法将轨迹转换为自适应失败分类法,该分类法能作为共享反馈接口,在智能体系统搜索、运行时及轨迹选择等方面改进智能体,提升准确率和分辨率。
中文摘要 AI 辅助
智能体系统的执行轨迹记录了其失败方式,一些在不改变模型权重的情况下改进该系统的过程(轨迹选择、提示和工作流程优化、运行时监控)会读取这些轨迹以获取反馈。然而原始轨迹并非积累反馈的良好媒介,它们冗长、特定于实例且缺乏用于反复出现的失败的稳定词汇表。我们认为智能体系统应维护一个由自身行为诱导的、关于其失败方式的明确表示,且在需要失败反馈的任何地方都可复用。AdaMAST通过将目标系统的轨迹转换为紧凑的、基于证据的失败分类法来构建这种表示:沿三个固定轴(系统级、特定角色和特定领域)组织的命名失败代码,每个名称、定义和证据模式均由轨迹诱导;无需人工编写代码,也无需人工注释轨迹。该分类法不仅是事后诊断,还是一个共享反馈接口,从三个方面改进智能体。在智能体系统搜索中,对失败候选者的分类法编码诊断在我们测试的所有五个基准上均优于自由形式的反思。在运行时,分类法反馈将SWE智能体在SWE基准验证迷你版上的分辨率从自由文本反思时的60%提高到70%,并将Claude代码作为运行时技能从64.0%提高到70.7%。在轨迹选择中,基于诱导代码构建的验证器AdaMAST-Judge在终端基准2.0上的最佳5准确率比Pass@1提高了8 - 15分。该词汇表本身紧凑(数量级压缩,保留轨迹差异)、忠实于人类(比手工制作的参考词汇表更紧密地匹配专家失败注释)且自适应(为不同领域诱导的分类法共享的代码很少)。自适应失败分类法闭合了智能体产生的轨迹与改进它们的过程之间的循环。
英文摘要
An agent system's execution traces record how it fails, and procedures that improve such a system without changing model weights (trajectory selection, prompt and workflow optimization, runtime monitoring) read these traces for feedback. Yet raw traces are a poor medium for accumulating that feedback: long, instance-specific, and lacking a stable vocabulary for recurring failures. We argue that an agent system should instead maintain an explicit representation of how it fails, induced from its own behavior and reusable wherever failure feedback is needed. AdaMAST builds this representation by converting a target system's traces into a compact, evidence-grounded failure taxonomy: named failure codes organized along three fixed axes (system-level, role-specific, and domain-specific), with every name, definition, and evidence pattern induced from the traces; no code is hand-authored, no trace human-annotated. The taxonomy is not merely a post-hoc diagnostic but a shared feedback interface, improving agents in three ways. In agent-system search, taxonomy-coded diagnoses of failed candidates outperform free-form reflection on all five benchmarks we test. At runtime, taxonomy feedback raises SWE-agent's resolution on SWE-bench Verified Mini from 60% with free-text reflection to 70%, and improves Claude Code from 64.0% to 70.7% as a runtime skill. In trajectory selection, AdaMAST-Judge, a verifier built on the induced codes, improves best-of-5 accuracy on Terminal-Bench 2.0 by 8-15 points over Pass@1. The vocabulary itself is compact (an order-of-magnitude compression that preserves trace distinctions), human-faithful (matching expert failure annotations more closely than a hand-crafted reference vocabulary), and adaptive (taxonomies induced for different domains share few codes). Adaptive failure taxonomies close the loop between the traces agents produce and the procedures that improve them.
发表机构
- University of California, Berkeley(加州大学伯克利分校)
- Apple(苹果公司)
- Bespoke Labs(Bespoke实验室)
机构由 AI 辅助整理,请以论文原文为准。