模型还是控制?一种以交互为中心的智能体故障定位分类法
Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent Failures
浏览论文内容
中文总结 AI 辅助
该研究提出以交互为中心的智能体故障分类法,将故障定位到交互及责任组件,适用于多类智能体架构,经评估具有良好可重复性与结构一致性。
中文摘要 AI 辅助
现有评估常将智能体故障简化为系统级结果,模糊了故障来源及可改进智能体系统的干预措施,引发修复分配问题:同一可见故障因来源不同,可能需模型后训练、控制(harness)工程、环境重新设计或基准修复。智能体行为源于模型、控制(harness)、用户、工具、记忆与环境的交互,结果级标签往往不足以支撑改进。多数故障分类法因特定于基准且缺乏共享结构,难以解决该问题。我们提出一种以交互为中心的分类法,将故障定位到其起源的交互中并识别责任组件,通过将41种故障模式分配给两个组件间的边及指示修复归属的故障侧来组织。该分类法具有可操作性:模型侧故障确定后训练目标,控制(harness)侧故障指向脚手架及工具集成修复,环境或评估器故障揭示需重新设计的评估条件。该架构适用于从编码助手到长程个人助手、多智能体系统的各类智能体架构。我们通过公开基准、模型系统卡片、已发表报告及记录的智能体轨迹中的实例验证该分类法,并使用独立推理智能体作为评判者评估其可重复性。在四个前沿模型中,最强评判者与人工类别标签的Cohen's κ=0.76,表明这些类别捕捉到共享结构而非标注者特定偏好。
英文摘要
Existing evaluations often reduce agent failures to system-level outcomes, obscuring where the fault originated and which intervention would improve the agent system. This creates a repair-assignment problem: the same visible failure may call for model post-training, harness engineering, environment redesign, or benchmark repair depending on its source. Because agent behavior emerges from interactions among models, harnesses, users, tools, memory, and environments, outcome-level labels are often insufficient for improvement. Most failure taxonomies do little to resolve this problem because they are benchmark-specific and lack a shared structure. We introduce an interaction-centric taxonomy that localizes failures to the interactions in which they originate and identifies the responsible component. It organizes 41 failure modes by assigning each to an edge between two components and a fault side indicating where the repair belongs. This makes the taxonomy actionable: model-side failures identify targets for post-training, harness-side failures point to scaffolding and tool-integration fixes, and environment or grader failures reveal evaluation conditions requiring redesign. The schema applies across agent architectures, from coding assistants to long-horizon personal assistants and multi-agent systems. We ground the taxonomy in worked examples from public benchmarks, model system cards, published reports, and logged agent trajectories, and evaluate its reproducibility using independent reasoning agents as judges. Across four frontier models, the strongest judge reaches Cohen's $κ=0.76$ against human category labels, suggesting that the categories capture shared structure rather than annotator-specific preferences.
发表机构
- Scale
机构由 AI 辅助整理,请以论文原文为准。