发表机构
Gaoling School of Artificial Intelligence, Renmin University of China; Kuaishou Technology(中国人民大学高瓴人工智能学院; 快手科技)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对长周期智能体监督难题,提出证据支撑行为图(EBG)方法,通过将证据分组为行为并构建关系图,帮助监控器识别重要决策并定位证据,在基准测试中显著提升效果。
AI 中文摘要
随着智能体承担长周期任务,用户从做出个体决策转向监督自主执行。然而,智能体活动的大量性以及支持证据的碎片化,使得难以确定哪些决策值得用户验证。我们研究能够识别重要决策并定位证据以帮助用户评估其影响的监控器。我们引入了AgentMonBench,一个软件工程基准,包含三个子集,覆盖两个互补维度:需求与行为之间的一致性,以及对于验证而言重要自主决策的感知。为支持这些判断,我们提出了证据支撑行为图(EBG),一种无需训练的方法,将来源关联的证据分组为行为,并将其关系组织成图。EBG呈现该图的任务导向视图,以帮助监控器在上下文中解释行为。跨八个模型的实验表明,与直接访问原始上下文相比,EBG在大多数设置中提高了决策识别和证据定位能力。进一步实验表明,EBG的证据定位增益在输入规模和超参数设置下持续存在,而实际应用则展示了其对于人类监督的实用价值。
英文摘要
As agents take on long-horizon tasks, users shift from making individual decisions to overseeing autonomous execution. Yet the volume of agent activity and the fragmentation of supporting evidence make it difficult to determine which decisions warrant user verification. We study monitors that identify consequential decisions and locate evidence to help users assess their implications. We introduce AgentMonBench, a software-engineering benchmark comprising three subsets that cover two complementary dimensions: alignment between requirements and behavior, and awareness of consequential autonomous decisions for verification. To support these judgments, we propose the Evidence-Grounded Behavior Graph (EBG), a training-free method that groups source-linked evidence into behaviors and organizes their relationships into a graph. EBG presents task-oriented views of this graph to help monitors interpret behavior in context. Experiments across eight models show that EBG improves decision identification and evidence localization in most settings compared with direct access to the original context. Further experiments show that EBG's evidence-localization gains persist across input scales and hyperparameter settings, while real-world applications illustrate its practical value for human oversight.