发表机构
Northeastern University; Inclusion AI, Ant Group; University of Maryland, College Park(东北大学; 蚂蚁集团 Inclusion AI; 马里兰大学帕克分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对终端智能体自我验证可信度低的问题,提出诊断框架量化其验证与修复行为,并设计SCVD蒸馏方法,在TerminalBench2.1上显著提升Pass@1,同时避免分布外退化。
AI 中文摘要
终端智能体在与命令行环境交互解决任务时,依赖自我验证来评估和纠正其解决方案。然而,这种自我验证的可信度仍未被充分理解。为系统性地探究这一问题,我们引入了一个诊断框架,该框架识别每条轨迹中第一个完整的解决方案,判断其是否客观正确,并利用这一真实标签来量化智能体后续的验证和恢复行为。将该框架应用于TerminalBench2.1上的十个终端智能体,我们发现,在形成完整候选方案后,验证几乎普遍发生,但仅有61.43%的错误候选方案被检测到,且仅有49.36%的已检测错误被成功修复。这些结果表明,自我验证的主要弱点不在于启动验证,而在于错误检测与修复。基于这些发现,我们提出了学生条件验证蒸馏(SCVD),该方法先让学生生成候选解决方案,并从相同的交互上下文中蒸馏出更强教师的后续验证与恢复过程。在三个Qwen3.5骨干模型上,SCVD在TerminalBench2.1上将Pass@1相对于对应基础模型提升了9.74至16.85个百分点,相对于标准全轨迹蒸馏提升了4.49至8.61个百分点,同时避免了全轨迹蒸馏在SWE-bench Verified上显著的分布外退化。
英文摘要
Terminal agents rely on self-verification to assess and correct their solutions as they solve tasks through interaction with command-line environments. Yet how trustworthy such self-verification is remains poorly understood. To investigate this question systematically, we introduce a diagnostic framework that identifies the first complete solution in each trajectory, determines whether it is objectively correct, and uses this ground truth to quantify the agent's subsequent verification and recovery behavior. Applying it to ten terminal agents on TerminalBench2.1, we find that verification is nearly universal after a complete candidate is formed, yet only 61.43\% of incorrect candidates are detected and only 49.36\% of detected errors are successfully repaired. These results show that the main weakness in self-verification lies not in initiating verification, but in detecting and repairing errors. Motivated by these findings, we propose Student-Conditioned Verification Distillation (SCVD), which lets the student first produce a candidate solution and distills a stronger teacher's subsequent verification and recovery from the same interaction context. Across three Qwen3.5 backbones, SCVD improves \textsc{Pass@1} on TerminalBench2.1 by 9.74--16.85 percentage points over the corresponding base models and by 4.49--8.61 points over the standard full-trajectory distillation, while avoiding the pronounced out-of-distribution degradation of full-trajectory distillation on SWE-bench Verified.