Mid-Harness:在模型与执行框架之间扩展终端智能体的动作
Mid-Harness: Scaling Actions Between Model and Harness for Terminal Agents
AI总结:
Mid-Harness通过在模型与执行框架间采样并验证候选动作,提升终端智能体的动作可靠性,实验表明动作扩展能有效提高成功率并降低令牌成本。
AI中文摘要:
终端智能体通过随机模型生成来执行动作,然而生成有用动作的能力并不能确保其可靠执行。一个糟糕的命令(例如,错误的软件包安装)可能以阻碍后续进展的方式改变环境,即使模型本可以生成更好的替代方案。我们研究了在模型-执行框架边界分配测试时计算是否能提高动作可靠性和轨迹成功率,以及什么使这种分配有效。为研究这些问题,我们引入了Mid-Harness,它在转发一个候选动作执行之前对其进行采样和验证,同时保持生成器和执行框架不变。使用TMAX-9B生成器,在弱验证下增加动作采样收益甚微,而一个能力强的验证器则能利用同一生成器产生的有用替代方案。在TerminalBench-Lite上,GPT-5.6 Sol验证器将Pass@1从基础智能体的50.00%提高到采样8个动作时的68.03%。当同一TMAX-9B模型作为验证器时,成对验证在所评估的验证机制中表现最佳。将较强验证器的响应蒸馏到TMAX-9B中,进一步提高了Pass@1,同时保持动作生成器不变。使用TMAX-9B在TerminalBench-Lite上,结合动作和轨迹扩展比单独生成更多轨迹在更低的估计令牌成本下达到更高的成功率。Mid-Harness还在其他模型、基准和执行框架上提升了性能。这些发现将动作扩展确定为终端智能体中测试时计算扩展的一个有前景的目标。
英文摘要:
Terminal agents act through stochastic model generations, yet the ability to generate a useful action does not ensure its reliable execution. A poor command (e.g., wrong package install) can change the environment in ways that hinder subsequent progress, even when the model could generate a better alternative. We investigate whether allocating test-time compute at the model-harness boundary can improve action reliability and trajectory success, and what makes this allocation effective. To study these questions, we introduce Mid-Harness, which samples and verifies candidate actions before forwarding one for execution, while keeping the generator and harness unchanged. With a TMAX-9B generator, more action sampling yields little benefit under weak verification, whereas a capable verifier can exploit useful alternatives from the same generator. On TerminalBench-Lite, a GPT-5.6 Sol verifier raises Pass@1 from 50.00% for the base agent to 68.03% with 8 sampled actions. When the same TMAX-9B model serves as the verifier, pairwise verification performs best among the evaluated verification mechanisms. Distilling responses from the stronger verifier into TMAX-9B further improves Pass@1, while leaving the action generator unchanged. With TMAX-9B on TerminalBench-Lite, combining action and trajectory scaling reaches higher success at lower estimated token cost than generating more trajectories alone. Mid-Harness also improves performance across additional models, benchmarks, and harnesses. These findings identify action scaling as a promising target for test-time compute scaling in terminal agents.