发表机构
University of Edinburgh; The Chinese University of Hong Kong(爱丁堡大学; 香港中文大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过零售和航空实验,对比固定计划与乱序文本,量化了智能体框架中规划指导与验证器对成功率、错误接受率和成本的影响,发现其价值取决于错误接受损失的高低。
AI 中文摘要
智能体框架提供规划指导、组织执行并检查完成情况。我们在τ²-bench中的两个零售实验和一个航空试点中研究了这些组件如何影响成功率、错误接受率和成本。主要比较将预先编写的任务特定计划(固定)与字数匹配的乱序策略文本(虚假)配对,以隔离指导内容的贡献。在265个匹配单元中,固定计划将oracle验证的成功率提高了7.17个百分点(90%任务聚类自助法区间,1.15--13.36个百分点),且收益集中在较高复杂度的任务中。一个只读终端验证器拒绝了61%的零售oracle无效情节,同时扣留了17%的正确情节,每个情节的额外成本不到一美分。哪个组件更重要取决于分配给错误接受率的损失:在低责任下,规划收益占主导;在高责任下,验证器避免的错误通过占主导——而独立验证器以一小部分成本捕获了完整规划加验证堆栈几乎所有的错误通过收益。
英文摘要
Agent harnesses supply planning guidance, organize execution, and check completion. We study how these components affect success, erroneous acceptance, and cost in two Retail experiments and an Airline pilot in $τ^2$-bench. The primary comparison pairs prewritten task-specific plans (Fixed) with shuffled policy text matched in word count (Sham), isolating the contribution of guidance content. Across 265 matched cells, Fixed improves oracle-verified success by 7.17 percentage points (90\% task-clustered bootstrap interval, 1.15--13.36 points), with gains concentrated in higher-complexity tasks. A read-only terminal verifier rejects 61\% of Retail oracle-invalid episodes while withholding 17\% of correct ones, at less than one cent of additional cost per episode. Which component matters more depends on the loss assigned to erroneous acceptance: at low liability the planning gain dominates; at high liability the verifier's avoided false passes dominate---and a standalone verifier captures nearly all the false-pass benefit of the full planning-plus-verification stack at a fraction of its cost.