发表机构
Carnegie Mellon University; Tsinghua University(卡内基梅隆大学; 清华大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
ControlScope研究LLM智能体应修订工作流的程度,比较继续生成、编辑参数和替换未完成工作流三种方式,通过嵌套权限和多项任务评估,揭示修复访问与实际选择及执行的关系。
AI 中文摘要
一个语言模型智能体应该修订正在运行的工作流中的多少内容?ControlScope比较了继续生成代码、编辑下一个工具调用的数据参数,以及从相同的公共执行状态替换未完成的工作流。嵌套权限将可用的修复与智能体选择的动作区分开来。我们在文件系统任务、ALFWorld和AppWorld上评估了一次性和重复性审查。在20个文件系统任务中,每个任务两个源程序和三次推理审查者抽取下,FULL完成了15-16个任务,而KEEP完成了13个;在四次快速抽取下,FULL完成了10-13个,而KEEP完成了13个。新的学生记录确认复现了一次批量读取修复。ALFWorld快速面板在52个场景的87个任务上产生KEEP/ARG/FULL分数85/86/87,在四个场景的134个任务上产生134/134/127;在87个任务队列上的推理也产生85/86/87,但审查成本显著。AppWorld V1官方测试面板包含来自195个场景模板的585个任务实例,显示净差异很小。冻结重放暴露了在两次失败的文件组织运行中,由后续修订中断的、可行的智能体编写的替换。离线源轨迹中点比较显示,后续审查完成了一次不充分的修复。五次调用保护在20次新鲜源运行中节省了19.4%的日志模型输出,但损失了一次成功。一个仅参数快捷方式表明,更广泛的采样策略可能忽略两种操作集中都存在的更便宜的、成功的编辑。这些结果将修复访问与实际选择和后续执行联系起来。
英文摘要
How much of a running workflow should a language model agent revise? ControlScope compares continuing generated code, editing the next tool call's data arguments, and replacing the unfinished workflow from the same public execution state. The nested permissions separate available repairs from the actions an agent selects. We evaluate one-time and repeated reviews across filesystem tasks, ALFWorld, and AppWorld. Across two source programs per task and three reasoning-reviewer draws on 20 filesystem tasks, FULL completes 15-16 tasks versus 13 for KEEP; across four fast draws it completes 10-13 versus 13. Fresh student-record confirmation reproduces a batch-read repair. ALFWorld fast panels yield KEEP/ARG/FULL scores of 85/86/87 on 87 tasks across 52 scenes and 134/134/127 on 134 tasks across four scenes; reasoning on the 87-task cohort also yields 85/86/87 with substantial review cost. An AppWorld V1 official-test panel of 585 task instances from 195 scenario templates shows small net differences. Frozen replays expose viable agent-written replacements interrupted by later revision in two failed file-organization runs. An offline source-trajectory midpoint comparison shows later reviews completing an insufficient repair. Five-call protection saves 19.4% of logged model output and loses one success across 20 fresh source runs. An argument-only shortcut shows that the broader sampled policy can overlook a cheaper successful edit available in both operation sets. These outcomes tie repair access to actual choices and subsequent execution.