发表机构
The Hong Kong Polytechnic University(香港理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本综述通过梳理348项研究,区分LLM智能体中组件替换的任务层面收益与局部决策质量贡献,提出八项报告条目,强调仅凭在线执行或指标增益不足以确立归因。
AI 中文摘要
背景。语言模型智能体中的组件替换会改变执行轨迹,可能影响后续观察、资源使用和恢复机会。评估其在任务层面的收益以及局部决策质量的贡献,需要不同的证据。方法。本关键范围综述梳理了348项研究,并审查了90条比较记录:其中88条来自40项纳入研究,2条来自补充研究。八个目的性选择的案例围绕替换决策、执行条件、测量可比性、对照和剩余解释构建了综合框架。结果。在报告局部决策指标的222项研究中,142项同时报告了测量的任务终点,49项报告了代理指标。这些计数识别出同时报告两种测量类型的研究,但并未确立这些测量来自匹配比较。结果监控器报告了包级完成增益,但其归因于检测器质量的程度有限;首块选择报告了针对离线代理终点评估的局部改进;携带证据的终止报告了更少的过早无依据终止和完成非劣效性,但未确立完成优越性。跨案例分析确定了三种候选机制,涉及恢复与中断、干预时机和下游使用。归因和部署取决于比较对照、标签定义和控制器可用的信息。结论。本综述区分了组件替换的任务层面收益与局部决策质量的贡献,并推导出八项针对具体主张的报告条目。仅凭在线执行或局部与任务指标的同时增益,并不能单独确立更好的局部决策解释了任务层面的增益。
英文摘要
Background. A component replacement in a language-model agent changes an execution trajectory, potentially altering later observations, resource use, and recovery opportunities. Different evidence is needed to assess its task-level benefit and the contribution of local decision quality. Methods. This critical scoping review maps 348 studies and examines 90 comparison records: 88 from 40 included studies and two from supplementary studies. Eight purposively selected cases structure the synthesis around the replaced decision, executed conditions, measurement comparability, controls, and remaining explanations. Results. Of 222 studies reporting local decision metrics, 142 also report measured task endpoints and 49 report proxies. These counts identify studies that report both types of measurement, without establishing that the measurements come from matched comparisons. Outcome Monitors reports a package-level completion gain whose attribution to detector quality remains limited; First-chunk selection reports a local improvement assessed against an offline proxy endpoint; Evidence-Carrying Termination reports fewer premature unsupported terminations and completion non-inferiority, without establishing completion superiority. Cross-case analysis identifies three candidate mechanisms involving recovery and disruption, intervention timing, and downstream use. Attribution and deployment depend on the comparison controls, label definitions, and information available to the controller. Conclusions. The review distinguishes the task-level benefit of a component replacement from the contribution of local decision quality and derives eight claim-specific reporting items. Neither online execution nor simultaneous gains in local and task metrics alone establish that better local decisions explain the task-level gain.
Comments36 pages, 3 figures. The authors contributed equally