技能实际做了什么?LLM智能体中工具与技能使用的估计目标与评估效度:一项批判性综述
What Does a Skill Actually Do? Estimands and Evaluation Validity for Tool and Skill Use in LLM Agents: A Critical Review
- The Hong Kong Polytechnic University(香港理工大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本综述批判性分析LLM智能体工具与技能评估的估计目标与效度,从六维度区分不同设计,指出任务配对等常见做法无法识别调用效应,并提出报告清单以提升评估可比性。
AI中文摘要:
大型语言模型(LLM)智能体日益依赖外部工具和可复用技能,这些工具和技能在运行时从包含数千条目的库中选择。关于检索器、路由器或技能库“改进”智能体的报告,可能指检索召回率、启用库带来的成功变化、仅限于触发任务的配对对比,或近似预算约束下的增益。本批判性综述探讨了每种设计比较了什么以及基于何种假设。基于估计目标(estimand)的智能体评估方法,我们从六个维度描述工具和技能设计:处理对比、目标人群、结果、预算约束、汇总度量及识别假设。十三项核心实证研究锚定了证据综合,辅以相关方法论工作和更广泛文献的设计层面解读。我们的贡献在于明确区分了某些原始作者已通过触发条件配对分解、分析性反例和跨研究比较所承认的差异。基于任务配对本身并不能识别调用效应;配对增益和回归计数衡量的是协议特定的不一致性,而非预期结果恶化的任务占比;模块部署的总效应回答的是与预算约束效率不同的问题。我们比较了精选技能提供与检索器替换、触发子集与全任务结果、观察到的成本降低与预算约束比较。一份报告清单和实例将上述区分与研究可报告的信息联系起来。本综述未开展新实验;实证结果来自引用的研究,数值玩具示例为分析性说明。
英文摘要:
Reported improvements from tools and reusable skills in large language model agents refer to different comparisons. This critical narrative review examines what these evaluations estimate and which conclusions their designs support. The review checks the roles of one hundred cited papers and extracts focal evaluation designs in detail from thirty-five studies. Targeted readings of thirty-five additional published or accepted studies broaden coverage of tool creation, memory, interactive benchmarks, reliability, and risk. Designs are characterized by treatment contrast, target population, outcome, budget constraint, summary measure, and identification assumptions. Analytic decompositions and counterexamples show that pairing runs on the same task does not itself identify an invocation effect when evaluation conditions on a trigger within the treated run. Paired gain and regression counts describe discordance under the coupling protocol rather than the share of tasks whose expected outcomes worsen. Total effects of deploying a module answer a different question from efficiency under a common budget. Comparisons across studies distinguish curated skill provision from retriever replacement, task populations from triggered subsets, and preparation costs from marginal usage costs. Publication status and reading depth are recorded. The review provides a methodological synthesis and a reporting checklist to help align claims about tools and skills with the comparisons their evaluation designs support.