发表机构
Soongsil University(崇实大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对支持检索的LLM智能体,提出RAE指标衡量实际技能使用,发现聚合评估存在悖论,部分模型全局检索有提升但实际检索任务性能受损,RAE可准确评估工具使用的真实效果。
AI 中文摘要
大语言模型(LLM)智能体越来越依赖外部技能,但标准评估方法无法明确这些技能的检索是否真正有帮助。聚合指标通常会对比使用检索和未使用检索的任务,这会引入严重的选择偏差,且无法分离出技能使用的真实效果。为衡量这种实际使用能力——我们将其形式化为技能遵循(Skill Following, SF)——我们提出了检索触发的实际使用效应(Retrieval-Invoked Actual-Use Effect, RAE)。RAE 仅在智能体主动检索技能的任务上,计算匹配的启用技能与未启用技能执行之间的相同任务结果差异。我们在编码和数学领域评估了17个LLM,发现了一个显著的评估悖论:模型经常表现出正的聚合检索提升,但RAE为负。在MBPP+上,多个看似在全局范围内受益的模型,在检索发生的精确任务上实际损害了自身性能。这些发现表明,聚合平均值会产生关于工具使用熟练度的误导性错觉,而RAE直接衡量检索到答案的流程是否真正挽救了更多结果而非造成损害。
英文摘要
Large Language Model (LLM) agents increasingly rely on external skills, yet standard evaluations obscure whether retrieving these skills actually helps. Aggregate metrics often compare retrieved versus non-retrieved tasks, introducing severe selection bias and failing to isolate the true effect of skill use. To measure this actual-use capability-which we formalize as Skill Following (SF)-we introduce the Retrieval-Invoked Actual-Use Effect (RAE). RAE computes the same-task outcome difference between matched skill-enabled and skill-disabled executions, conditioned exclusively on tasks where the agent actively retrieved a skill. Evaluating 17 LLMs across coding and mathematical domains, we uncover a stark evaluation paradox: models frequently show positive aggregate retrieval lift but negative RAE. On MBPP+, multiple models that appear to benefit system-wide actually harm their own performance on the exact tasks where retrieval occurred. These findings demonstrate that aggregate averages can create a misleading illusion of tool-use proficiency, whereas RAE directly measures whether the retrieval-to-answer pipeline genuinely rescues more outcomes than it harms.
CommentsAccepted to Findings of EMNLP 2026