arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

技能遵循:评估支持检索的大语言模型智能体的实际技能使用情况

Skill Following: Evaluating Actual Skill Use in Retrieval-Enabled LLM Agents

Seonghyeon Cho, Chanjun Park

arXiv 2609.00549首次发表:更新:

发表机构

Soongsil University(崇实大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对支持检索的LLM智能体,提出RAE指标衡量实际技能使用,发现聚合评估存在悖论,部分模型全局检索有提升但实际检索任务性能受损,RAE可准确评估工具使用的真实效果。

AI 中文摘要

大语言模型(LLM)智能体越来越依赖外部技能,但标准评估方法无法明确这些技能的检索是否真正有帮助。聚合指标通常会对比使用检索和未使用检索的任务,这会引入严重的选择偏差,且无法分离出技能使用的真实效果。为衡量这种实际使用能力——我们将其形式化为技能遵循(Skill Following, SF)——我们提出了检索触发的实际使用效应(Retrieval-Invoked Actual-Use Effect, RAE)。RAE 仅在智能体主动检索技能的任务上,计算匹配的启用技能与未启用技能执行之间的相同任务结果差异。我们在编码和数学领域评估了17个LLM,发现了一个显著的评估悖论:模型经常表现出正的聚合检索提升,但RAE为负。在MBPP+上,多个看似在全局范围内受益的模型,在检索发生的精确任务上实际损害了自身性能。这些发现表明,聚合平均值会产生关于工具使用熟练度的误导性错觉,而RAE直接衡量检索到答案的流程是否真正挽救了更多结果而非造成损害。

英文摘要

Large Language Model (LLM) agents increasingly rely on external skills, yet standard evaluations obscure whether retrieving these skills actually helps. Aggregate metrics often compare retrieved versus non-retrieved tasks, introducing severe selection bias and failing to isolate the true effect of skill use. To measure this actual-use capability-which we formalize as Skill Following (SF)-we introduce the Retrieval-Invoked Actual-Use Effect (RAE). RAE computes the same-task outcome difference between matched skill-enabled and skill-disabled executions, conditioned exclusively on tasks where the agent actively retrieved a skill. Evaluating 17 LLMs across coding and mathematical domains, we uncover a stark evaluation paradox: models frequently show positive aggregate retrieval lift but negative RAE. On MBPP+, multiple models that appear to benefit system-wide actually harm their own performance on the exact tasks where retrieval occurred. These findings demonstrate that aggregate averages can create a misleading illusion of tool-use proficiency, whereas RAE directly measures whether the retrieval-to-answer pipeline genuinely rescues more outcomes than it harms.

CommentsAccepted to Findings of EMNLP 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑