arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

回归税:剖析技能对大语言模型智能体产生帮助和伤害的原因

The Regression Tax: Decomposing Why Skills Help and Hurt LLM Agents

Darshan Tank, Baran Nama

arXiv 2607.22520首次发表:更新:

发表机构

Sentient Labs(森蒂恩实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究剖析大语言模型智能体中技能产生帮助和伤害的原因,通过对比实验区分回归和残余失败两种结果,识别出回归的三个原因,发现现有技能过度强调程序指导,指出应分解技能净效应评估,还识别出应避免的回归模式及可靠性的关键因素。

AI 中文摘要

在大语言模型智能体中添加程序技能通常通过任务成功率的平均提升来评估。然而,这一指标掩盖了一个重要成本:技能也可能使智能体表现更差。我们通过在跨越两个办公自动化基准和三个模型框架栈的近6000次运行中比较有技能和无技能的智能体来衡量这两方面。这使我们能区分两种结果:回归是指无技能时能解决但添加技能后失败的任务;残余失败是指有技能和无技能时都失败的任务。我们发现回归情况很显著,最佳性能的技能主要通过更少回归而非更多提升来超越其他技能。我们识别出回归的三个原因:技能描述渗透、基础位移和验证位移。分析持续失败揭示了相同的潜在模式。现有技能过度强调程序指导而忽视基础和验证。纠正评估工件并研究痕迹后,我们发现许多回归和持续失败可通过更好的基础和验证来恢复。程序技能应通过将其净效应分解为提升和回归来评估,而不仅仅是总体改进。我们识别出技能应避免的三种回归模式,并发现可靠性更多取决于基础和验证而非程序技能选择。

英文摘要

Adding procedural skills to an LLM agent is typically evaluated by average improvement in task success. However, this metric hides an important cost: skills can also make agents worse. We measure both sides by comparing agents with and without skills across nearly 6,000 runs spanning two office automation benchmarks and three model harness stacks. This allows us to distinguish two outcomes. A regression is a task solved without skills but failed after skills are added. A residual failure is a task that fails both with and without skills. We find that regressions are substantial enough that the best performing skills outperform others primarily by regressing less, not by gaining more. We identify three causes of regression: (i) skill description osmosis, a skill changes an agent's behavior simply by being present in context, even when it is never invoked; (ii) grounding displacement, a skill's prescribed procedure overrides how the agent interprets its inputs; and (iii) verification displacement, where the procedure suppresses checks the agent would otherwise perform on its outputs. Analysing persistent failures reveals the same underlying pattern. Existing skills overemphasize procedural guidance the stage least often responsible for failure while under supporting grounding and verification, the dominant sources of remaining errors. After correcting evaluation artifacts and studying traces, we find many regressions and persistent failures recoverable through better grounding and verification. Procedural skills should be evaluated by decomposing their net effect into gains and regressions, not by aggregate improvement alone. We identify three regression modes skills should avoid, and find that reliability depends more on grounding and verification than on procedural skill choice.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑