给定增量的埃塔:用边际工具效用定义大语言模型工具效率
Eta Given Delta: Defining LLM Tool Efficiency With Marginal Tool Utility
浏览论文内容
中文总结 AI 辅助
研究提出用于评估LLM智能体轨迹中工具调用率的工具效率及边际工具效用指标,用LLM-as-a-Judge确定边际工具效用符号,区别于间接测效率的 prior work,直接测效率,为LLM评估研究及相关工程做贡献。
中文摘要 AI 辅助
本文介绍了工具效率,这是一种用于评估大语言模型(LLM)智能体轨迹中有用工具调用率的新定量指标。为确保工具效率定义明确,我们还引入了边际工具效用,这是一个针对每个工具调用定义的新定量指标,表明工具是否有用,或者在不影响准确性且提高工具效率的情况下是否可以从工具套件中安全移除。在本文中,我们使用“大语言模型即评判器”来确定轨迹中每个工具调用的边际工具效用的符号。虽然之前已经做了很多工作来开发改进大语言模型工具使用的技术,并设计以准确性为代理间接测量效率的评估方法,但我们的工作集中在事后轨迹分析中通过本文提出的定量指标直接测量效率。我们希望这项工作能为大语言模型评估研究前沿做出贡献,作为未来基准设计和智能体利用工程(特别是关于创建精简工具套件)的跳板,这些工程针对与准确性互补但不同的指标进行优化。
英文摘要
This paper introduces tool efficiency, a new quantitative metric to evaluate the rate of useful tool calls in an LLM agent trajectory. To ensure that tool efficiency is well-defined, we also introduce marginal tool utility, a new quantitative metric defined per tool call indicating whether a tool is useful or whether it can be safely removed from the tool suite without affecting accuracy while increasing tool efficiency; in this paper, we determine the sign of marginal tool utility for each tool call in a trajectory using LLM-as-a-Judge. While much prior work has been done to develop techniques that improve tool use by LLMs and design evaluation methods measuring efficiency indirectly using accuracy as a proxy, our work is centered on measuring efficiency directly via the quantitative metric proposed in this paper in post hoc trajectory analyses. It is our intention that this work contributes to the frontier of LLM evaluation research as a springboard for future benchmark designs and agent harness engineering (specifically with regards to creating lean tool suites) that optimize for metrics that complement but are distinct from accuracy.
发表机构
- Foam(泡沫)
机构由 AI 辅助整理,请以论文原文为准。