调用来自模型内部:研究基于探针的大型语言模型工具调用错误检测方法
The Calls are Coming from Inside the Model: Investigating Probe-based Detection of Tool-Calling Errors in LLMs
- Pacific Northwest National Laboratory(太平洋西北国家实验室)
- National Security Agency(美国国家安全局)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究探究基于线性探针的工具调用错误检测方法,在18个工具调用LLM上验证其效能,发现探针可捕捉多种工具调用错误,还能泛化到新型错误,为LLM工具调用错误检测提供有效手段。
AI中文摘要:
已知大型语言模型(LLM)的隐藏状态包含与模型知识和行为相关的丰富信息,仅通过检查输入和输出很难提取这些信息。随着基于LLM的系统越来越多地与外部世界交互,检测工具的不正确或不当使用成为一个值得关注的领域。受此启发,我们研究使用线性探针检测不正确工具调用的有效性,在伯克利函数调用排行榜上评估了18个工具调用LLM,以测量探针的效能。总体而言,我们发现探针是捕捉各种不同工具调用错误的有效手段,包括使用参数值错误但类型正确的错误,这类错误可能不会被标准日志框架记录。成功的重要因素包括模型规模、探针层和模型后训练类型。我们还表明,探针能够泛化到新型错误,这对于实际部署至关重要。
英文摘要:
The hidden states of large language models (LLMs) are known to capture rich information relating to model knowledge and behavior that can be hard to extract from examination of input and output alone. As LLM-based systems increasingly interface with the external world, one area of concern is detecting incorrect or improper use of tools. Motivated by this, we study the effectiveness of using linear probes to detect incorrect tool-calls, measuring probe efficacy across 18 tool-calling LLMs evaluated on the Berkeley Function Calling Leaderboard. Overall, we find that probing is an effective means to catch a range of different tool-calling errors, including errors arising from using an argument that has the wrong value but the correct type, which might not be recorded by standard logging frameworks. Important factors in success include model size, probing layer, and model post-training type. We also show that probes are capable of generalizing to novel types of errors, which is critical in real world deployments.