面向基于SLM的智能体任务-工具意图匹配
Toward SLM-based agentic task-tool intent matching
- Cisco Systems(思科系统)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究探讨小型语言模型作为任务-工具相关性分类器,用于验证智能体工具调用是否符合任务意图,并通过提示优化、监督微调和强化学习提升性能。
AI中文摘要:
配备工具的人工智能智能体使用工具调用来访问数据并对外部系统采取行动。智能体系统的横向增长增加了这些交互的数量,并进一步推动了对自动化、每次调用监督的需求,这种监督可以在低延迟和/或本地部署下运行。传统的授权方案可以确定智能体是否被允许调用工具,但无法评估智能体的底层认知,具体来说,即工具选择是否代表了一个逻辑上、相关的步骤,以满足任务的意图。因此,一个被允许的调用仍可能偏离任务的意图:一个恶意智能体可能会偏离调用,或促使其他智能体进行一系列不符合任务意图的调用。因此,每次调用都需要被验证。在本研究中,我们探讨了小型语言模型(SLMs)在此目的上的适用性:一个SLM充当任务-工具相关性分类器,独立评估每个选定的工具相对于分配的任务,并为下游执行返回相关性信号。配备了一个新颖的数据集,其中包含多工具任务,所需工具跨越不同的模型上下文协议(MCP)服务器,我们使用了提示优化、监督微调和通过GRPO的强化学习来优化和专门化SLMs。
英文摘要:
Tool-equipped AI agents use tool calls to access data and act on external systems. Horizontal growth of agentic systems increases the number of these interactions, and further motivates the need for automated, per-call oversight that can operate at low latency and/or on-prem. Conventional authorization schemes can determine whether an agent is allowed to invoke a tool, but cannot assess the agent's underlying cognition, specifically, whether the tool selection represents a logical, relevant step toward satisfying the intent of the task or not. Consequently, an allowed call may still deviate from the task's intent: a rogue agent might deviate the calls or nudge other agents to make a combination of calls that would not align with the intent of the task. Therefore, every call needs to be verified. In this study we investigate the applicability of Small Language Models (SLMs) to this purpose: an SLM functions as a task-tool relevance classifier that evaluates every selected tool independently against the assigned task and returns a relevance signal for downstream enforcement. Equipped with a novel dataset with multi-tool tasks whose required tools span distinct Model Context Protocol (MCP) servers, we used prompt-optimization, supervised fine-tuning, and reinforcement learning through GRPO to optimize and specialize SLMs.