发表机构
Sharif University of Technology(谢里夫理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对小语言模型智能体,提出不确定性感知的步骤级切换框架STEPGATE,动态升级困难步骤至强模型,在匹配云动作率下显著提升任务成功率并减少远程传输。
AI 中文摘要
小语言模型(SLM)作为本地智能体控制器具有吸引力,因为它们减少了远程推理、延迟和部署足迹,但结构化工具错误可能导致智能体步骤失败。现有的路由器通常为每个查询选择一次模型。然而,智能体暴露了顺序决策点,其难度会根据中间观察动态变化。我们提出了STEPGATE,一个不确定性感知的切换框架,它对每个本地SLM动作进行评分,并选择性地将具有挑战性的步骤升级到更强的模型。在52个任务、保留的单步骤BFCL派生测试分割上,Qwen2.5-1.5B/7B组合达到了82.7%的任务成功率,升级率为30.8%,而仅本地为67.3%,随机升级(使用33.8%的升级率)为75.4%。在另一个多轮评估中,STEPGATE仅使用30.0%的云动作就实现了69.0%的轨迹成功率和84.0%的动作成功率,而仅本地为48.0%/70.5%,随机升级为60.0%/78.2%,查询级路由为57.0%/77.1%(仅强模型在100%云动作下达到82.0%的轨迹成功率)。这些结果表明,在匹配的云动作率下,步骤级升级恢复了与更强的Qwen2.5-7B后端之间的大部分性能差距,同时远程传输更少的令牌。然而,我们的评估仅限于一个模型家族、一个更强的后端和脚本化任务。此外,测试集很小,多轮比较依赖于配对区间和统计检验,我们的风险等级作为研究注释而非正式的安全保证。
英文摘要
Small language models (SLMs) are attractive as local agent controllers because they reduce remote inference, latency, and deployment footprint, yet structured tool errors can cause an agent step to fail. Existing routers typically select a model once per query. However, agents expose sequential decision points whose difficulty dynamically changes based on intermediate observations. We propose STEPGATE, an uncertainty-aware handoff framework that scores each local SLM action and selectively escalates challenging steps to a stronger model. On a 52-task held-out single-step BFCL-derived test split, the Qwen2.5-1.5B/7B pair attains 82.7% task success with 30.8% escalation, versus 67.3% local-only and 75.4% random escalation (which uses 33.8% escalation). In a separate multi-turn evaluation, STEPGATE achieves 69.0% trajectory success and 84.0% action success using only 30.0% cloud actions, compared with 48.0%/70.5% local-only, 60.0%/78.2% random escalation, and 57.0%/77.1% query-level routing (strong-only achieves 82.0% trajectory success at 100% cloud actions). These results suggest that step-level escalation recovers a large share of the performance gap to the stronger Qwen2.5-7B backend at a matched cloud-action rate while transmitting fewer tokens remotely. However, our evaluation is limited to one model family, a single stronger backend, and scripted tasks. Furthermore, the test sets are small, multi-turn comparisons rely on paired intervals and statistical tests, and our risk tiers serve as research annotations rather than formal safety guarantees.
CommentsAccepted at the 40th Conference on Neural Information Processing Systems (NeurIPS 2026). Workshop: SLMs for Agentic Systems, Paris, France, 2026