发表机构
Tongji University; The University of Sydney; University of Science and Technology of China; Nankai University; Johns Hopkins University; City University of Hong Kong; University of Surrey(同济大学; 悉尼大学; 中国科学技术大学; 南开大学; 约翰斯·霍普金斯大学; 香港城市大学; 萨里大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对基于技能的LLM智能体执行决策难题,提出RADEG轻量型决策层,通过学习代理模型预测执行效用,在288个rollout的评估中减少不必要执行并优于基线方法
AI 中文摘要
智能体技能正日益用于为大语言模型(LLM)智能体配备可复用的过程性知识。尽管由于技能库的扩充,近期工作已大幅改进了技能检索,但检索到合理的技能组合并不保证执行该组合是值得的。由于每个技能条件下的 rollout(执行过程)计算成本高昂,决定是否执行检索到的组合已成为一项日益重要的挑战。为此,我们引入了奖励感知动态执行门控(RADEG),这是一个轻量型、与检索器无关的决策层,位于技能检索和智能体执行之间。RADEG 学习一个低成本的代理模型,在启动昂贵的 rollout 之前预测查询-组合对的执行效用。为在控制任务难度的同时获得有信息量的监督,我们通过删除、添加或替换一个技能,对每个检索到的组合进行局部扰动,生成匹配的同查询 rollout,以分离组合构成对验证器奖励的影响。在部署期间,随着新的验证器反馈可用,RADEG 仅更新一个热启动的逻辑回归头,无需重新训练检索器或智能体即可实现执行/跳过边界的低成本适配。在对收集的 288 个 rollout 进行的查询级留出评估中,RADEG 在保留大部分下游验证器奖励的同时,大幅减少了不必要的智能体执行。它在不同执行预算下始终优于基于相关性的门控和随机门控,表明感知执行的代理建模为技能检索提供了实用且具成本效益的补充。
英文摘要
Agent skills are increasingly used to equip large language model (LLM) agents with reusable procedural knowledge. Although recent work has substantially improved skill retrieval due to the increasing skill libraries, retrieving a plausible skill bundle does not guarantee that executing it is worthwhile. Since every skill-conditioned rollout is computationally expensive, deciding whether a retrieved bundle should be executed has become an increasingly important challenge. To this end, we introduce the Reward-Aware Dynamic Execution Gate (RADEG), a lightweight, retriever-agnostic decision layer between skill retrieval and agent execution. RADEG learns a low-cost surrogate model that predicts the execution utility of a query--bundle pair before the expensive rollout is launched. To obtain informative supervision while controlling for task difficulty, we locally perturb each retrieved bundle by deleting, adding, or replacing one skill, producing matched same-query rollouts that isolate the effect of bundle composition on verifier reward. During deployment, RADEG updates only a warm-started logistic head as new verifier feedback becomes available, enabling inexpensive adaptation of the execute/skip boundary without retraining either the retriever or the agent. Under a query-level held-out evaluation on 288 collected rollouts, RADEG substantially reduces unnecessary agent executions while preserving a large fraction of the downstream verifier reward. It consistently outperforms relevance-based and random gating across different execution budgets, demonstrating that execution-aware surrogate modeling provides a practical and cost-effective complement to skill retrieval.