arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.00011cs.IRcs.AIcs.SE

SkillSelect-Serve:面向小型LLM代理的预算可控且QoS感知的技能服务推荐与组合

SkillSelect-Serve: QoS-Aware Budgeted Skill Service Recommendation for LLM Agents

Jingyuan Zheng, Dongjing Wang, Xin Zhang, Hao Chen, Youhuizi Li, Xudong Shen, Haiping Zhang, Butian Huang, Dongjin Yu, Guandong Xu

AI总结:

提出SkillSelect-Serve框架,将技能选择建模为服务推荐与组合问题,通过双粒度效用建模和预算约束优化,在35,353个技能和586个任务查询上优于固定top-k检索基线。

AI中文摘要:

可复用技能库正成为大型语言模型(LLM)代理的重要基础设施,然而现有选择方法通常将技能视为可检索文档并返回固定的top-k列表。本文提出SkillSelect-Serve,一个预算可控且QoS感知的框架,将代理技能选择形式化为技能服务推荐与组合。SkillSelect-Serve将原始技能表示为结构化的技能服务,包含功能描述、依赖关系、上下文成本、风险及QoS相关属性。本地微代理需求规划器将自然语言任务转换为结构化服务需求,而共享发现骨干从大型注册表中检索候选服务。该框架随后执行双粒度效用建模,包括技能级边际适用性估计和包级校准,以权衡覆盖率、冗余、成本和风险。在35,353个技能和586个任务查询上的实验表明,SkillSelect-Serve在相同预算下始终优于固定top-k检索基线的包召回率和平均效用。

英文摘要:

Reusable agent skills are emerging as a service-oriented capability layer for Large Language Model (LLM) agents. Unlike plain retrieval items, a skill exposes functional capabilities, input-output assumptions, tool dependencies, context cost, and risk metadata. Selecting skills is particularly challenging for small LLM agents, which can load only a few capability units under restricted context, tool availability, and risk tolerance. Existing fixed Top-k methods rank skills by textual relevance and overlook requirement satisfaction, deliverability, and operational constraints. We present SkillSelect-Serve, a QoS-aware, budget-constrained Skill Service recommendation framework. Raw skills are profiled as structured Skill Services, the task is converted into a structured requirement object, and candidates discovered from a large-scale registry are ranked by a calibrated task-conditioned suitability estimator and packed by a constrained projection enforcing token-budget, aggregated-risk, and tool-availability constraints, using only deployment-observable features. On a registry of 35,353 skills with pooled multi-positive relevance judgments verified by two independent assessors, the unconstrained top-5 recommendation fits a realistic 4,000-token context for only 9.1% of tasks; the constrained projection restores 100% deliverability at a cost of only 1.14 points of hit rate, outperforming retrieve-and-rerank, budget truncation, and diversity-based selection under identical budgets. The same mechanism halves delivered risk exposure and eliminates the 44-81% tool-violation rates of tool-agnostic recommendation. At an identical three-service budget, hit rate improves from 0.8864 to 0.9091 over fixed Top-3 retrieval. The results support managing reusable agent skills as discoverable, comparable, and constraint-aware service units instead of plain retrievable documents.

补充信息

↑