arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.20389cs.AI

表征影响检索:多模态智能体框架中技能发现与路由的案例研究

Representation Affects Retrieval: A Case Study of Skill Discovery and Routing in a Multimodal Agent Harness

Kevin Dela Rosa

首次发表
浏览论文内容

中文总结 AI 辅助

该研究以多模态智能体框架Tinycloud为案例,发现提示内技能的部分暴露会产生词汇竞争,抑制正确技能选择,其为大规模技能路由提供了小规模视角。

中文摘要 AI 辅助

生产级智能体框架必须从不断增长的技能库中发现并排序出最适合用户任务的技能。在小规模场景下,这种选择在上下文内完成:大语言模型(LLM)规划器在其系统提示中暴露的技能表征中进行选择,无需显式的基于嵌入的检索步骤。我们将这种上下文内选择视为大规模场景下基于嵌入的技能检索的小规模对应物,并对生产级多模态视频智能体框架Tinycloud如何为规划器表征其技能进行案例研究。该框架以两种重复出现的表征形式提供技能:工具技能(tool-skills),其封装单个外部API或系统工具,作为基础词汇;以及工作流技能(workflow-skills),其编排工具技能调用并结合模板渲染以生成指定交付物。框架通过系统提示中的两个界面暴露这些技能:自动加载技能的内联体界面(完整指令、脚本、模板),以及按需技能的单行列表。针对三种暴露模式(全部开启、默认模式、全部关闭)的六项任务选择 ablation 实验显示:完整自动加载在所有任务中均选对最优技能;全部关闭会减慢执行速度并导致严重的发现失败;而生产默认模式则因词汇信号与自动加载的工具技能冲突,使规划器注意力从列出的工作流技能转移,导致一项任务路由错误。核心发现是,提示内的技能暴露并非单调有益:部分暴露会产生词汇竞争,抑制正确选择。我们将这一小规模观察与近期大规模场景下基于检索的技能路由工作相联系,并将本研究的贡献定位为案例研究而非基准测试。

英文摘要

A production agent harness must discover and rank, from a growing library of skills, the one most appropriate for a user's task. At small scale this selection happens in context: the LLM planner chooses among skill representations exposed in its system prompt, without an explicit embedding-based retrieval step. We treat this in-context selection as the small-N counterpart to embedding-based skill retrieval at scale, and present a case study of how Tinycloud, a production multimodal video agent harness, represents its skills for the planner. The harness ships skills under two recurring representations: tool-skills that wrap a single external API or system tool and serve as primitive vocabulary, and workflow-skills that orchestrate tool-skill calls plus a template render to produce one named deliverable. The harness exposes them via two surfaces in the system prompt: an inlined-body surface (full instructions, scripts, templates) for autoloaded skills, and a one-line listing for on-demand skills. A six-task selection ablation across three exposure regimes (all-on, default, all-off) shows that full autoload selects the gold skill on every task; all-off slows execution and produces hard discovery failures; and the production default misroutes one task because its lexical signal collides with an autoloaded tool-skill that pulls planner attention away from a listed workflow-skill. The headline finding is that in-prompt exposure of skills is not monotonically helpful: partial exposure can create lexical competition that suppresses correct selection. We connect this small-N observation to recent retrieval-based skill-routing work at large scale, and frame this contribution as a case study rather than a benchmark.

发表机构

  • Cloudglue USA(美国Cloudglue公司)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑