arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SIGIL:将智能体技能编译为类型化工具框架

SIGIL: Skill Compilation for Reliable and Efficient Agent Execution

Jayanaka Dantanarayana, Savini Kashmira, Lingjia Tang, Jason Mars

arXiv 2607.27309首次发表:更新:

发表机构

University of Michigan(密歇根大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

SIGIL通过Skill Compilation将散文式智能体技能编译为类型化工具框架,提升了智能体执行技能要求步骤的比例、完整过程的频率并减少token使用,且效果与模型无关。

AI 中文摘要

集成AI的智能体越来越多地从技能中获取能力:技能是加载到模型上下文并由工具调用循环运行的散文式过程文件。技能被描述给运行时,但从未编码在其中,因此模型每次运行时都要重新推导其控制流,可能会跳过要求的验证步骤。在30项技能和两代模型中,散文式智能体仅执行其自身技能要求的56%步骤,同时生成通过输出检查的产物。解决方案是编写工具框架(harness),其中过程是程序结构。然而,手动编写工具框架繁琐且会丢失使技能成功的创作界面。为解决此限制,我们引入Skill Compilation(技能编译),在SIGIL中实现,该工具将散文式技能编译为可执行工具框架。其核心是AG-IR,一种类型化智能体中间表示,将模型拥有的认知与代码拥有的机制分离。编译后的工具框架执行86%的要求步骤,完成完整过程的频率是原来的2.3倍,所需token数为原来的0.58倍。值得注意的是,该保证与模型无关:工具框架在两代模型中均达到86%,而散文式智能体的表现从56%波动至68%。

英文摘要

Agent skills describe reusable procedures, but runtime models must still interpret their instructions and coordinate execution. We introduce Skill Compilation, which translates the procedure prescribed by an authored skill into an executable harness while preserving decisions left to the model. Our compiler, SIGIL, translates skills and their resources into a typed intermediate representation, validates it, and deterministically generates the harness. We evaluate SIGIL on 11 compliance-critical skills, where following the prescribed procedure is part of correctness. SIGIL improves adherence across all four runtime models, achieving up to 100% measured mean Skill Adherence compared with 26.2-54.0% for direct SKILL.md execution. On this suite, SIGIL reduces total runtime tokens by 21-45% for three of the four models. On SkillsBench, SIGIL also improves task performance across all four models, with absolute gains of 3.5-35.7 percentage points. These findings suggest that compiling reusable skill procedures into executable harnesses can improve adherence, task completion, and runtime efficiency.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑