SkillSeam:审计智能体技能集合的六项原则
SkillSeam: Six Principles for Auditing Agent Skill Collections
- X32 Studio(X32工作室)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
SkillSeam提出六项原则审计智能体技能集合的接缝失效,通过扰动测试量化各原则的失效机制,将编写建议转化为可测试的系统属性。
AI中文摘要:
一个包含称职技能的文件夹尚不能构成一个可靠的系统。技能很少单独失败;它们往往在集合的接缝处失败。随着智能体技能库的增长,程序之间相互竞争注意力,别名导致双重加载,边界变得模糊,而粒度不当的技能将路由错误转化为任务失败。我们提出SkillSeam,一种审计个体技能文件之间关系以使其成为系统的方法。它将每个集合级原则映射到一个失效机制、其最强可观测信号以及一个受控的扰动测试。从一个封闭的技能系统出发,SkillSeam扰动六项设计原则:持久性梯度、系统一致性、机制门控、正交覆盖、流程和粒度纪律。关键在于,每项原则都通过其失效机制所预测的渠道进行评估,而非仅通过准确性。将持久性层级扁平化会使加载的技能令牌增加60%;悬空的锚点使总令牌增加64%,并使准确性下降3.1个百分点;在技能数量和上下文大小固定的情况下,用同义别名替换无关对照会使非规范路由从0/32增加到15/32,并使一半的匹配释义对发生翻转;在候选所有权审计中,重叠的通道使报告的所有权冲突从0/16增加到14/16;平淡的触发器使路由冲突从3/32增加到30/32,并将加载的技能令牌膨胀3.7倍;而一个粒度错配导致最大的准确性下降,达12.5个百分点。这些结果将六条编写建议转化为可测试的系统属性,而无需将每次探测都视为确认。我们发布了字节差分变体、任务切片、汇总以及一屏设计检查清单,以便其他技能系统能够测量相同的失效渠道。
英文摘要:
A folder of competent skills is not yet a reliable system. Skills rarely fail alone; they fail at the seams of a collection. As an agent's skill library grows, procedures compete for attention, aliases double-load, boundaries blur, and poorly sized skills turn routing errors into task failures. We introduce SkillSeam, a method that audits the relationships through which individual skill files become a system. It maps each collection-level principle to a failure mechanism, its strongest observable, and a controlled perturbation test. From one sealed skill system, SkillSeam perturbs six design principles: persistence gradient, system coherence, regime gating, orthogonal coverage, flow, and granularity discipline. Crucially, each principle is evaluated through the channel its failure mechanism predicts rather than through accuracy alone. Flattening the persistence hierarchy increases loaded-skill tokens by 60%; a dangling anchor raises total tokens by 64% and shifts accuracy by -3.1pp; with skill count and context size held fixed, replacing an unrelated control with a synonymous alias raises noncanonical routes from 0/32 to 15/32 and flips half of matched paraphrase pairs; in a candidate-ownership audit, overlapping lanes raise reported ownership conflicts from 0/16 to 14/16; bland triggers drive routing conflicts from 3/32 to 30/32 and inflate loaded-skill tokens 3.7x; and one granularity mis-mix produces the largest accuracy drop, -12.5pp. These outcomes turn six pieces of authoring advice into testable system properties without treating every probe as confirmation. We release the byte-differenced variants, task slices, rollups, and a one-screen design checklist so that other skill systems can measure the same failure channels.