arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SkillSpec: 基于共识门控与表示特化的智能体技能演化

SkillSpec: Consensus-Gated Agent Skill Evolution via Representation Specialization

Huancheng Chen, Xiaodi Sun, Zhaoqiong Huang, Shenyang Huang Shreya Singhal, Jingwen Lu

arXiv 2610.00704首次发表:更新:

发表机构

Microsoft AI(微软人工智能)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

SkillSpec通过共识门控演化与表示特化两阶段框架,提升LLM智能体技能的可重用性与成功率,在六个基准上平均提升6.89%。

AI 中文摘要

自然语言技能是文本形式的过程性记忆,大型语言模型(LLM)智能体通过它们保留可复用的任务知识,而无需更新模型权重。现有方法通常将技能视为静态工件或使用聚合验证分数作为反馈进行优化的整体文档。然而,将技能表示为整体文档限制了优化仅限于其文本内容,无法显式建模过程性知识被检索和执行的架构。我们识别出学习保留哪些知识与决定如何组织这些知识之间的关键区别:文本更新应首先通过执行证据进行验证,之后保留的知识应根据其过程依赖性和检索需求进行结构化。为此,我们引入了SkillSpec,一个包含共识门控演化和表示特化两个阶段的框架。在共识门控阶段,互补的编辑意图生成完整的候选技能。仅当配对评估达到共识时,更新才会被提交,要求每次重复评估中都有足够的整体改进和非负的聚合配对增益。在特化阶段,从完整优化轨迹(包括接受和拒绝的候选)中导出的过程和冗余敏感性信号,指导选择扁平、图或混合表示。在六个基准和三个目标语言模型上,SkillSpec相比SkillOpt在三个模型上的平均成功率提高了6.89%。这些结果表明,可靠的技能演化和表示特化解决了互补的目标:决定保留哪些知识以及如何为推理组织这些知识。

英文摘要

Natural-language skills are textual procedural memories through which large language model (LLM) agents retain reusable task knowledge without updating model weights. Existing methods typically treat skills as either static artifacts or monolithic documents optimized using aggregate validation scores as feedback. However, representing a skill as a monolithic document restricts optimization to its textual content, without explicitly modeling the structure through which procedural knowledge is retrieved and executed. We identify a key distinction between learning what knowledge to retain and determining how to organize it: textual updates should first be validated through execution evidence, after which the retained knowledge should be structured according to its procedural dependencies and retrieval requirements. To this end, we introduce SkillSpec, a two-phase framework comprising consensus-gated evolution and representation specialization. In the consensus-gated phase, complementary editing intents generate complete candidate skills. An update is committed only when paired evaluations reach consensus, requiring sufficient overall improvement and non-negative aggregate paired gain in every repeated evaluation. In the specialization phase, signals of process and redundancy sensitivity derived from the full optimization trajectory, including accepted and rejected candidates, guide the selection of a flat, graph, or hybrid representation.Across six benchmarks and three target language models, SkillSpec improves average success rate over SkillOpt by 6.89%, averaged across the three models. These results demonstrate that reliable skill evolution and representation specialization address complementary objectives: deciding what knowledge to retain and how to structure it for inference.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑