发表机构
Tencent Cloud Andon; Zhejiang University(腾讯云Andon; 浙江大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
SkillEvo通过多轮用户模拟生成反馈、独立治理层修复退化,在6类云服务等数据集上,相比两种基线方法显著提升了智能体技能进化效果。
AI 中文摘要
智能体技能目前要么是人工编写,要么是通过大型语言模型(LLM)的单次生成过程产生,因此不存在可通过自身交互产生的失败来实现改进的闭环。近期研究虽构建了该闭环,但其反馈源于单轮问答评估,这导致了严重的不对称性:首轮修复了单次交互可暴露的缺陷后,进化梯度便会衰减,仅在多轮交互中才会显现的缺陷仍无法被察觉,进化陷入停滞。此类系统的治理同样由端到端验证分数驱动,该标量闸门可拒绝退化的候选方案,但无法定位或修复其结构性原因。本文认为,持续技能进化的核心约束并非编辑能力或迭代次数,而是评估反馈能否持续提供可信的进化梯度。我们提出SkillEvo,其中可信反馈生成梯度,可控治理层约束其方向。第一组件将多轮用户模拟从评估端点重构为反馈生成器:后续提问逐层暴露缺陷,使每轮修订既消耗反馈又产生新反馈。第二组件将标量闸门的被动拒绝替换为独立治理层,主动修复事实退化与结构冗余,防止梯度随退化累积而漂移。在6类云服务、9个生产技能及98份技能参考文件上,SkillEvo相比基于自反思的进化提升了23.0个百分点,相比单轮问答驱动的进化提升了15.4个百分点。
英文摘要
Agent Skills are today either hand-authored or produced in a single LLM generation pass, and consequently possess no closed loop through which they might improve from the interaction failures they actually cause. Recent work does close this loop, but derives its feedback from single-turn question-answering evaluation. The consequence is a sharp asymmetry: once the first round has patched the gaps that a single exchange can reveal, the evolution gradient decays, the defects that surface only across multiple turns remain invisible, and evolution stalls. Governance in these systems is likewise driven by an end-to-end verification score, a scalar gate that can reject a degraded candidate but can neither localize nor repair its structural cause. We argue that the binding constraint on sustained skill evolution is neither editing capability nor the number of iterations, but whether the evaluation feedback keeps supplying trustworthy evolution gradients. We introduce SkillEvo, in which trustworthy feedback generates the gradient and controllable governance constrains its direction. The first component recasts multi-turn user simulation from an evaluation endpoint into a feedback generator: follow-up questions expose defects layer by layer, so that every round of revision both consumes feedback and produces new feedback. The second replaces the passive rejection of a scalar gate with an independent governance layer that actively repairs factual degradation and structural bloat, preventing the gradient from drifting as degradation accumulates. Across six categories of cloud services, 9 production Skills, and 98 skill-reference files, SkillEvo surpasses self-reflection-based evolution by 23.0 points and single- turn-QA-driven evolution by 15.4 points.