arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.22610cs.AI

面向自进化智能体的联盟感知技能可靠性

Coalition-Aware Skill Reliability for Self-Evolving Agents

Qiyan Zhao, Xiaofeng Zhang, Bo Liu, Minda Chen, Wei Xiong, Jingyang Chen, Guanting Ye, Wenhao Yu, Xiaosong Yuan, Shijie Han, Da-Han Wang, Jianmin Ji, Fei Huang, Xu-Yao Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对自进化智能体技能库的可靠性问题,提出CASS与u-SMCO两类干预措施,在多数据集上提升了任务性能与跨域泛化能力,还降低了强化学习对噪声奖励的敏感性。

中文摘要 AI 辅助

智能体技能是从交互轨迹中提炼的结构化人工制品,可从技能库中动态复用,已成为支持基于大语言模型(LLM)的自进化智能体从过往经验学习的核心机制。然而现有工作大多聚焦于技能的获取、进化与检索等操作层面,却未解决一个更基础的可靠性问题:智能体技能库中积累的技能是否确实能在机制层面做出正向贡献?我们通过在不同技能库组成与部署领域开展系统性技能库审计,测量智能体行为的变化来探究该问题。这些审计揭示了两类反复出现的可靠性失效:联盟污染,即库层面的收益掩盖了联盟层面技能的负向贡献;跨域效用反转,即源域有益的技能在迁移后效果反转。这些发现催生了两类可靠性干预措施:技能积累阶段的联盟感知技能选择,以及迁移后的无标签技能掩码。联盟感知技能选择(CASS)利用采样Shapley边际为当前技能库选择更可靠的候选技能;无监督技能掩码联盟优化器(u-SMCO)会掩码那些在无标签目标域数据上排除后可提升检索质量的迁移技能。在LoCoMo、LongMemEval、HotpotQA与ALFWorld上开展的智能体实验表明,CASS与u-SMCO相较于强大的基于技能的自进化智能体基线,可持续提升任务性能与跨域泛化能力。除准确性外,联盟条件可靠性建模还降低了强化学习对噪声结果-奖励波动的敏感性,并揭示了基于孤立的技能评估的局限性。

英文摘要

Agent skills, structured artifacts distilled from interaction trajectories and dynamically reused from skill banks, have become a central mechanism for enabling large language model (LLM)-based self-evolving agents to learn from past experience. Yet existing work has largely focused on the operational aspects of skills, such as acquisition, evolution, and retrieval, while leaving a more fundamental reliability question unresolved: Do accumulated skills in an agent's skill bank actually make positive mechanistic contributions? We investigate this question through systematic skill-bank audits across alternative bank compositions and deployment domains, measuring the resulting changes in agent behavior. These audits reveal two recurring reliability failures: coalition pollution, where bank-level gains conceal negative coalition-level skill contributions, and cross-domain utility reversal, where source-beneficial skills reverse their effects after transfer. These findings motivate two reliability interventions: coalition-aware skill selection during skill accumulation and label-free skill masking after transfer. Coalition-Aware Skill Selection (CASS) selects more reliable candidate skills for the current bank using sampled Shapley marginals. Unsupervised Skill-Masked Coalition Optimizer (u-SMCO) masks transferred skills whose exclusion improves retrieval quality on unlabeled target-domain data. Agentic experiments on LoCoMo, LongMemEval, HotpotQA, and ALFWorld show that CASS and u-SMCO consistently improve task performance and cross-domain generalization over strong skill-based self-evolving agent baselines. Beyond accuracy, coalition-conditioned reliability modeling reduces sensitivity to noisy outcome-reward fluctuations during reinforcement learning and exposes the limits of isolation-based skill evaluation.

↑