发表机构
Meituan; School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences; Institute of Automation, Chinese Academy of Sciences; Zhongguancun Academy(美团; 中国科学院大学先进交叉科学学院; 中国科学院自动化研究所; 中关村学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对LLM智能体仅从成功轨迹提取技能会陷入「技能模仿陷阱」的问题,提出边界感知技能记忆(BASM),通过补充边界字段提升工具使用可靠性,在多基准测试中显著优于基线方法。
AI 中文摘要
从过往成功中提取技能对大语言模型(LLM)智能体的高效进化至关重要。主流的智能体自进化范式通常依赖一个核心假设:为LLM配备源自成功轨迹的技能记忆,将单调提升其问题解决能力。然而,探测分析显示,仅从成功轨迹中提取技能会使模型陷入「技能模仿陷阱」。对于与过往成功相似但需要不同工具的任务,检索更多技能反而会悖论式地提升模型对错误工具调用的置信度——过程技能使错误工具调用的边际比无记忆基线提高了47%。为克服此局限,我们提出**边界感知技能记忆(Boundary-Aware Skill Memory, BASM)**,它为每项技能补充显式边界字段:适用条件、风险提示、规避规则及恢复说明。这些字段将每项检索到的技能从无条件动作模板转化为状态条件指导:智能体在条件满足时应用该技能,在不满足时抑制不合适的工具调用,在执行失败时发起针对性修复。在三个智能体基准和四个模型规模下,BASM始终优于成功蒸馏技能记忆基线:它在AppWorld上将任务成功率提高了多达23.8%,在BFCL上将准确率提高了多达5.0%,在AgentDojo上将攻击成功率降低了4.6%,同时相对于无记忆基线将AppWorld的平均步骤数减少了多达6.6%。
英文摘要
Extracting skills from past successes is critical for the efficient evolution of Large Language Model (LLM) agents. Prevailing agent self-evolution paradigms typically rely on a core assumption: equipping LLMs with skill memories derived from successful trajectories will monotonically improve their problem-solving capabilities. However, probe analyses reveal that extracting skills solely from successful trajectories traps the model in a \textbf{Skill Imitation Trap}. For tasks that resemble past successes but require different tools, retrieving more skills paradoxically increases the model's confidence in wrong tool calls---procedure skills raise the wrong-tool margin by $47\%$ over a memory-free baseline. To overcome this limitation, we propose \textbf{Boundary-Aware Skill Memory} (BASM), which augments each skill with explicit boundary fields---applicability conditions, risk cues, avoidance rules, and recovery notes. These fields transform each retrieved skill from an unconditional action template into state-conditioned guidance: the agent applies the skill when its conditions hold, suppresses inapplicable tool calls when they do not, and issues targeted repairs when execution fails. Across three agent benchmarks and four model scales, BASM consistently outperforms success-distilled skill-memory baselines: it improves task success rate by up to $23.8\%$ on AppWorld, accuracy by up to $5.0\%$ on BFCL, and reduces attack success rate by $4.6\%$ on AgentDojo, while simultaneously reducing average AppWorld steps by up to $6.6\%$ relative to the memory-free baseline.
CommentsAccepted by EMNLP2026 Findings