arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.22793cs.CLcs.AI

TRACE:面向一致性、极限感知的自进化技能库

TRACE: A Self-Evolving Skill Bank for Consistent, Limit-Aware LLM Agents

Wenhao Wu, Menghao Zhang, Xin Wang, Zhi Wang, Kun Shao, Jian Luan

首次发表
浏览论文内容

中文总结 AI 辅助

TRACE是一种无需修改模型权重的自进化技能库方法,通过轨迹对比进化优化技能,在CAR-bench任务上显著提升LLM智能体的一致性性能,缩小潜在与可靠性能的差距。

中文摘要 AI 辅助

LLM智能体在面向用户的产品中的可靠部署不取决于原始任务解决能力,而取决于一致性和极限感知:在重复试验中表现一致,且能识别请求无法或暂无法安全满足的情况。CAR-bench揭示了车载助手领域的可靠性差距:由LLM模拟的用户发出不完整或模糊的请求,要求智能体通过多轮对话和工具使用解决不确定性,同时严格遵守领域策略。即便前沿模型在至少能解决一次的任务(Pass@3)与能在所有试验中一致解决的任务(Pass^k)之间仍存在显著差距。我们通过TRACE(TRAjectory-Contrastive Evolution,轨迹对比进化)弥合这一差距,该方法可迭代提升基于技能的智能体的行为知识,无需修改模型权重。该知识被组织为模块化、可检索的技能库,每个技能编码一组独立的工具使用规则和行为准则。TRACE通过智能体自进化循环进化该库:每次评估轮次后,它按调用的技能分组轨迹,并通过对比成功与失败行为优化每个技能。更新后的库指导后续轮次,部署期间,执行器(Actor)每轮执行基于状态的技能编排。在GPT-5.5上,TRACE将一致性指标(Pass^3)提升34.6个百分点,从59.9%升至94.5%,同时将潜在性能与可靠性能的差距缩小至仅4.0个百分点。在官方隐藏测试集上,TRACE使用GPT-5.6-Sol获得第一名,Pass^3分数达70%,较基线提升40%。这些结果表明,TRACE可将模型的高潜力转化为稳定、一致的性能提升。项目主页:this https URL

英文摘要

Reliable deployment of LLM agents in user-facing products depends not on raw task-solving ability but on consistency and limit-awareness: behaving the same way across repeated trials, and recognizing when a request cannot, or cannot yet, be safely fulfilled. CAR-bench exposes this reliability gap in the domain of in-car assistants: an LLM-simulated user issues incomplete or ambiguous requests, requiring the agent to resolve uncertainty through multi-turn dialogue and tool use while strictly adhering to domain policies. Even frontier models show a substantial gap between what they can solve at least once (Pass@3) and what they solve consistently across trials (Pass^k). We bridge this gap with TRACE (TRAjectory-Contrastive Evolution), which iteratively improves a skill-based agent's behavioral knowledge without modifying model weights. This knowledge is organized as a Skill Bank of modular, retrievable skills, each encoding a self-contained set of tool-use rules and behavioral guidelines. TRACE evolves this bank through an agentic self-evolution loop: after each evaluation round, it groups trajectories by the skills invoked and refines each skill by contrasting successful and failed behaviors. The updated bank then guides subsequent rounds, while during deployment the Actor performs state-conditioned skill orchestration at every turn. On GPT-5.5, TRACE improves consistency (Pass^3) by 34.6 points, from 59.9% to 94.5%, while shrinking the gap between potential and reliable performance to just 4.0 points. On the official hidden set, TRACE achieved first place using GPT-5.6-Sol, attaining a Pass^3 score of 70%-a 40% relative improvement over the baseline. These results show that TRACE converts high model potential into stable, consistent performance gain. Project homepage: https://darwin-agent.github.io/Car-bench-TRACE.

发表机构

  • Xiaomi Inc.(小米公司)
  • Nanjing University(南京大学)
  • Beijing University of Posts and Telecommunications(北京邮电大学)
  • Tsinghua University(清华大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑