发表机构
Huawei Technologies, Co., Ltd.(华为技术有限公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
RefCon通过顺序自精炼与并行自对比,无需黄金标签即可从嘈杂经验中提取高质量记忆,在多个基准上显著提升上下文演化智能体性能,并泛化至软件工程任务。
AI 中文摘要
长时程智能体交互会产生有用但嘈杂的经验,而重新训练模型来吸收这些经验成本高昂。因此,上下文演化智能体需要记忆提取方法,这些方法能随着测试时计算量的增加而改进,且不依赖黄金标签。我们提出了RefCon,它将顺序自精炼与并行自对比相结合,以在无黄金标签的情况下提取更高质量的记忆。在AppWorld和BFCL-V3上,跨多种上下文演化智能体框架的评估表明,RefCon带来了强劲且一致的改进,包括在ACE上相对提升21.6%,在ReMe上相对提升16.6%(相较于无扩展基线),而一个注重多样性的变体(DivCon)在ReasoningBank上实现了35.5%的提升。RefCon在无真实标签的情况下持续优于现有基线,并泛化到不同模型规模及软件工程任务,在这些任务中甚至超越了使用真实标签的基线。我们进一步分析了准确率-令牌权衡和扩展行为,显示RefCon保持了良好的效率,并随着使用更多轨迹而持续改进,这与仅注重多样性的扩展(其更早饱和)不同。
英文摘要
Long-horizon agent interactions generate useful but noisy experience, and retraining models to absorb it is expensive. Context-evolving agents therefore need memory extraction methods that improve with more test-time compute without relying on gold labels. We propose RefCon, which combines sequential self-refinement with parallel self-contrast to extract higher-quality memories without gold labels. Evaluated on AppWorld and BFCL-V3 across multiple context-evolving agent frameworks, RefCon delivers strong and consistent gains, including relative improvements of 21.6% on ACE and 16.6% on ReMe over no-scaling baselines, while a diversity-focused variant (DivCon) achieves a 35.5% gain on ReasoningBank. RefCon consistently outperforms existing baselines without ground-truth labels, and generalizes across model scales and to software engineering tasks, where it surpasses even ground-truth baselines. We further analyze the accuracy-token trade-off and scaling behavior, showing RefCon maintains favorable efficiency and continues to improve as more trajectories are used, unlike diversity-only scaling which saturates earlier.