AI 中文总结
CONTRAMEM是无需训练的自演进过程记忆框架,通过对比多模型轨迹提炼功能卡与技能卡,在GAIA2/ARE等任务上大幅提升计算机使用智能体成功率,且具备跨模型迁移性。
AI 中文摘要
自主计算机使用智能体越来越多地被应用于需要协调应用调用、持久状态跟踪和验证器敏感写入的长时序任务,但它们仍然容易出现过程性故障:误读应用状态、工具语义或任务进度。过程记忆有望带来更一致的决策和更少的冗余探索,但在不进行模型训练的情况下构建高质量记忆仍然具有挑战性。我们提出了CONTRAMEM,一种源灵活、无需训练的自演进过程记忆框架,该框架将相同任务的结果差异视为监督:正确性、效率、恢复能力和失败模式的差异揭示了与结果相关的过程区分,这些区分被提炼为紧凑的应用级功能卡和任务级技能卡库,该库通过局部整理而非仅追加积累或整体重写进行演进。在保留的GAIA2/ARE计算机使用任务上,CONTRAMEM使三个源模型目标的成功率提高了一倍以上(从26.2%升至55.3%),每个模型均获得一致提升(GPT-5.5:27.5升至61.0;Claude Sonnet 4.6:28.0升至52.5;DeepSeek V4 Pro:23.0升至52.5)。相同的库原封不动地迁移到未见过的Qwen3.7 Plus(从18.5升至35.5),表明是可迁移的过程知识而非模型特定行为。相同的构建方式原封不动地迁移到AppWorld,在两个公开测试分割上击败了无记忆基线及其自身的单源自记忆变体,适用于所有三个中端智能体。在匹配的轨迹预算下,异构多模型轨迹产生的记忆比自模型或同模型多轮次记忆更强:优势来自对比行为多样性,而非更强的源智能体或更多采样。
英文摘要
Autonomous computer-use agents are increasingly applied to long-horizon tasks requiring coordinated application calls, persistent state tracking, and verifier-sensitive writes, yet they remain prone to procedural failures: misreading application state, tool semantics, or task progress. Procedural memory promises more consistent decisions and less redundant exploration, but constructing high-quality memory without model training remains challenging. We introduce CONTRAMEM, a source-flexible, training-free framework for self-evolving procedural memory that treats same-task outcome variation as supervision: differences in correctness, efficiency, recovery, and failure modes expose outcome-relevant procedural distinctions, distilled into a compact bank of app-level Function Cards and task-level Skill Cards that evolves through localized curation rather than append-only accumulation or whole-bank rewriting. On held-out GAIA2/ARE computer-use tasks, CONTRAMEM more than doubles the success rate across the three source-model targets (26.2% to 55.3%), with consistent per-model gains (GPT-5.5: 27.5 to 61.0; Claude Sonnet 4.6: 28.0 to 52.5; DeepSeek V4 Pro: 23.0 to 52.5). The same bank transfers unchanged to the unseen Qwen3.7 Plus (18.5 to 35.5), indicating transferable procedural knowledge rather than model-specific behavior. The same construction carries over unchanged to AppWorld, beating both no memory and its own single-source self-memory variant for all three mid-tier agents on both public test splits. Under a matched trajectory budget, heterogeneous multi-model trajectories yield stronger memory than self- or same-model multi-rollout memory: the margin comes from contrastive behavioral diversity, not stronger source agents or more sampling.
Comments35 pages, 7 figures; includes technical appendix