发表机构
The Memory Company(记忆公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究人工智能冻结权重智能体如何从部署反馈中持续学习,核心方法是将冻结模型与外部记忆配对,主要贡献是通过该方法在银行领域任务中提升成功率,解决多个基线未解决任务,且结果可跨模型,记忆可迁移,相关内容已发布。
AI 中文摘要
人工智能智能体在每次运行中都会遇到学习机会,但几乎全部被丢弃:底层模型在部署时被冻结,所以今天解决难题的智能体明天再次遇到时需从零开始。然而日常操作已产生结果判定和事后修正等形式的反馈。我们表明,当冻结模型与将每段情节提炼为可检索自然语言规则的外部记忆配对时,这种反馈是持续学习的充分信号。在τ-bench的银行领域,与检索完整策略语料库的静态RAG控制相比,从单比特结果判定学习将单次试验成功率提高到基线的1.6倍,从修正学习提高到2.6倍,解决了基线从未解决的84个任务中的22个。该结果在不同模型上都有体现,积累的记忆也可迁移。相关工具、协议和数据已发布。
英文摘要
AI agents encounter learning opportunities in every episode they run, and discard nearly all of them: the underlying models are frozen at deployment, so an agent that resolves a difficult request today starts from zero when it recurs tomorrow. Yet ordinary operation already produces feedback, in the form of outcome verdicts and after-the-fact corrections. We show that this feedback is a sufficient signal for continual learning when the frozen model is paired with an external memory that distils each episode into retrievable natural-language rules. On the banking domain of $τ$-bench, against a static-RAG control retrieving over the complete policy corpus, learning from the one-bit outcome verdict lifts single-trial success to 1.6$\times$ the baseline, and learning from corrections to 2.6$\times$, converting 22 of the 84 tasks the baseline never solves. The result spans the deployment spectrum, measured on Mistral Large, an open-weights model that organisations with data sovereignty requirements can self-host, and replicated on a frontier model, Claude Sonnet 5. The accumulated memory also transfers: each model, reading the store built by the other, rises above its own no-memory baseline. The harness, protocol, and data are released.