接地智能体记忆:面向企业智能体的环境探测式策展
Grounding Agent Memory: Environment-Probing Curation for Enterprise Agents
浏览论文内容
中文总结 AI 辅助
提出环境探测式策展,为智能体记忆策展提供只读世界工具,无需重训练即可提升CLBench通过率至73%并降低成本,优于现有基线。
中文摘要 AI 辅助
持久记忆正进入面向生产的智能体平台,以帮助长时程智能体跨会话积累经验。然而,一个仅限已完成轨迹的后任务策展智能体可能保留错误、过度概括部分证据或保留过时知识。我们引入环境探测式策展,这是一种部署兼容的扩展,为现有的异步策展智能体提供最小权限、只读的世界工具,以检查、界定和刷新候选记忆。它无需模型重训练,并保持任务智能体、检索器、记忆表示和生产写入权限不变。在基于其SDK构建的类生产GitHub Copilot (GHCP) 测试台中,我们在CLBench数据库探索和90个改编的APEX管理咨询任务上比较了无状态执行、全上下文学习、GHCP + Mem和GHCP + Mem(带环境探测)。在CLBench上,探测将通过率从39%提升至73%,通过折扣奖励从8.60提升至22.60,同时将每问题查询次数从8.8降至4.7,任务智能体成本从3.38美元降至1.68美元。在六个APEX世界中,所有18项记忆与基线平均奖励比较均为正向,任务智能体工具调用减少16%至75%;探测在五个世界中提供了最佳的任务智能体每美元奖励增益。探测在Sonnet 4.6和Opus 4.7上均获得比GHCP + Mem更高的平均奖励,且无模式漂移。因此,环境探测将现有的智能体记忆策展转变为环境知情、可审计的过程,同时保持紧凑的任务时接口。
英文摘要
Persistent memory is entering production-oriented agent platforms to help long-horizon agents accumulate experience across sessions. Yet a post-task curator agent restricted to completed trajectories can preserve errors, overgeneralize partial evidence, or retain stale knowledge. We introduce environment-probing curation, a deployment-compatible extension that gives an existing asynchronous curator agent least-privilege, read-only world tools to check, scope, and refresh candidate memories. It requires no model retraining and leaves the task agent, retriever, memory representation, and production write authority unchanged. In a production-like GitHub Copilot (GHCP) harness built on its SDK, we compare stateless execution, full in-context learning, GHCP + Mem, and GHCP + Mem (w/ Env Probing) on CLBench database exploration and 90 adapted APEX management-consulting tasks. On CLBench, probing raises pass rate from 39% to 73% and pass-discounted reward from 8.60 to 22.60 while reducing queries from 8.8 to 4.7 per question and task-agent cost from \$3.38 to \$1.68. Across six APEX worlds, all 18 memory-versus-baseline mean reward comparisons are positive and task-agent tool calls fall by 16--75%; probing gives the best task-agent reward gain per dollar in five worlds. Probing also attains higher mean reward than GHCP + Mem on both Sonnet 4.6 and Opus 4.7 without schema drift. Environment probing therefore turns existing agent-memory curation into an environment-informed, auditable process while preserving a compact task-time interface.
发表机构
- Microsoft Corporation(微软公司)
机构由 AI 辅助整理,请以论文原文为准。