AutoMem:面向自动化记忆架构搜索的文本梯度递归自改进框架
AutoMem: A Text-Gradient Recursive Self-Improvement Framework for Automated Memory Architectures Search
- School of Computer Science and Technology, East China Normal University(华东师范大学计算机科学与技术学院)
- Shanghai AI Laboratory(上海人工智能实验室)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
AutoMem是面向任务自适应记忆架构搜索的文本梯度递归自改进框架,通过经验引导架构搜索与失败引导模块诊断,在多基准测试中性能优于人工设计基线,实现了准确率与效率的良好权衡。
AI中文摘要:
长期记忆对大型语言模型(LLM)智能体的重要性日益提升,但记忆设计仍是一个高度耦合的架构问题:编码内容、存储方式、检索方法及管理策略会随任务和主干模型的不同而存在显著差异。我们构建了包含5种编码器、5种存储器、6种检索器和4种管理器的离散搜索空间,结果表明不存在始终占优的单一记忆架构:不同任务偏好不同的模块组合,导致性能存在显著差距。基于此,我们提出AutoMem,一种面向任务自适应记忆架构搜索的文本梯度递归自改进框架。AutoMem通过两个组件在分解空间中进行优化:经验引导的架构搜索,从历史搜索轨迹和累积反思中提出候选架构;以及失败引导的模块诊断,将与记忆相关的失败定位到特定模块,并转化为针对性的文本反馈。在GAIA、WebWalkerQA和xBench-DeepSearch数据集上,针对两种LLM主干模型的实验表明,AutoMem始终能发现优于最强人工设计记忆基线的任务自适应记忆架构,在6个基准-主干设置中平均准确率提升2.8个百分点。进一步分析显示,AutoMem实现了良好的准确率-效率权衡,在Qwen3.5-122B-A10B模型下,相比最强准确率基线减少14.3%的token成本,且仅需少量引导迭代就能找到比规模大得多的随机搜索更强的架构。
英文摘要:
Long-term memory is increasingly central to LLM agents, yet memory design remains a highly coupled architecture problem: what to encode, how to store it, how to retrieve it, and how to manage it can vary substantially across tasks and backbone models. We construct a discrete search space with 5 encoders, 5 stores, 6 retrievers, and 4 managers, and show that no single memory architecture consistently dominates: different tasks favor different module combinations, leading to substantial performance gaps. Motivated by this, we propose \textsc{AutoMem}, a text-gradient recursive self-improvement framework for task-adaptive memory architecture search. \textsc{AutoMem} optimizes over the factored space through two components: Experience-Guided Architecture Search, which proposes candidate architectures from historical search trajectories and accumulated reflections, and Failure-Guided Module Diagnosis, which localizes memory-related failures to specific modules and converts them into targeted textual feedback. Experiments on GAIA, WebWalkerQA, and xBench-DeepSearch across two LLM backbones show that \textsc{AutoMem} consistently discovers task-adaptive memory architectures that outperform the strongest human-designed memory baselines, improving accuracy by $2.8$ points on average across six benchmark-backbone settings. Further analysis shows that \textsc{AutoMem} achieves a favorable accuracy-efficiency trade-off, reducing token cost by $14.3\%$ over the strongest accuracy baselines under Qwen3.5-122B-A10B, while also finding stronger architectures than substantially larger random searches within only a few guided iterations.