arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.06830cs.CLcs.AIcs.LG

MemPilot:为LLM智能体编排按需多模态记忆策展

MemPilot: Orchestrating On-Demand Multimodal Memory Curation for LLM Agents

  • Nanyang Technological University(南洋理工大学)
  • Tsinghua University(清华大学)
  • University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

机构由 AI 辅助整理,请以论文原文为准。

Haozhen Zhang, Haodong Yue, Quanyu Long, Jianzhu Bao, Qingyuan Liu, Tao Feng, Bohan Liu, Weida Liang, Wenya Wang

AI总结:

MemPilot通过强化学习优化多步LLM策略,在性能、成本和延迟偏好下按需策展多模态记忆,并在五个基准上取得更优权衡。

AI中文摘要:

记忆已成为LLM智能体生态系统中不可或缺的一部分,支持跨交互的信息保留和重用。然而,现有的大多数智能体记忆系统以查询无关的方式构建记忆,这可能会产生不必要的预处理成本,并丢弃后来证明至关重要的细节。近期研究开始将记忆处理转向运行时自适应,但通常专注于特定操作或固定处理方案,对性能、成本和延迟的灵活控制仍鲜有探索。为解决这一挑战,我们提出了MemPilot,一个灵活的框架,在不同性能-成本-延迟偏好下编排按需记忆策展。具体而言,我们通过强化学习优化多步LLM策略,以迭代地在从查询无关记忆中检索和将原始多模态历史的查询特定策展委托给异构LLM和VLM之间进行选择。该策略联合控制证据量、策展指令、模型选择和视觉访问,实现对运行时计算的细粒度分配。为了在竞争目标下优化该策略,我们采用目标级优势解耦,分别估计每个目标的优势后再聚合。此外,我们引入基于前缀的边际效用估计,用于多步轨迹中的细粒度信用分配。在五个多模态智能体记忆基准上的实验表明,在不同优化偏好下取得了有利的性能-成本-延迟权衡,偏好扫描产生的边界比现有权衡感知基线更广。

英文摘要:

Memory has become integral to the LLM agent ecosystem, supporting information retention and reuse across interactions. However, most existing agent memory systems construct memory in a query-agnostic manner, which can incur unnecessary preprocessing cost and discard details that later prove essential. Recent studies have begun shifting memory processing toward runtime adaptation, but typically specialize in particular operations or fixed processing schemes, leaving flexible control over performance, cost, and latency largely underexplored. To address this challenge, we present \textbf{MemPilot}, a flexible framework that orchestrates on-demand memory curation under different performance--cost--latency preferences. Specifically, we optimize a multi-step LLM policy via reinforcement learning to iteratively choose between retrieving from query-agnostic memory and delegating query-specific curation of raw multimodal history to heterogeneous LLMs and VLMs. The policy jointly controls evidence amount, curation instructions, model selection, and visual access, enabling fine-grained allocation of runtime computation. To optimize this policy under competing objectives, we adapt objective-wise advantage decoupling by separately estimating each objective's advantage before aggregation. Moreover, we introduce prefix-based marginal utility estimation for fine-grained credit assignment across multi-step rollouts. Experiments on five multimodal agent-memory benchmarks demonstrate favorable performance--cost--latency trade-offs across optimization preferences, with preference sweeps yielding broader frontiers than existing trade-off-aware baselines.

补充信息

↑