MESA:面向长视野智能体记忆的任务自适应多结构证据选择
MESA:Task-Adaptive Multi-Structure Evidence Selection for Long-Horizon Agent Memory
- Microsoft Research Asia(微软亚洲研究院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
MESA是面向长视野智能体记忆的多结构证据选择框架,通过自适应选择融合互补记忆结构,在AMA-Bench上性能优于基线且减少了证据标记使用量。
AI中文摘要:
长视野智能体积累的轨迹包含数百个交织的推理、动作与观测步骤,回答查询可能依赖于历史中深埋的证据。外部存储器将这些轨迹存储为结构化表示,但每种结构提供的视角独特且不完整。现有多记忆系统要么为每个查询读取固定的结构集合,导致上下文膨胀并引入噪声;要么将每个查询路由到单一结构,无法组合互补证据。对AMA-Bench的控制分析显示,最优记忆配置通常既非单一结构,也非全部结构的并集,而是随查询和任务需求变化的多结构记忆定制组合。受这些发现启发,我们提出结构级动态选择:从专用记忆结构库中选择并融合查询自适应子集。我们提出MESA(面向长视野智能体的多结构证据选择框架),它为每条轨迹构建五个互补结构视图,并从端到端答案级反馈中学习,为冻结的答案模型选择并融合查询特定子集。为在这种弱监督下学习,MESA采用带先验引导搜索和UCB引导调度的 harness 优化,以平衡探索与利用。在AMA-Bench上,MESA比最强基线性能提升8.5%,同时使用的证据标记比全结构替代方案少41%。
英文摘要:
Long-horizon agents accumulate trajectories spanning hundreds of interleaved reasoning, action, and observation steps, where answering a query may depend on evidence buried far back in the history. External memory stores such trajectories as structured representations, yet each structure provides a distinct and incomplete view. Existing multi-memory systems either read a fixed set of structures for every query, inflating context and introducing noise, or route each query to a single structure, preventing the composition of complementary evidence. A controlled analysis on AMA-Bench shows that the optimal memory configuration is typically neither a single structure nor the full union, but a tailored composition of multiple structural memories that varies with query and task demands. Motivated by these findings, we formulate structure-level dynamic selection: selecting and fusing a query-adaptive subset from a library of specialized memory structures. We propose MESA (a Multi-structure Evidence Selection framework for long-horizon Agent), which builds five complementary structure views of each trajectory and learns from end-to-end answer-level feedback to select and fuse a query-specific subset for a frozen answer model. To learn under this weak supervision, MESA employs harness optimization with prior-guided search and UCB-guided scheduling to balance exploration and exploitation. On AMA-Bench, MESA outperforms the strongest baseline by 8.5% while using 41% fewer evidence tokens than the all-structure alternative.