arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

当历史是多模态时:重新思考长视野智能体的上下文管理

When History Is Multimodal: Rethinking Context Management for Long-Horizon Agents

Jiaqi Su, Cong Pang, Jiawei Hong, Tiankuo Yao, Zixuan Chen, Xin Lou, Lewei Lu

arXiv 2608.29897首次发表:更新:

发表机构

National University of Singapore; ShanghaiTech University; SenseTime Research(新加坡国立大学; 上海科技大学; 商汤科技研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出视觉渲染(VR)作为上下文管理器,并构建无需训练的VERA策略,在多模态任务中保留原生视觉观测,相比无压缩减少31.5%-63.1%的累计非缓存令牌,在多模态基准中准确率最高。

AI 中文摘要

长视野智能体需要一个上下文管理器,通过被动策略或决定如何访问和重组记忆的主动策略,将不断增长的交互历史压缩为有界的工作上下文。同时,现有光学记忆工作主要将像素视为文本化历史的密集编解码器,通常假设将上下文渲染为光学记忆会相对于文本产生显著的性能下降,因此将这种表示与SFT(监督微调)、自蒸馏或强化学习结合以缩小差距,但仍未解决两个问题:(i)在公平、受控的对比下,视觉渲染作为上下文管理器的表现如何;(ii)当历史本质上是多模态时,这种载体是否具有天然优势。本文将上下文管理形式化为预算约束下的历史转换,并引入视觉渲染(Visual Rendering, VR)作为表示性上下文管理器。在共享的测试环境、策略模型、触发机制和任务域下,我们在4个以文本为中心的基准和3个多模态基准上,将VR与4个基线(无压缩、全部丢弃、滑动窗口、摘要)进行评估,发现视觉记忆是天然的视觉证据载体。基于这一发现,我们提出VERA(Visual Evidence-Retaining strategy for long-horizon Agents,长视野智能体的视觉证据保留策略),这是一种无需训练的上下文管理器,基于确定性渲染构建,无暴露的记忆操作:在以文本为中心的基准上,它像VR一样将文本历史渲染为视觉形式;在多模态基准上,它保留原生视觉观测而非将其转换为文本。在几乎所有基准中,VERA相比无压缩将累计非缓存令牌减少了31.5%-63.1%,在以文本为中心的任务上与现有管理器表现相当,且在多模态任务中达到所有基线中的最高准确率,支持长视野上下文管理的模态保留观点。

英文摘要

Long-horizon agents need a context manager to compress growing interaction histories into a bounded working context, via passive strategies or active strategies that decide how memory is accessed and reorganized. Meanwhile, prior optical-memory work mainly treats pixels as a dense codec for textualized histories, often presupposing that rendering context into optical memory incurs a significant performance drop relative to text, thus coupling this representation with SFT, self-distillation, or reinforcement learning to close this gap, leaving unresolved (i) how visual rendering performs as a context manager under a fair, controlled comparison, and (ii) whether this carrier offers a native advantage when history is inherently multimodal. In this paper, we formulate context management as a budget-constrained history transformation and introduce Visual Rendering (VR) as a representational context manager. Under a shared harness, policy model, trigger, and task domain, we evaluate VR on four text-centric and three multimodal benchmarks against four baselines (No Compression, Discard-All, Sliding Window, Summarization), finding visual memory is a natural carrier of native visual evidence. Building on this finding, we propose VERA (Visual Evidence-Retaining strategy for long-horizon Agents), a training-free context manager built on deterministic rendering with no exposed memory operations: on text-centric benchmarks it renders textual history as VR does, while on multimodal benchmarks it retains native visual observations instead of translating them into text. Across nearly all benchmarks, VERA cuts cumulative non-cache tokens by 31.5%-63.1% versus No Compression, matches existing managers on text-centric tasks, and achieves the highest accuracy among all baselines on multimodal tasks, supporting a modality-preserving view of long-horizon context management.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑