发表机构
Shandong University; Southeast University; Kuaishou; Hong Kong University of Science and Technology; Harbin Institute of Technology (Shenzhen)(山东大学; 东南大学; 快手; 香港科技大学; 哈尔滨工业大学(深圳))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对多模态智能体检索中的上下文爆炸问题,提出无需训练的MM-ContextFold框架,通过按需加载图像并折叠分支结果,在七个基准上将平均准确率提升6.3个百分点,上下文长度减少27.5%。
AI 中文摘要
多模态智能体检索(MAR)要求智能体通过迭代调用外部工具来解决复杂的信息寻求任务。典型的框架如ReAct将原始多模态输入和不断累积的交互历史保存在单一且持续增长的上下文中,导致上下文爆炸问题。虽然现有方法通过压缩冗余文本来缓解这一问题,但针对管理令牌密集型视觉内容的有效策略仍 largely 未被探索。为弥补这一空白,我们首先对约10,000条轨迹进行了系统的实证研究。结果表明,随着视觉线索通过外部工具逐步提取并文本化到上下文中,原始图像变得越来越冗余。持续保留图像与更高的输出熵相关,甚至可能降低任务准确率。受这些发现的启发,我们提出MM-ContextFold,一种无需训练的框架,仅在需要时加载原始图像。它维护一个持久的纯文本主上下文用于高层规划,并为依赖图像的子任务生成临时分支上下文。在每个分支中,智能体加载相关图像,完成子任务,并将结果作为简洁的文本摘要折叠回主上下文;随后丢弃图像和分支轨迹。在五个骨干模型上的七个MAR基准上的实验表明,MM-ContextFold相比ReAct将平均准确率提高了6.3个百分点,同时将工作上下文长度减少了27.5%。
英文摘要
Multimodal Agentic Retrieval (MAR) requires agents to solve complex information-seeking tasks by iteratively invoking external tools. Typical frameworks such as ReAct maintain raw multimodal inputs and the accumulating interaction history in a single, ever-growing context, leading to the context explosion problem. While existing methods alleviate this issue by compressing redundant text, effective strategies for managing token-intensive visual content remain largely underexplored. To address this gap, we first conduct a systematic empirical study of approximately 10,000 trajectories. The results show that as visual cues are progressively extracted through external tools and textualized into the context, raw images become increasingly redundant. Continued image retention is associated with higher output entropy and can even degrade task accuracy. Motivated by these findings, we propose MM-ContextFold, a training-free framework that loads raw images only when needed. It maintains a persistent, text-only main context for high-level planning and spawns ephemeral branch contexts for image-dependent subtasks. Within each branch, the agent loads the relevant images, completes the subtask, and folds the result back into the main context as a concise textual summary; the images and branch trace are then discarded. Experiments on seven MAR benchmarks across five backbone models show that MM-ContextFold improves average accuracy by 6.3 percentage points over ReAct while reducing the working context length by 27.5\%.