arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

阅读并非推理:弥合视觉-文本压缩中的智能体策略差距

Reading is not Reasoning: Bridging the Agentic Policy Gap in Vision-Text Compression

Cheng Fan, Junyi Zhou, Tingzhang Luo, RongJian Xu, Qiyanhui Lu, Mingjian Zhu, Hanting Chen, Jianyuan Guo

arXiv 2608.08960首次发表:更新:

AI 中文总结

该研究针对视觉-文本压缩导致的智能体策略差距,提出CAPS跨模态智能体策略自蒸馏框架,在SearchQA、ALFWorld等数据集上显著提升性能并大幅降低上下文成本。

AI 中文摘要

多步骤语言模型智能体会反复处理不断增长的交互历史,导致产生大量的上下文成本。视觉-文本压缩通过将历史记录渲染为图像来降低这些成本,但由此产生的模态偏移会造成明显的能力差距。通过对历史恢复、匹配状态决策以及完整轨迹进行受控评估,我们发现该差距无法仅用光学字符识别(OCR)的质量来解释。视觉历史智能体在动作选择、查询构建、终止判断和证据使用方面表现出系统性偏差,这揭示了存在一种智能体策略差距。我们提出了CAPS,这是一个两阶段跨模态智能体策略自蒸馏框架,利用同一模型更强的文本历史策略来监督其视觉历史对应策略。离线轨迹自蒸馏将成功的文本策略行为迁移至视觉历史输入,而在线策略自蒸馏则在强化学习过程中对视觉历史策略访问的状态提供密集监督。在SearchQA数据集上,采用3B和7B骨干模型的CAPS分别比AgentOCR提升了5.0%和3.4%;在完整历史的ALFWorld数据集上,对应提升幅度为15.6%和14.5%。在所有设置下,CAPS与匹配的文本历史策略相比,平均记忆上下文成本降低了多达63.3%,峰值成本降低了多达83.4%。这些结果表明,显式的跨模态策略自蒸馏能够在视觉-文本压缩下保留智能体的能力,我们的代码将在后续版本中公开提供。

英文摘要

Multi-step language-model agents repeatedly process growing interaction histories, leading to substantial context costs. Vision--text compression reduces these costs by rendering history as images, but the resulting modality shift creates a marked capability gap. Through controlled evaluations of history recovery, matched-state decisions, and complete trajectories, we show that this gap cannot be explained by OCR quality alone. Visual-history agents exhibit systematic drift in action selection, query formulation, stopping, and evidence use, revealing an agentic policy gap. We introduce \textbf{CAPS}, a two-stage \textbf{C}ross-modal \textbf{A}gentic \textbf{P}olicy \textbf{S}elf-distillation framework that uses the same model's stronger text-history policy to supervise its visual-history counterpart. Offline trajectory self-distillation transfers successful text-policy behavior to visual-history inputs, while online policy self-distillation provides dense supervision on states visited by the visual-history policy during reinforcement learning. On SearchQA, CAPS improves over AgentOCR by 5.0\% and 3.4\% with 3B and 7B backbones, respectively. On full-history ALFWorld, the corresponding gains are 15.6\% and 14.5\%. Across settings, CAPS reduces average memory-context cost by up to 63.3\% and peak cost by up to 83.4\% relative to matched text-history policies. These results show that explicit cross-modal policy self-distillation can preserve agent capability under vision--text compression. Our code will be made publicly available in a future release.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑