arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.17921cs.AI

多智能体VLM系统的协作记忆

Collaborative Memory for Multi-Agent VLM Systems

Huixin Zhang, Shao-Jun Xia, Di Wang, Liangxi Liu, Hainan Xiong, Zihao Wang

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出协作记忆框架,通过记忆层次、跨智能体共享和一致性机制,解决多智能体VLM系统中共享视觉上下文与协调解释的问题,为构建可靠高效的智能体团队奠定基础。

中文摘要 AI 辅助

视觉语言模型(VLM)智能体结合专门的感知、工具和推理能力来处理复杂的视觉任务。在多智能体环境中,不同的智能体检查不同的图像区域、视频帧或视觉表示,因此协作不仅涉及分布式推理,还涉及分布式感知。这使得共享视觉上下文成为VLM智能体协作中的核心问题。在本文中,我们围绕协调解释和更新依赖推理的需求,构建了记忆层次结构、跨智能体共享和一致性机制。有效的协作要求智能体能够基于其他智能体的贡献进行构建、恢复缺失的视觉上下文,并在新证据出现时协调不同的解释。共享视觉记忆不仅保留图像或文本摘要,还保留观察结果、智能体解释和后续推理之间的依赖关系。这些设计考虑共同塑造了信息在VLM智能体之间的流动和演化方式。所提出的框架为构建可靠且资源高效的智能体团队奠定了基础。

英文摘要

Vision-language model (VLM) agents combine specialized perception, tools, and reasoning to address complex visual tasks. In multi-agent settings, different agents inspect different image regions, video frames, or visual representations, so collaboration extends beyond distributed reasoning to distributed perception. This makes shared visual context a central problem in VLM agent collaboration. In this paper, we frame memory hierarchy, cross-agent sharing, and consistency mechanisms around the need to reconcile interpretations and update dependent reasoning. Effective collaboration requires agents to build on contributions from other agents, recover missing visual context, and reconcile differing interpretations as new evidence emerges. Shared visual memory preserves not only images or textual summaries but also the dependencies among observations, agent interpretations, and subsequent reasoning. Together, these design considerations shape how information flows and evolves across VLM agents. The proposed framework provides a foundation for building reliable and resource-efficient agent teams.

发表机构

  • Texas A&M University(德克萨斯A&M大学)
  • Duke University(杜克大学)
  • Foxconn(富士康)
  • Northeastern University(东北大学)
  • Harvard University(哈佛大学)
  • Meta

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑