arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ViSAGE:为长视频理解构建自修正记忆

ViSAGE: Constructing Self-Correcting Memories for Long-Form Video Understanding

Xinkui Zhao, Enbo Chen, Yifan Zhang, Chang Liu, Guanjie Cheng, Naibo Wang, Yueshen Xu

arXiv 2607.28678首次发表:更新:

发表机构

Zhejiang University; Xidian University(浙江大学; 西安电子科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

ViSAGE是一种多模态智能体记忆框架,通过跨模态绑定、双向记忆精调及多智能体交叉验证解决长视频理解中的实体混淆等问题,较最强基线准确率提升5.9%。

AI 中文摘要

在长时序环境中运行的多模态智能体必须构建并持续更新多媒体记忆,以支持实体一致、时间定位的推理。然而,现有智能体记忆方法常因过度压缩和分段处理丢弃细粒度实体线索,且过度依赖向量相似性检索,会返回语义相关但实体不匹配的证据,导致实体混淆、错误传播和答案幻觉。我们提出ViSAGE,一种构建自修正、以实体为中心记忆的多模态智能体记忆框架。具体而言,ViSAGE通过跨模态绑定在长时序范围内锚定实体身份,随后应用双向记忆精调传播延迟的实体证据,回溯性统一历史记录并改进未来推理。我们还引入多智能体交叉验证,在身份-证据对齐约束下评估检索证据,当证据缺失时支持弃权(不执行)而非无依据答案。大量结果表明,ViSAGE始终优于最强基线,准确率提升5.9%。

英文摘要

Multimodal agents operating in long-horizon environments must build and continually update multimedia memories to support entity-consistent, temporally grounded reasoning. However, existing agentic memory approaches often discard fine-grained dentity cues under aggressive compression and segment-wise processing. They also rely heavily on vector similarity retrieval, which can surface semantically related yet identity-mismatched evidence, leading to entity confusion, error propagation, and hallucinated answers. We propose ViSAGE, a multimodal agentic memory framework that constructs self-correcting, entity-centric memories. Specifically, ViSAGE anchors entity identity via cross-modal binding over long temporal ranges. It then applies bidirectional memory refinement to propagate delayed identity evidence, retroactively unifying historical records and improving future reasoning. We also introduce multi-agent cross-verification to assess retrieved evidence under an identity-evidence alignment onstraint, enabling abstention instead of unsupported answers when evidence is missing. Extensive results demonstrate that ViSAGE consistently outperforms the strongest baseline, achieving 5.9% higher accuracy.

CommentsAccept by ACMMM 2026

DOI:10.1145/3767308.3835852

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑