arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.00685cs.MA

M2Note: 通过错误笔记本学习实现视觉语言模型的持续演化

M2Note: Continual Evolution of Vision Language Models via Mistake Notebook Learning

Haiwen Li, Jing Tang, Rui Chen, Lei Sun, Xiangxiang Chu

首次发表
浏览论文内容

中文总结 AI 辅助

提出M2Note,一种无需训练的持续演化框架,通过将失败轨迹转化为可编辑的笔记,利用多模态检索增强生成在推理时引导模型,避免重复错误,并在多个基准上提升性能。

中文摘要 AI 辅助

视觉语言模型(VLM)在多模态推理任务中展现出显著能力,但仍存在反复失败,例如跳过关键视觉检查、误用领域规则和幻觉不支持的概念。现有解决方案大多依赖监督微调(SFT)和强化学习(RL),这些方法迭代成本高,且在分布偏移下可能脆弱。为此,我们提出多模态错误笔记本学习(M2Note),一种无需训练的持续演化框架,将学习外部化为可编辑的记忆。M2Note将失败轨迹转化为紧凑的主题-指导笔记:主题总结潜在领域和概念,指导提供可在未来推理中复用的可操作验证步骤。测试时,M2Note通过多模态检索增强生成(RAG)检索相关笔记并将其附加到模型上下文中,引导推理远离先前观察到的陷阱。为了稳定持续演化,我们采用带有回滚的批量级后验证,仅当笔记编辑在同一批次上提升性能时才提交,减少噪声更新并防止回归。M2Note支持自我演化(同一VLM作为求解器和监督者)和跨模型演化(更强的监督者指导较弱的求解者),无需权重更新即可实现能力迁移。在六个多模态推理基准上的实验表明,该方法在领域和骨干网络上均有一致改进,同时实现了强大的成本和样本效率,并与思维链(CoT)提示互补。

英文摘要

Vision Language Models (VLMs) have demonstrated remarkable capabilities in multimodal reasoning tasks, yet they still suffer from recurring failures, such as skipping key visual checks, misapplying domain rules, and hallucinating unsupported concepts. Most existing solutions rely on supervised fine-tuning (SFT) and reinforcement learning (RL), which are expensive to iterate and can be brittle under distribution shift. To this end, we propose Multimodal Mistake Notebook Learning (M2Note), a training-free continual evolution framework that externalizes learning into an editable memory. M2Note transforms failed trajectories into compact subject-guidance notes: the subject summarizes the underlying domain and concept, while the guidance provides actionable verification steps that can be reused in future inference. At test time, M2Note retrieves relevant notes via multimodal retrieval-augmented generation (RAG) and appends them to the model context, steering reasoning away from previously observed pitfalls. To stabilize continual evolution, we adopt batch-level post-verification with rollback, which commits notebook edits only if they improve performance on the same batch, reducing noisy updates and preventing regressions. M2Note supports both self-evolving, where the same VLM acts as solver and supervisor, and cross-model evolving, where a stronger supervisor guides a weaker solver, enabling capability transfer without weight updates. Experiments on six multimodal reasoning benchmarks show consistent improvements across domains and backbones, while achieving strong cost and sample efficiency and remaining complementary to Chain-of-Thought (CoT) prompting.

发表机构

  • Beijing University of Posts and Telecommunications(北京邮电大学)
  • AMAP, Alibaba Group(高德,阿里巴巴集团)

机构由 AI 辅助整理,请以论文原文为准。

↑