arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Where-OPD:基于合成场景的多模态大语言模型空间引导在线策略自蒸馏

Where-OPD: Spatially Guided On-Policy Self-Distillation of MLLMs with Synthetic Scenes

Sophia Sirko-Galouchenko, Monika Wysoczanska, Andrei Bursuc, Nicolas Thome, Spyros Gidaris

arXiv 2610.02117首次发表:更新:

发表机构

Sorbonne Université; CNRS; ISIR; Institut universitaire de France (IUF); ILLS(索邦大学; 法国国家科学研究中心; 智能系统与机器人研究所; 法国大学研究院; 国际学习系统研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出Where-OPD方法,通过空间定位的文本指导进行在线策略自蒸馏,利用合成场景训练多模态大模型,提升计数与文档理解,并实现3.23点真实基准性能提升。

AI 中文摘要

在线策略自蒸馏近期已成为一种提升语言模型推理能力的有效方法,其通过让学生模型监督自身的一个冻结或指数移动平均(EMA)版本(该版本接收特权信息)来实现。然而,其在多模态大语言模型(MLLMs)中的应用仍 largely 未被探索。近期方法利用特权视觉信息(例如与问题对应的图像裁剪)来改善细粒度感知,但其收益局限于受益于此类视觉缩放的特定任务,并且需要人工标注的定位数据或外部教师模型。我们引入了一种不同形式的用于多模态大语言模型的在线策略自蒸馏,该方法为教师模型提供文本形式的、空间定位的指导,以识别与查询相关的视觉元素。我们使用程序化生成的场景,这些场景自动提供对象身份和空间坐标,从而实现可扩展且无需标注的后训练。教师模型利用这种空间指导来定位并整合来自多个相关图像区域的证据,而学生模型则学习仅从图像和问题中复现由此产生的行为。我们的方法在多个模型上的计数、文档和图表理解基准中持续提升了性能。重要的是,尽管后训练仅使用合成场景,但由此产生的改进能够迁移到真实世界的感知基准,在CVBench、V*、ZoomBench、BLINK、HR-Bench和MME-RealWorld上平均性能提升了3.23个百分点。这些结果表明,空间定位的特权信息可以通过在线策略自蒸馏引发更广泛的感知能力,实现超越后训练所用任务和数据分布的实质性合成到真实迁移。项目页面:此 https URL

英文摘要

On-policy self-distillation has recently emerged as an effective approach for improving language-model reasoning by supervising students with a frozen or EMA version of themselves that receives privileged information. Its application to multimodal large language models (MLLMs), however, remains largely unexplored. Recent approaches use privileged visual information, such as image crops corresponding to a question, to improve fine-grained perception, but their gains are confined to tasks that benefit from such visual zooming and require either human-annotated grounding data or external teacher models. We introduce a different form of on-policy self-distillation for MLLMs that provides the teacher with textual, spatially grounded guidance identifying the visual elements relevant to a query. We use procedurally generated scenes with automatically available object identities and spatial coordinates, enabling scalable and annotation-free post-training. The teacher uses this spatial guidance to locate and integrate evidence from multiple relevant image regions, while the student learns to reproduce the resulting behavior from the image and question alone. Our approach consistently improves performance on counting, document and chart understanding benchmarks across multiple models. Importantly, although post-training uses only synthetic scenes, the resulting improvements transfer to real-world perception benchmarks, yielding a 3.23-point gain in average performance across CVBench, V*, ZoomBench, BLINK, HR-Bench, and MME-RealWorld. These results show that spatially grounded privileged information can induce broader perceptual capabilities through on-policy self-distillation, enabling substantial synthetic-to-real transfer beyond the task and data distribution used for post-training. Project page: https://github.com/sirkosophia/Where-OPD

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑