arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Mage-VL:一种高效的编解码器原生流式多模态基础模型

Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model

Senqiao Yang, Kaichen Zhang, Zhaoyang Jia, Jinghao Guo, Yifei Shen, Xinjie Zhang, Xiaoyi Zhang, Haoqing Wang, Xiao Li, Peng Zhang, Xiang An, Yin Xie, Zhening Liu, Xun Guo, Jiahao Li, Shicheng Zheng, Jinglu Wang, Zongyu Guo, Wenxuan Xie, Zihan Zheng, Yuxuan Luo, Bin Li, Yan Lu

arXiv 2607.24904首次发表:更新:

发表机构

Microsoft(微软)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对标准视觉语言模型在流式感知任务的不足,提出Mage-VL。核心方法是用Mage-ViT选择性编码关键区域,减少视觉令牌消耗。该模型在多模态理解和交互上表现出色,推理速度加快,还有多项关键实证发现。

AI 中文摘要

标准视觉语言模型(VLMs)存在莫拉维克悖论:在复杂离线视觉推理方面表现出色,但在简单流式感知任务中存在困难且处理效率低下。我们提出了Mage-VL,一种用于实时多模态理解和交互的高效编解码器原生流式基础模型。其核心的自定义分词器Mage-ViT通过使用运动向量和跨稀疏锚(I)和预测(P)帧的残余能量选择性地编码动态、高熵区域,取代均匀帧采样。在16x16补丁级别运行,减少了超过75%的视觉令牌消耗,同时保留时空上下文。在约560M未标记图像和100M未标记视频帧上从头开始训练,Mage-ViT匹配或优于在数十亿图像-文本对上训练的旗舰编码器。我们建立了AI4AI数据管道,包括用于多模态字幕的提示-代码联合优化和人工智能驱动的性能诊断以指导训练方法。此外,通过受生物启发的双系统架构——轻量级系统1事件门和因果系统2解码器,Mage-VL实现了主动流式感知。广泛评估表明,Mage-VL-4B在静态任务上与Qwen3-VL-4B相当,在视频理解和2D/3D空间推理方面有显著提升,推理速度加快高达3.5倍,全面超越15B Phi-4推理视觉基线。我们还给出了七个关键实证发现。

英文摘要

Standard vision-language models (VLMs) suffer from Moravec's paradox: they excel at complex offline visual reasoning but struggle with simple streaming perception tasks and process them inefficiently. We present Mage-VL, an efficient codec-native streaming foundation model for real-time multimodal understanding and interaction. At its core, our custom tokenizer, Mage-ViT, replaces uniform frame sampling by selectively encoding dynamic, entropy-rich regions using motion vectors and residual energy across sparse anchor (I) and predicted (P) frames. Operating at a 16 x 16 patch level, this reduces visual token consumption by over 75% while preserving spatiotemporal context. Trained from scratch on approximately 560M unlabeled images and 100M unlabeled video frames, Mage-ViT matches or outperforms flagship encoders trained on billions of image-text pairs. We establish AI4AI data pipelines encompassing prompt-code joint optimization for multimodal captioning and AI-driven performance diagnosis to guide training recipes. Furthermore, through a bio-inspired dual-system architecture - a lightweight System 1 event gate and a causal System 2 decoder - Mage-VL enables proactive streaming perception. Extensive evaluations show that Mage-VL-4B matches Qwen3-VL-4B on static tasks while achieving strong gains in video understanding and 2D/3D spatial reasoning, with up to a 3.5x wall-clock inference speedup, and comprehensively surpasses the 15B Phi-4-reasoning-vision baseline. Beyond model artifacts, we deliver seven key empirical findings covering pre-training data efficiency, variable-resolution scaling, codec system acceleration, VideoQA SFT redundancy, motion-spatial synergy, AI4AI data pipelines, and Zero-Vision SFT for multimodal RL.

CommentsProject page: https://microsoft.github.io/Mage

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑