发表机构
Microsoft(微软)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对标准视觉语言模型在流式感知任务的不足,提出Mage-VL。核心方法是用Mage-ViT选择性编码关键区域,减少视觉令牌消耗。该模型在多模态理解和交互上表现出色,推理速度加快,还有多项关键实证发现。
AI 中文摘要
标准视觉语言模型(VLMs)存在莫拉维克悖论:在复杂离线视觉推理方面表现出色,但在简单流式感知任务中存在困难且处理效率低下。我们提出了Mage-VL,一种用于实时多模态理解和交互的高效编解码器原生流式基础模型。其核心的自定义分词器Mage-ViT通过使用运动向量和跨稀疏锚(I)和预测(P)帧的残余能量选择性地编码动态、高熵区域,取代均匀帧采样。在16x16补丁级别运行,减少了超过75%的视觉令牌消耗,同时保留时空上下文。在约560M未标记图像和100M未标记视频帧上从头开始训练,Mage-ViT匹配或优于在数十亿图像-文本对上训练的旗舰编码器。我们建立了AI4AI数据管道,包括用于多模态字幕的提示-代码联合优化和人工智能驱动的性能诊断以指导训练方法。此外,通过受生物启发的双系统架构——轻量级系统1事件门和因果系统2解码器,Mage-VL实现了主动流式感知。广泛评估表明,Mage-VL-4B在静态任务上与Qwen3-VL-4B相当,在视频理解和2D/3D空间推理方面有显著提升,推理速度加快高达3.5倍,全面超越15B Phi-4推理视觉基线。我们还给出了七个关键实证发现。
英文摘要
Standard vision-language models (VLMs) suffer from Moravec's paradox: they excel at complex offline visual reasoning but struggle with simple streaming perception tasks and process them inefficiently. We present Mage-VL, an efficient codec-native streaming foundation model for real-time multimodal understanding and interaction. At its core, our custom tokenizer, Mage-ViT, replaces uniform frame sampling by selectively encoding dynamic, entropy-rich regions using motion vectors and residual energy across sparse anchor (I) and predicted (P) frames. Operating at a 16 x 16 patch level, this reduces visual token consumption by over 75% while preserving spatiotemporal context. Trained from scratch on approximately 560M unlabeled images and 100M unlabeled video frames, Mage-ViT matches or outperforms flagship encoders trained on billions of image-text pairs. We establish AI4AI data pipelines encompassing prompt-code joint optimization for multimodal captioning and AI-driven performance diagnosis to guide training recipes. Furthermore, through a bio-inspired dual-system architecture - a lightweight System 1 event gate and a causal System 2 decoder - Mage-VL enables proactive streaming perception. Extensive evaluations show that Mage-VL-4B matches Qwen3-VL-4B on static tasks while achieving strong gains in video understanding and 2D/3D spatial reasoning, with up to a 3.5x wall-clock inference speedup, and comprehensively surpasses the 15B Phi-4-reasoning-vision baseline. Beyond model artifacts, we deliver seven key empirical findings covering pre-training data efficiency, variable-resolution scaling, codec system acceleration, VideoQA SFT redundancy, motion-spatial synergy, AI4AI data pipelines, and Zero-Vision SFT for multimodal RL.
CommentsProject page: https://microsoft.github.io/Mage