发表机构
University of Science and Technology of China; Alibaba Token Hub, Alibaba Group; The Chinese University of Hong Kong; Zhejiang University(中国科学技术大学; 阿里巴巴集团,阿里巴巴Token Hub; 香港中文大学; 浙江大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
VisionWeave通过门控空间池化器和粒度路由器,使多模态大模型能自适应分配视觉表示,在保持98.9%性能的同时平均节省43%令牌,并提升2.3倍吞吐量。
AI 中文摘要
多模态大语言模型已成为视觉理解的主导范式,但通过将输入编码为密集、固定大小的补丁令牌而产生了大量成本。然而,视觉信息分布不均:某些区域需要细粒度细节,而其他区域则可接受紧凑表示。下采样牺牲了这种细节,而现有的令牌剪枝和自适应方法在内容自适应粒度、任务泛化以及与现代MLLMs和服务基础设施的集成方面仍然有限。克服这些限制需要基础模型以端到端方式学习在何处以及以何种粒度分配视觉表示,我们将这种原生能力称为弹性视觉表示编织。我们引入了VisionWeave,通过大规模训练在前沿级MLLMs中建立这一能力。它结合了两个组件:一个门控空间池化器在共享的MRoPE坐标内构建粗粒度表示以及原生细粒度表示,而一个粒度路由器学习它们的内容自适应分配。仅通过自蒸馏,我们在Qwen3.5-4B上验证了这一能力,并扩展到Qwen3.8-27B,使用了超过30K A100 GPU小时。基于Qwen3.8-27B,VisionWeave自适应调整令牌节省以适应视觉内容,平均节省43.0%的令牌,同时在八个基准上保持98.9%的原生性能,而固定50%节省目标的令牌剪枝基线仅保留88%的性能。广泛评估确认了在不同任务、分辨率和视频帧上的稳健效率-质量权衡。当部署在SGLang服务引擎上时,我们的方法实现了2.3倍的吞吐量提升,同时将平均TTFT降低了54.4%,平均TPOT降低了60.6%。总之,我们相信这些结果将弹性视觉编织定位为下一代多模态模型的有前景能力。
英文摘要
Multimodal large language models have become the dominant paradigm for visual understanding, but incur substantial costs by encoding inputs into dense, fixed-size patch tokens. However, visual information is unevenly distributed: some regions require fine-grained detail, while others admit compact representations. Downsampling sacrifices this detail, while existing token pruning and adaptive approaches remain limited in content-adaptive granularity, task generalization, and integration with modern MLLMs and serving infrastructure. Overcoming these limitations calls for foundation models that learn, end to end, where-and at what granularity-to allocate visual representations, a native capability we term elastic visual representation weaving. We introduce VisionWeave, establishing this capability in frontier-level MLLMs through large-scale training. It combines two components: a gated spatial pooler constructs coarse-grained representations alongside native fine-grained representations within a shared MRoPE coordinate, while a granularity router learns their content-adaptive allocation. Through self-distillation alone, we validate this capability on Qwen3.5-4B and scale to Qwen3.8-27B with over 30K A100 GPU-hours. Based on Qwen3.8-27B, VisionWeave adaptively adjusts token savings to visual content, saving 43.0% tokens on average while retaining 98.9% native performance across eight benchmarks, versus only 88% performance preserved for token pruning baselines with a fixed 50% savings target. Extensive evaluations confirm robust efficiency-quality trade-offs across diverse tasks, resolutions and video frames. When deployed on SGLang serving engine, our method achieves a 2.3x throughput gain while reducing mean TTFT by 54.4% and mean TPOT by 60.6%. Together, we believe these results position elastic visual weaving as a promising capability for next-generation multimodal models.