MOSAIC:用于高效异构视觉语言模型的自适应层间组合
MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models
浏览论文内容
中文总结 AI 辅助
研究针对视觉语言模型中异构结构依赖手工静态混合模式的问题,提出硬件感知搜索方法MOSAIC,通过多目标混合整数规划确定最优配置,并引入两阶段参数恢复过程,提升了模型推理效率且降低训练成本。
中文摘要 AI 辅助
视觉语言模型(VLMs)使用同构Transformer处理多媒体数据取得了成功。近期研究表明,交织线性注意力等高效机制的异构结构在性能和推理延迟上优于同构设计。然而,这些方法依赖手工静态混合模式,次优且难适配特定硬件。为此提出MOSAIC,一种硬件感知搜索方法,将多种效率机制整合到统一搜索空间,通过多目标混合整数规划确定最优配置。为减轻结构转换的性能下降,引入两阶段参数恢复过程。通过MOSAIC - 4B验证,结果表明其在多基准测试中性能与基线匹配,训练成本不到原来的2%,推理效率大幅提升。
英文摘要
Vision-Language Models (VLMs) have achieved success using homogeneous Transformers to process multimedia data. Recent studies show that heterogeneous structures interleaving efficient mechanisms, like linear attention, improve both performance and inference latency over homogeneous designs. However, these efforts rely on handcrafted static mixing patterns, which are sub-optimal and difficult to adapt to specific hardware. To bridge this gap, we propose Multi-Objective Search for Adaptive Inter-layer Composition (MOSAIC), a hardware-aware search method that automatically transforms homogeneous models into optimized heterogeneous architectures. MOSAIC integrates diverse efficiency mechanisms--including linear, sparse, and low-rank operators--into a unified search space. By formulating the selection as a multi-objective Mixed Integer Programming (MIP) problem, our method identifies optimal configurations that maximize downstream performance under strict hardware latency constraints. To mitigate performance degradation from structural transitions, we introduce a two-stage parameter recovery process: global off-policy distillation to stabilize internal representations, followed by a dual-teacher on-policy distillation leveraging a 235B oracle for knowledge expansion and the original 4B teacher for distributional stability. We validate MOSAIC through MOSAIC-4B, derived from Qwen3-VL-4B-Instruct. Results demonstrate that MOSAIC-4B matches the baseline's performance across multiple benchmarks while requiring less than 2% of the original training cost. Furthermore, it substantially improves inference efficiency, achieving 1.76x prefilling and 2.54x decoding speedups.
发表机构
- LiAuto Inc.(理想汽车公司)
机构由 AI 辅助整理,请以论文原文为准。