作为辅助监督的生成:通过解耦嵌入预测以零推理开销增强视觉理解
Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction
- ByteDance(字节跳动)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究提出GAS框架,将视觉生成作为辅助监督,通过解耦MoT架构的NEP实现零推理开销,提升了多模态理解尤其是感知与空间理解能力。
AI中文摘要:
尽管多模态大语言模型(Multimodal Large Language Models, MLLMs)已取得显著进展,但视觉理解与生成通常被视为不同的目标。现有统一框架常依赖离散视觉标记化或扩散目标,其生成目标与视觉理解模型所用的连续表示存在差异,导致直接迁移以增强现有预训练MLLMs并非易事。本研究提出GAS,一种生成引导的训练框架,将视觉生成重新解释为表示学习的辅助监督。具体而言,GAS在解耦的混合专家Transformer(Mixture-of-Transformers, MoT)架构内采用下一个嵌入预测(Next Embedding Prediction, NEP)作为跨模态生成范式。通过维护共享的底层主干与并行的上层网络,GAS使生成损失以更精细的空间精度和更强的视觉保留能力丰富共享视觉通路,同时屏蔽上层理解层免受直接生成梯度影响。为最大化该协同效应,我们进一步构建高度相关的生成任务,这些任务需深度认知基础而非仅通用合成。在不同模型规模和训练阶段,GAS提升了整体多模态理解能力,其最可靠的增益体现在感知与空间理解方面。关键在于,由于辅助生成分支在训练后被丢弃,这些增益不会产生额外推理开销。大量受控对比实验和表示层面分析进一步阐明了生成引导训练何时及为何有益于理解,并证明生成引导训练是实现更强多模态理解的可行实用路径。
英文摘要:
While Multimodal Large Language Models (MLLMs) have achieved remarkable progress, visual understanding and generation are typically treated as divergent objectives. Existing unified frameworks often rely on discrete visual tokenization or diffusion objectives whose generative targets differ from the continuous representations consumed by visual understanding models, making direct transfer to enhance existing pretrained MLLMs non-trivial. In this work, we present GAS, a generation-guided training framework that reinterprets visual generation as auxiliary supervision for representation learning. Concretely, GAS adapts Next Embedding Prediction (NEP) as a cross-modal generation paradigm within a decoupled Mixture-of-Transformers (MoT) architecture. By maintaining a shared lower trunk and parallel upper layers, GAS lets generation losses enrich the shared visual pathway with finer spatial precision and stronger visual retention while shielding the upper understanding layers from direct generation gradients. To maximize this synergy, we further construct highly correlated generation tasks that demand deep cognitive grounding rather than generic synthesis alone. Across model scales and training stages, GAS improves aggregate multimodal understanding, with its most reliable gains on perception and spatial comprehension. Crucially, because the auxiliary generation branch is discarded after training, these gains incur zero inference overhead. Extensive controlled comparisons and representation-level analyses further clarify when and why generation-guided training benefits understanding, and demonstrate the feasibility of generation-guided training as a practical route to stronger multimodal understanding.