AI 中文总结
LoopVL将循环Transformer扩展到视觉-语言模型,通过模块与模型循环迭代更新统一状态,在多模态理解和视觉推理基准上超越非循环模型,并观察到视觉注意力显著转变的“视觉顿悟时刻”。
AI 中文摘要
我们引入LoopVL,以研究循环Transformer能否有效扩展到视觉-语言模型。LoopVL结合了模块循环和模型循环计算,通过共享模块迭代更新统一的视觉-语言状态。我们从零开始训练LoopVL,包括语言预训练、多模态训练和后训练。在多模态理解和视觉推理基准上,LoopVL优于一系列规模相近和更大的非循环模型。我们还观察到LoopVL中的“视觉顿悟时刻”,其特征是跨循环的视觉注意力发生显著转变。LoopVL为循环视觉-语言建模提供了实践证据,并为共享参数如何支持在不断演化的视觉-语言状态上进行更深层次的多模态计算提供了直观视角。
英文摘要
We introduce LoopVL to study whether Loop Transformers can be effectively extended to vision- language models. LoopVL combines Module-Loop and Model-Loop computation to iteratively update a unified vision-language state through shared modules. We train LoopVL from scratch through language pre-training, multimodal training, and post-training. LoopVL outperforms a range of similarly sized and larger non-recurrent models on multimodal understanding and visual reasoning benchmarks. We also observe Visual Aha Moments in LoopVL, characterized by pronounced shifts in visual attention across loops. LoopVL provides practical evidence for recurrent vision-language modeling and offers an intuitive perspective on how shared parameters can support deeper multimodal computation over continuously evolving visual-language states.