AI 中文总结
本研究提出开源视觉-语言模型MOSS-VL,通过全栈协同设计实现实时交互能力,在流式基准测试中表现优异,大幅领先同类开源模型,发布了相关检查点与代码。
AI 中文摘要
我们提出了MOSS-VL,这是一个开源视觉-语言模型家族,将实时交互(边说话边感知)视为核心能力。该模型在全栈层面协同设计:语言解码器仅通过门控交叉注意力关注视觉信息,因此模型在生成时可自然感知输入帧;合成的交互语料库用于监督何时说话、何时保持沉默以及何时修正;分阶段课程将所有实时特定训练集中在强大离线基础之上的一个轻量最终阶段。离线场景下,MOSS-VL-Instruct在可比规模上具有竞争力,且在时间推理视频数据集上表现领先。在四个流式基准测试中,MOSS-VL-Realtime在开源流式模型中,三个基准的平均性能最佳(第四个基准排名第二),在严格测试主动行为的三个子集上大幅领先——在OmniMMI主动警报任务中,其得分66.0,而最佳基线得分仅为37.5。MOSS-VL拥有113亿参数,且视觉令牌位于解码序列之外,随着视觉上下文增长,其与同骨干的Qwen3-VL-8B相比,首次生成令牌时间的优势从2.8倍扩大至5.1倍。我们发布了全部五个检查点、训练课程以及实时推理代码,链接为:https://arxiv.org/abs/...(注:原链接保留)
英文摘要
We present MOSS-VL, an open vision-language model family that treats real-time interaction -- perceiving while it speaks -- as a first-class capability. It is co-designed across the stack: the language decoder attends to vision only through gated cross-attention, so the model can naturally see incoming frames while generating; a synthesized interaction corpus supervises when to speak, when to stay silent, and when to revise; and a staged curriculum concentrates all real-time-specific training in one light final stage over a strong offline foundation. Offline, MOSS-VL-Instruct is competitive at comparable scale and leads temporal-reasoning video sets. Across four streaming benchmarks, MOSS-VL-Realtime posts the best average on three (second on the fourth) among open-source streaming models, sweeping the three subsets that squarely test proactive behavior -- 66.0 vs. 37.5 for the best baseline on OmniMMI Proactive Alerting. With 11.3B parameters but visual tokens outside the decoded sequence, MOSS-VL widens its time-to-first-token advantage over same-backbone Qwen3-VL-8B from 2.8x to 5.1x as visual context grows. We release all five checkpoints, the training curriculum, and the real-time inference code at https://github.com/OpenMOSS/MOSS-VL.
Comments22 pages. Project page: https://openmoss.ai/MOSS-VL/