arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Qwen3.8-Omni:迈向原生全模态智能体

Qwen3.8-Omni: Towards Native Omni-Modal Agents

Qwen Team

arXiv 2609.25611首次发表:更新:

AI 中文总结

本文提出原生多模态智能体模型 Qwen3.8-Omni-Flash,通过协同训练和百万级上下文窗口提升理解、推理与长时程任务能力,并配套开源插件框架和实时智能体框架,支持视频编辑等生产应用。

AI 中文摘要

我们推出 Qwen3.8-Omni-Flash,一个面向真实世界多模态生产力的原生多模态智能体模型。与以往主要强调感知和交互的全模态模型相比,Qwen3.8-Omni-Flash 显著提升了多模态理解与推理能力,以及在长时程智能体任务上的表现。这些能力得益于一种原生多模态协同训练策略,该策略在保留强大文本域能力的同时,促进了智能体能力从文本任务向音频和视频任务的迁移。该模型继承了 Qwen3.8-Next 的稀疏混合专家(MoE)架构,并将上下文窗口扩展至一百万 token,支持长上下文多模态推理和长时程规划。这些进展使其能够作为主智能体或专门的子智能体集成到生产工作流中,支持视频编辑、长音频和长视频翻译、音乐驱动的音乐视频或电影生成,以及基于视频的笔记或全技能创建。为解决现有智能体框架中缺乏原生音频和视频支持的问题,我们发布了 Qwen-MM-Plugins,一个用于多模态生产力的轻量级开源插件框架。我们进一步将实时多模态交互视为一个系统级挑战,需要协调上下文与记忆管理、工具使用和子智能体委派。为此,我们发布了 Qwen-Live-Harness,一个基于 Qwen3.8-Omni-Flash 构建响应式、实时多模态智能体的开源框架。广泛评估表明,Qwen3.8-Omni-Flash 在多模态理解、推理、长时程智能体执行和视频生产力任务上均取得了强劲性能。这些结果以及配套的开源工具支持 Qwen3.8-Omni-Flash 作为在研究和生产中部署原生多模态智能体的实用基础。

英文摘要

We introduce Qwen3.8-Omni-Flash, a natively multimodal agentic model for real-world multimodal productivity. Compared with previous omni models, which primarily emphasized perception and interaction, Qwen3.8-Omni-Flash substantially improves multimodal understanding and reasoning, as well as performance on long-horizon agentic tasks. These capabilities are supported by a native multimodal co-training strategy that preserves strong text-domain capabilities while facilitating the transfer of agentic capabilities from text to audio and video tasks. The model inherits the sparse mixture-of-experts (MoE) architecture of Qwen3.8-Next and extends the context window to one million tokens, supporting long-context multimodal reasoning and long-horizon planning. These advances enable integration into production workflows as a primary agent or a specialized sub-agent, supporting video editing, long-form audio and video translation, music-conditioned music video or movie generation, and video-based note or omni-skill creation. To address the lack of native audio and video support in existing agent harnesses, we release Qwen-MM-Plugins, a lightweight open-source plugin framework for multimodal productivity. We further frame real-time multimodal interaction as a system-level challenge requiring orchestration of context and memory management, tool use, and sub-agent delegation. Accordingly, we release Qwen-Live-Harness, an open-source framework for building responsive, real-time multimodal agents based on Qwen3.8-Omni-Flash. Extensive evaluations demonstrate that Qwen3.8-Omni-Flash achieves strong performance across multimodal understanding, reasoning, long-horizon agentic execution, and video productivity tasks. These results and the accompanying open-source tools support Qwen3.8-Omni-Flash as a practical foundation for deploying natively multimodal agents in research and production.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑