arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.15863cs.CV

LynnReal-Omni:面向智能体视觉工作流的原生多模态视频生成

LynnReal-Omni: Native multi-modal Video Generation for Agentic Visual Workflows

发表机构LynnReal AI
查看机构详情
  • LynnReal AI

机构由 AI 辅助整理,请以论文原文为准。

Xiaofeng Mao, Peijia Lin, Shaohao Rui, Yibo Zhang, Haibin Wan, Weijie Ma

首次发表
浏览论文内容

中文总结 AI 辅助

LynnReal-Omni提出统一多模态视频生成框架,结合智能体参考控制与扩散模型,实现稳定高质量生成,并通过Flash加速实现实时渲染。

中文摘要 AI 辅助

视频扩散模型具有随机性且难以控制:精确的内容往往需要反复采样,且无法保证成功,而长时程场景在外观、交互和时间连贯性上会出现漂移。智能体视觉创作提供了明确的参考、可编辑的3D场景或可执行的游戏状态以实现稳定控制,但本身并不能保证高保真的物体或角色一致性。将两者结合可以实现稳定、高质量的生成。为实现这一结合,我们提出了LynnReal-Omni,一个基于32B共享多模态扩散Transformer的原生多模态视频生成框架,统一了文本到视频、图像条件生成、参考引导生成、结构控制、编辑、退化视频修复和长视频生成。它接受异构视觉输入,包括外观参考、可编辑的3D渲染和游戏录制,使智能体能够在统一模型内组合视觉条件。我们还训练了一个专门的27B Flash共享多模态扩散Transformer用于实时渲染。我们构建了一个系统化的数据管道,用于视频清理、主体关联、多模态标注和对齐控制构建,生成了一个多镜头视听片段的精选语料库,并引入了MSAVP,一个包含100个提示、20个指标的评估设计,分别评估指令遵循、生成合理性、视觉质量、时间行为和音频协调。LynnReal-Omni-Flash通过模型和解码加速进一步降低了推理成本,包括一个轻量级VAE解码器;在一张H100上,使用LynnReal-Omni生成并解码22帧540p视频的暖启动耗时843毫秒,而使用Flash则为377毫秒。这些结果为实时流式视频生成奠定了基础,使LynnReal-Omni成为智能体视觉创作的统一、可控且高效的基础。

英文摘要

Video diffusion models are stochastic and hard to control: precise content often requires repeated sampling without guaranteed success, and long-horizon scenes drift in appearance, interactions, and temporal coherence. Agentic visual creation provides explicit references, editable 3D scenes, or executable game states for stable control, but does not by itself guarantee high object or character fidelity. Combining the two can enable stable, high-quality generation. To realize this combination, we present LynnReal-Omni, a native multimodal video generation framework built on a 32B shared multimodal diffusion transformer that unifies text-to-video, image-conditioned generation, reference-guided generation, structural control, editing, degraded video restoration, and long-video generation. It accepts heterogeneous visual inputs, including appearance references, editable 3D renders, and game recordings, allowing agents to compose visual conditions within a unified model. We also train a dedicated 27B Flash shared multimodal diffusion transformer for real-time rendering. We build a systematic data pipeline for video cleaning, subject association, multimodal annotation, and aligned control construction, yielding a curated corpus of multi-shot audiovisual segments, and introduce MSAVP, a 100-prompt, 20-metric evaluation design that separates instruction following, generating plausibility, visual quality, temporal behavior, and audio coordination. LynnReal-Omni-Flash further reduces inference cost through model and decoding acceleration, including a lightweight VAE decoder; on one H100, warm generation and decoding of a 22-frame 540p video take 843 ms with LynnReal-Omni and 377 ms with Flash. These results provide a foundation for real-time streaming video generation, making LynnReal-Omni a unified, controllable, and efficient basis for agentic visual creation.

↑