arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.10720cs.AIcs.CLcs.CV

Ex-Omni-2D:具备原生视觉存在的高表达全模态对话模型

Ex-Omni-2D: Expressive Omni-Modal Dialogue Models with Native Visual Presence

Haoyu Zhang, Zhipeng Li, Xiaoying Tang, Tianshu Yu, Yiwen Guo

首次发表
浏览论文内容

中文总结 AI 辅助

Ex-Omni-2D是一种全模态对话框架,可生成含文本、个性化语音与参考条件视频的协同响应,通过特定机制实现高效增量生成,在指定分辨率下达成1.293的端到端RTF,提供实用的质量效率平衡点。

中文摘要 AI 辅助

全模态对话模型可理解多模态输入并生成语音回复,但其响应仍缺乏视觉实体感。本文提出Ex-Omni-2D,这是一种全模态对话框架,可生成包含文本、个性化语音及参考条件视频的协同响应。给定多模态查询、参考图像和参考音频,该模型会预测结构化的视觉思考计划(VTP),描述场景、情感和动作,随后生成响应文本及原生多码本语音单元。这些单元构成共享的声学-时间接口:它们被解码为语音并与视频帧在线对齐。该接口使响应和化身通路能从异构的语音、对话及化身视频数据中学习,无需大规模查询-文本-语音-视频监督。全序列视频生成器作为主教师模型,为实现高效增量生成,我们进一步将其蒸馏为少步块因果流式学生模型,其前缀流式机制在连续块间传递干净隐变量,以减少后期块的累积退化。采用四步推理时,完整的四GPU流水线在分辨率为400×720/720×400时实现1.293的端到端实时因子(RTF),提供了实用的质量-效率操作点。

英文摘要

Omni-modal dialogue models can understand multimodal inputs and synthesize spoken replies, but a spoken answer still leaves the agent visually absent. We introduce \textbf{Ex-Omni-2D}, a framework that answers a multimodal query with coordinated text, personalized speech, and reference-conditioned video. The dialogue model first writes a structured \textit{Visual Thought Plan} (VTP) for scene, emotion, and motion, then generates the response text and multi-codebook speech units. These speech units are decoded into audio and aligned with video frames, giving the speech and avatar modules a common timing signal while allowing them to learn from different data sources. The video module is trained as a full-sequence Teacher conditioned on reference appearance, VTP semantics, and frame-aligned speech units. We further explore to distill it into a few-step block-causal \emph{Streaming Student}; its Prefix Streaming mechanism carries the previous clean latent into the next chunk and is analyzed as a partial mitigation for late-chunk subject drift. At $400\times720$/$720\times400$, the four-step four-GPU Student provides incremental output with lower startup latency than the full-sequence Teacher.

↑