arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.21465eess.AScs.AIeess.IV

OmniVChat:原生音视频对话的合成、基准测试与训练

OmniVChat: Synthesizing, Benchmarking, and Training for Native Audio-Visual Dialogue

Haolin He, Yunfei Chu, Qi Chen, Wen Huang, Yuan Feng, Muzhi Zhu, Zheqi Dai, Haoning Xu, Dongchao Yang, Chunyat Wu, Zining Liang, Zhengxi Liu, Xiquan Li, Xie Che… 展开作者

Haolin He, Yunfei Chu, Qi Chen, Wen Huang, Yuan Feng, Muzhi Zhu, Zheqi Dai, Haoning Xu, Dongchao Yang, Chunyat Wu, Zining Liang, Zhengxi Liu, Xiquan Li, Xie Chen, Xize Cheng, Qize Yang, Jin Xu, Qiuqiang Kong

首次发表
浏览论文内容

中文总结 AI 辅助

OmniVChat定义原生音视频对话任务,提出多智能体数据引擎合成对话构建基准,并设计强化学习奖励,提升全模态模型对话能力。

中文摘要 AI 辅助

我们将OmniVChat(全模态视频聊天)定义为用户与全模态模型之间的原生音视频对话任务。在OmniVChat中,全模态模型直接同时接收用户的音频和视频,并返回文本。用户的查询嵌入在音频和视频中,没有单独的文本问题、外部字幕或语音识别。直接音视频输入减少了外部延迟和计算量,同时保留了感知线索。然而,OmniVChat的研究面临两个限制:数据可用性和评估。人们使用自己设备录制的数据稀缺。此外,一个好的回复通常需要考虑用户的周围环境、面部表情和附近物体,并且此类回复可以以多种不同方式表达,使得关键词匹配在评估回复质量时不可靠。智能体系统和视频生成的最新进展使得为理解而生成变得可行,这意味着使用合成对话进行训练和评估。因此,我们提出了OmniVChat-Studio,一个用于合成单轮和多轮音视频对话的多智能体数据引擎。我们使用合成对话构建了OmniVChat-Bench,一个评估基准,用于评估全模态模型在五个能力类别中的基本对话能力。我们还提出了OmniVChat-RL,一种强化学习奖励设计,共同针对OmniVChat中的回复正确性、效率和风格。使用OmniVChat-RL在合成对话上训练Qwen3-Omni-Instruct,提高了其在OmniVChat-Bench和人工录制的OmniVChat-Bench-Human上的性能。这些收益验证了奖励设计,并展示了在训练和评估中向真实世界对话的迁移。

英文摘要

We define OmniVChat (Omni Video Chat) as the task of native audio-visual dialogue between a user and an omni model. In OmniVChat, omni models directly and simultaneously receive audio and video from a user and return text. The user's query is embedded in the audio and video, without a separate text question, external captioning, or speech recognition. Direct audio-visual input reduces external latency and computation while preserving perceptual cues. However, research on OmniVChat faces two constraints: data availability and evaluation. Recordings of people using their own devices are scarce. Furthermore, good replies often depend on multimodal context and can be phrased in many ways, making keyword matching unreliable for evaluation. Recent progress in agent systems and video generation makes generation for comprehension viable, which means using synthesized dialogues for training and evaluation. Therefore, we present OmniVChat-Studio, a multi-agent data engine for synthesizing single- and multi-turn audio-visual dialogues. We use synthesized dialogues to build OmniVChat-Bench, an evaluation benchmark that evaluates omni models' basic dialogue abilities across five ability categories. Replies are judged by a large language model based on explicit scoring criteria. We also present OmniVChat-RL, a reinforcement learning reward design that jointly targets reply correctness, efficiency, and style in OmniVChat. Training Qwen3-Omni-Instruct with OmniVChat-RL on synthesized dialogues improves its performance on both OmniVChat-Bench and the human-recorded OmniVChat-Bench-Human. These gains validate the reward design and show transfer to real-world dialogues in training and evaluation.

发表机构

  • The Chinese University of Hong Kong(香港中文大学)
  • Alibaba Token Hub, Alibaba Group(阿里巴巴集团通义实验室)
  • Shanghai Jiao Tong University(上海交通大学)
  • Shanghai Innovation Institute(上海创新研究院)
  • Zhejiang University(浙江大学)

机构由 AI 辅助整理,请以论文原文为准。

↑