发表机构
College of AI, Tsinghua University(清华大学人工智能学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出首个开源大规模交互式视听数据集InteracVid,解决现有生成模型监督信号缺失问题,实验证实其可提升多模态助手的交互规划与响应生成性能,为交互式多模态生成提供关键基础。
AI 中文摘要
大型语言模型已使文本成为人与AI交互的默认媒介,但仅文本无法表达多模态助手、虚拟化身及具身智能体所需的全部响应内容。尽管近期的视听生成模型可合成高保真同步内容,现有监督信号却多为描述性的:模型被训练生成字幕,而非产生由外部用户交互触发的视听响应。我们提出InteracVid——首个开源大规模数据集,用于解决这一缺失的监督问题,每个样本均包含前置视听上下文、外部刺激及后续真实交互式响应。我们设计了一种感知元数据的流水线,从冗长嘈杂的直播流中提取交互片段,最终从超过5.9万个直播视频中生成了超过45.4万个上下文-查询-响应三元组,覆盖以对话为中心、以对象为中心、流程式、具身及基于屏幕的各类场景。十名标注员参与的人类研究证实,提取的交互对真实及重构查询而言均具有因果性、自然性且时间上完整。在包含100个真实直播聊天查询的保留基准上,基于InteracVid进行微调可同时提升交互规划与视听响应生成性能,独立人类评估再现了系统排名及自动评判器得出的结论。这些结果表明,交互结构化数据是交互式多模态生成的关键基础。
英文摘要
Large language models have made text the default medium for human--AI interaction, buttext alone cannot express the full range of responses required by multimodal assistants,avatars, and embodied agents. While recent audio-video generative models can synthesizehigh-fidelity synchronized content, existing supervision is largely \emph{descriptive}:models are trained to render captions rather than to produce audio-visual responsescaused by external user interactions. We introduce \textbf{InteracVid}, \emph{the firstopen-source large-scale dataset that addresses this missing supervision}, so that everysample couples a preceding audio-visual context and an external stimulus with the realinteractive response that follows. We design a metadata-aware pipeline that extractsinteractive clips from long, noisy livestreams, yielding over \textbf{454K}context-query-response triplets from more than \textbf{59K} livestream videos andspanning conversation-centered, object-centric, procedural, embodied, and screen-basedscenarios. A ten-rater human study confirms that the extracted interactions are causal,natural, and temporally complete for both genuine and reconstructed queries. On aheld-out benchmark of \textbf{100} genuine live-chat queries, fine-tuning on InteracVidimproves both interaction planning and audio-video response generation, and anindependent human evaluation reproduces the system ranking and the conclusions obtainedwith our automatic judge. These results highlight interaction-structured data as acritical foundation for interactive multimodal generation.