发表机构
BAAI; PKU; Kling; THU; USTC(北京智源人工智能研究院; 北京大学; Kling; 清华大学; 中国科学技术大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对现有基准无法捕捉流式音视频生成特性的问题,推出首个综合基准StreamAV-Bench,构建含双赛道的评估框架并评估13个系统,揭示当前模型的缺陷并给出模型开发见解。
AI 中文摘要
生成模型的最新进展正推动视频生成向无界流式音视频生成发展,以适配实时交互场景。然而,现有基准测试主要评估已完成的序列,难以捕捉流式特性。为填补这一空白,我们推出StreamAV-Bench,首个专为流式音视频生成设计的综合基准测试集。StreamAV-Bench构建了统一评估框架,包含用于评估指令遵循度和长时稳定性的渐进式赛道,以及用于评估交互响应性与状态保留及复用的交互式赛道。通过覆盖32个细粒度维度的专家验证评估案例,我们对13个代表性系统开展了广泛评估。分析显示,当前模型在渐进式生成中存在时间漂移问题,在交互式控制中存在响应瓶颈。基于全面的失败分析,我们分享了推进原生联合流式音视频模型开发的见解。
英文摘要
Recent advancements in generative models are pushing video generation toward unbounded streaming audio-video generation for real-time interactive worlds. However, existing benchmarks primarily evaluate completed sequences and struggle to capture streaming properties. To bridge this gap, we introduce StreamAV-Bench, the first comprehensive benchmark tailored for streaming audio-video generation. StreamAV-Bench establishes a unified evaluation framework, including the progressive track for instruction adherence and long-horizon stability, and the interactive track for interactive response and state retention and reuse. With expert-verified evaluation cases across 32 fine-grained dimensions, we conduct an extensive evaluation of 13 representative systems. Our analysis reveals that current models suffer from temporal drift in progressive generation and responsiveness bottlenecks during interactive control. Based on a comprehensive failure analysis, we share insights to advance the development of native joint audio-video streaming models.