arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.02367cs.MMcs.CV

缺失的时间关联:面向脚本驱动的音视频生成的时间上下文路由

The Missing Temporal Link: Temporal Context Routing for Script-Driven Audio-Video Generation

  • PKU(北京大学)
  • Qwen Applications(通义千问应用公司)
  • HKUST(香港科技大学)
  • CUHK(香港中文大学)
  • SJTU(上海交通大学)

机构由 AI 辅助整理,请以论文原文为准。

Yichen Liu, Quanwei Zhang, Haozhe Wang, Donghao Zhou, Jiankun Zhang, Xiaojie Li, Yang Shi, Jiaming Liu, Ruihua Huang, Yingtian Zou, Daquan Zhou

AI总结:

针对脚本驱动音视频生成中时序对齐不足的问题,提出时间上下文路由(TCR)方法,将脚本时序映射到音视频共享时间轴,显著提升了镜头边界和对话时序的准确性,且保持了音视频质量与同步性。

AI中文摘要:

联合音视频生成模型在视觉质量和音视频同步方面已取得显著进展,但它们对镜头切换时机和对话发生时机的控制能力仍有限,这一限制制约了其在脚本驱动内容创作中的应用,时序错误会破坏叙事连贯性和观看体验。当前的联合生成器在共享时间轴上对齐视频和音频表示,但结构化提示中指定的镜头和对话的精确时序仅编码在提示的文本表示中,未与任一模态的时间坐标对齐,导致视频和音频虽彼此同步,却均未遵循脚本时间线。这种不匹配促使我们将时间对齐从视频和音频扩展至结构化脚本,因此引入时间上下文路由(Temporal Context Routing, TCR),该方法将脚本时序映射到音视频生成的共享时间轴,并将每个提示的引导信息路由到两种模态的对应位置。在200个测试脚本上与基线相比,TCR将镜头边界平均绝对误差(Shot Boundary MAE)降低96%,从1.11秒降至0.042秒,将对话准确率(Dialogue Acc@0.5 s)从28.3%提升至84.1%;同时TCR在保持与基线相当的视觉质量和音视频同步性的前提下实现了这些改进,用户研究进一步显示,参与者在所有五个评估维度上均更偏好TCR。

英文摘要:

Joint audio-video generation models have made substantial progress in visual quality and audio-visual synchronization. However, they still provide limited control over when shot transitions occur and dialogue is spoken. This limitation constrains their application in script-driven content creation, where timing errors can undermine narrative coherence and the viewing experience. Current joint generators align video and audio representations on a shared temporal axis, yet the precise timing of shots and dialogue specified in a structured prompt is encoded only in the prompt's text representation and remains unaligned with the temporal coordinates of either modality. Consequently, video and audio may remain synchronized with each other while both fail to follow the script timeline. This mismatch motivates us to extend temporal alignment beyond video and audio to include the structured script. We therefore introduce Temporal Context Routing (TCR), which maps the script timing onto the shared temporal axis of video and audio generation and routes each prompt's guidance to the corresponding positions in both modalities. Compared with the baseline on 200 test scripts, TCR reduces Shot Boundary MAE by 96%, from 1.11 s to 0.042 s, and raises Dialogue Acc@0.5 s from 28.3% to 84.1%. TCR achieves these improvements while maintaining visual quality and audio-visual synchronization comparable to those of the baselines. A user study further shows that participants prefer TCR on all five evaluated dimensions.

↑