AI 中文总结
TimeSteer是用于联合视听扩散模型的无训练推理时语音调度框架,通过源区间定位与区域感知潜在重映射实现指定区间语音控制,在提升区间可控性的同时保持生成质量,还推出了相关基准SpeechShift。
AI 中文摘要
尽管预训练的联合视听扩散模型能对生成内容的“是什么”提供丰富控制,但无法对“何时”生成语音进行显式控制。为解决该问题,本文研究了推理时语音调度这一新任务,该任务可在不微调骨干模型的前提下,将耦合的语音与视觉发音限定在用户指定的起止区间内。本文揭示了去噪过程中可支撑该任务的两个固有特性:其一,对时序敏感的文本到音频交叉注意力头会沿潜在时间线暴露每个语音的模型隐含源区间;其二,预测得到的干净潜在变量已对耦合的语音与视觉发音完成组织,无需重新生成内容即可编辑其时序位置。基于上述发现,本文提出了无训练框架TimeSteer,该框架通过源区间定位(Source Span Localization)定位每个语音的源区间,并通过区域感知潜在重映射(Region-Aware Latent Remapping)将关联的视听潜在内容从源区间迁移至指定目标区间。本文还推出了首个用于联合视听生成中间隔级语音调度的基准SpeechShift。在两个代表性骨干模型上开展的实验表明,TimeSteer相较无训练基线大幅提升了区间可控性,同时保持了具有竞争力的整体生成质量。
英文摘要
Although pretrained joint audio-visual diffusion models offer rich control over \emph{what} to generate, they provide no explicit control over \emph{when} an utterance should occur. To address this, we study \emph{inference-time speech scheduling}, a novel task that places coupled speech and visual articulation within user-specified begin--end intervals without finetuning the backbone model. We uncover two intrinsic properties of the denoising process that enable this task. First, a timing-sensitive text-to-audio cross-attention head exposes each utterance's model-implied source span along the latent timeline. Second, the predicted clean latent already organizes coupled speech and visual articulation, allowing their temporal placement to be edited without regenerating the content. Building on these discoveries, we propose \textbf{TimeSteer}, a training-free framework that localizes each utterance's source span through \textbf{Source Span Localization} and transfers the associated audio-visual latent content from the source interval to the specified target interval through \textbf{Region-Aware Latent Remapping}. We further introduce \textbf{SpeechShift}, the first benchmark for interval-level speech scheduling in joint audio-visual generation. Experiments across two representative backbones show that TimeSteer substantially improves interval controllability over training-free baselines while maintaining competitive overall generation quality.