arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

VoxAudio:基于多奖励自回归流匹配的发声音频合成

VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching

Wenxiang Guo, Changhao Pan, Ziyue Jiang, Zhou Zhao, Fei Wu

arXiv 2608.12951首次发表:更新:

发表机构

Zhejiang University(浙江大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

VoxAudio是一种多奖励自回归流匹配模型,通过架构、偏好、数据层面的优化,解决现有T2A系统发声控制不足的问题,在多基准上验证了有效性与效率。

AI 中文摘要

发声音频合成是指生成嵌入可理解语音的环境音景音频的任务,是播客制作、视频配音等应用的基础。现有文本到音频(T2A)系统要么将引用语音简化为不可理解的发声低语,要么将其交由单独的文本到语音(TTS)模型进行事后混合,这会失去对语音出现时机及其与场景交互方式的控制。我们提出VoxAudio,一种因果自回归流匹配模型,从三个互补方面解决该问题:架构层面,具有逐块独立噪声水平的分块因果分解,使音频可通过带KV缓存的滑动窗口流式推理生成,支持可变目标时长;为实现任意块粒度的推理,我们还通过随机块边界预训练模型。偏好层面,多奖励负感知微调(NFT)联合优化语义保真度、语言准确性、审美质量和时间定位。数据层面,为补充发声内容缺失的监督,我们构建了VoxCorpus——一个带时间间隔的嵌入语音逐字转录文本标注的大规模语料库,以及VoxBench——一个带时间定位指标的区间标注基准。在涵盖通用音频、语音和统一发声音频的四个基准上的实验验证了VoxAudio的有效性和效率,代码和演示可在https URL获取。

英文摘要

Vocalized audio synthesis, the task of generating audio in which intelligible speech is embedded within an environmental soundscape, underpins applications such as podcast production and video dubbing. Existing Text-to-Audio (T2A) systems either reduce quoted speech to unintelligible vocal murmur or delegate it to a separate TTS model with post-hoc mixing, which forfeits control over when speech occurs and how it interacts with the scene. We present VoxAudio, a causal autoregressive flow matching model that addresses this problem from three complementary aspects. At the architecture level, chunk-wise causal factorization with independent per-chunk noise levels lets audio be emitted through sliding-window streaming inference with KV caching at variable target durations; to enable inference at arbitrary chunk granularities, we further pretrain the model with randomized chunk boundaries. At the preference level, multi-reward Negative-aware FineTuning (NFT) jointly optimizes semantic fidelity, linguistic accuracy, aesthetic quality, and temporal grounding At the data level, to supply the missing supervision for vocal content, we build VoxCorpus, a large-scale corpus whose captions quote the verbatim transcript of embedded speech with time intervals, and VoxBench, an interval-annotated benchmark with a temporal-grounding metric. Experiments on four benchmarks spanning general audio, speech, and unified vocalized audio validate the effectiveness and efficiency of VoxAudio. Our code and demos are available at https://voxaudio.github.io.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑