arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

关系合成:面向拟音与检索增强音频生成的结构介导拼接合成

Relational Synthesis: Structure-Mediated Concatenative Synthesis for Foley and Retrieval-Augmented Audio Generation

Keren Shao, Ayaka Kawano, Shlomo Dubnov

arXiv 2610.05768首次发表:更新:

发表机构

University of California San Diego(加利福尼亚大学圣迭戈分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出关系合成方法,通过Gromov式结构代价将参考音频的时间结构与幅度运动转移到源音频颗粒拼接中,实现不复制片段且保持声学一致的拟音生成,并可与神经RAG自然集成。

AI 中文摘要

我们提出一个问题:给定检索到的源音频$S$和独立的参考音频$R$,能否从这对音频$(S,R)$中合成新颖的音频$Y$,使得$Y$在声学上与$S$保持一致,同时不持续复制$S$或$R$的片段?第一个子目标在拟音音频制作中是众所周知的,而第二个子目标则是神经检索增强生成(RAG)中当$S$和$R$被朴素地注入神经生成器时的一个已知问题。我们证明,这两个子目标可以通过我们称之为关系合成的方法同时解决,这是一种拼接合成的变体,其中目标代价被替换为一种关系型的Gromov式结构代价。关系合成并非模仿$R$的内容,而是从“山的另一侧”利用它:它将$R$的时间结构和定向幅度运动转移到$S$的颗粒重组与拼接中,以新颖的方式保护$S$的声学信息。我们的实验表明,关系合成能自然地与神经RAG集成,并产生在时间一致性、声学保真度和泄漏持续性指标上表现良好的拟音音频,同时保持分布级质量和文本对齐。

英文摘要

We ask: given a retrieved source audio $S$ and a separate reference audio $R$, can we synthesize novel audio $Y$ out of this pair $(S,R)$ such that $Y$ remains acoustically consistent with $S$, while not persistently copying segments of $S$ or $R$? The first clause is a well-known goal in Foley audio production, and the second is a well-known issue in neural RAG when $S$ and $R$ are naively injected into neural generators. We show that both clauses can be addressed simultaneously using a method we coin relational synthesis, a variation of concatenative synthesis where target cost is replaced by a relational Gromov-like structural cost. Rather than imitating the content of $R$, relational synthesis exploits it from the "other side of the hill": it transfers the temporal structure and directed amplitude motion of $R$ to reorganize and concatenate the grains of $S$ in a novel manner that protects $S$'s acoustic information. Our experiments show that relational synthesis integrates naturally with neural RAG and produces Foley audio that performs well on metrics measuring temporal agreement, acoustic fidelity, and leakage persistence, while maintaining distribution-level quality and text alignment.

Comments5 pages, 2 figures, 2 tables. Submitted to ICASSP 2027. Audio demo: https://smoothken.github.io/relational_synth_demo/. This work has been submitted to the IEEE for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessible

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑