arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.15240cs.SD

生成未曾听闻的声音:系统发育引导的潜在生成用于祖先声音重建

Generating the Unheard: Phylogeny-Guided Latent Generation for Ancestral Sound Reconstruction

  • University of Wisconsin–Madison(威斯康星大学麦迪逊分校)
  • Massachusetts Institute of Technology(麻省理工学院)

机构由 AI 辅助整理,请以论文原文为准。

Tianyi Xu, Shrinaath Narasimhan, Evan Gorstein, Santiago Perea, Yunyi Shen, Claudia Solís-Lemus

AI总结:

提出首个系统发育引导的潜在生成框架,通过VAE编码、性状投影和锚定逆提升,为鸟类祖先节点生成全新且合理的发声,并在两个支系上验证了生成质量与系统发育一致性。

AI中文摘要:

一种祖先鸟类物种听起来是什么样的?现有的祖先状态重建方法可以在系统发育树的内部节点推断低维性状,例如形态特征,但尚未有人尝试生成丰富的感知信号,如音频。一些挑战包括推断的表示要么维度过低而无法解码,要么位于非生成性特征空间中,因此迄今为止没有方法能够生成祖先音频。我们引入了第一个能够生成合理祖先发声的框架。我们的流程将鸟类录音编码到VAE潜在空间中,学习一个与系统发育距离对齐的低维性状投影,在该性状空间中执行祖先推断,并通过锚定逆提升恢复可解码的潜在表示,然后为每个祖先节点生成新的波形。由于整个流程保持在可解码的潜在空间中,每个内部节点都获得一个真正新的音频输出,代表检索式替代方法无法获得的合理中间祖先声音。在两个系统发育上相距较远的鸟类支系——21种霸鹟科和19种山雀科——上的实验表明,我们的方法是唯一在两个数据集上同时实现真正生成、系统发育一致性和自然音频质量的方法。

英文摘要:

What did an ancestral bird species sound like? Existing ancestral state reconstruction methods can infer low-dimensional traits such as morphological characters at internal nodes of a phylogenetic tree, but no one has tried to produce rich perceptual signals such as audio. Some of the challenges include inferred representations that are either too low-dimensional to decode or lie in non-generative feature spaces, so no method to date can produce ancestral audio. We introduce the first framework that generates plausible ancestral vocalizations. Our pipeline encodes bird recordings into a VAE latent space, learns a low-dimensional trait projection aligned with phylogenetic distances, performs ancestral inference in this trait space, and recovers decodable latents through an anchored inverse lift before emitting novel waveforms for each ancestral node. Because the entire pipeline stays within a decodable latent space, every internal node receives a genuinely new audio output representing plausible intermediate ancestral sounds unavailable to retrieval-based alternatives. Experiments on two phylogenetically distant bird clades, 21-species Tyrannidae and 19-species Paridae, show that our method is the only approach that simultaneously achieves genuine generation, phylogenetic consistency, and naturalistic audio quality across both datasets.

补充信息

↑