会发声的图像:在单一画布上合成图像与声音
Images that Sound: Composing Images and Sounds on a Single Canvas
浏览论文内容
中文总结 AI 辅助
本文提出一种零样本方法,利用共享潜在空间中的文本到图像与文本到频谱图扩散模型并行去噪,合成同时具有自然图像外观和自然音频听感的频谱图。
中文摘要 AI 辅助
频谱图是声音的二维表示,其外观与我们视觉世界中的图像截然不同。而自然图像若被当作频谱图播放,会产生不自然的声音。本文表明,可以合成同时看起来像自然图像、听起来像自然音频的频谱图。我们将这些视觉频谱图称为“会发声的图像”。我们的方法简单且零样本,利用在共享潜在空间中运行的预训练文本到图像和文本到频谱图扩散模型。在反向过程中,我们同时使用音频和图像扩散模型对带噪潜在变量进行去噪,从而得到在两个模型下都具有高可能性的样本。通过定量评估和感知研究,我们发现该方法成功生成了既符合期望音频提示、又具有期望图像提示视觉外观的频谱图。视频结果请参见项目页面:https://ificl.github.io/images-that-sound/
英文摘要
Spectrograms are 2D representations of sound that look very different from the images found in our visual world. And natural images, when played as spectrograms, make unnatural sounds. In this paper, we show that it is possible to synthesize spectrograms that simultaneously look like natural images and sound like natural audio. We call these visual spectrograms images that sound. Our approach is simple and zero-shot, and it leverages pre-trained text-to-image and text-to-spectrogram diffusion models that operate in a shared latent space. During the reverse process, we denoise noisy latents with both the audio and image diffusion models in parallel, resulting in a sample that is likely under both models. Through quantitative evaluations and perceptual studies, we find that our method successfully generates spectrograms that align with a desired audio prompt while also taking the visual appearance of a desired image prompt. Please see our project page for video results: https://ificl.github.io/images-that-sound/
发表机构
- University of Michigan(密歇根大学)
机构由 AI 辅助整理,请以论文原文为准。