arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.09529cs.CV

面向表达性与忠实性的音频到图像生成:一个统一的多模态数据集与合成框架

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework

Dongxu Ge, Shansong Liu, Cheng Gong, Xiao-Lei Zhang, Chi Zhang, Xuelong Li

首次发表
浏览论文内容

中文总结 AI 辅助

针对音频到图像生成受限于传统数据集的问题,提出A2I-Set数据集与AudioCanvas模型,实现了更优的跨模态对齐与视觉表达性。

中文摘要 AI 辅助

作为跨模态生成的重要子领域,从音频合成图像形式的静态视觉内容,即音频到图像(A2I)生成,近年来受到越来越多的研究关注。然而,尽管现代文本到图像(T2I)模型的视觉质量已十分出色,A2I的性能仍受限于传统数据集,这类数据集往往同时缺乏高保真图像与精确的跨模态对齐。因此,现有方法通过微调强大的T2I模型仍难以实现高质量的音频到图像生成,从而限制了该领域的实际应用。针对这一缺口,我们推出A2I-Set,一个统一的高质量三模态数据集,包含32.3万对音频、图像及详细文本描述,专为视听研究(包括音频条件下的图像生成)设计。此外,我们通过人工监督为A2I任务构建了一个新的混合源测试集。我们进一步提出A2I模型AudioCanvas,在我们的A2I-Set上进行微调。实验表明,AudioCanvas实现了更具视觉表达性且跨模态对齐效果更好的结果,总体优于现有方法。我们的数据集和源代码可在该https网址获取。

英文摘要

As an important subfield of cross-modal generation, synthesizing static visual content in the form of images from audio, namely audio-to-image (A2I) generation, has attracted increasing research attention in recent years. Nevertheless, despite the remarkable visual quality of modern text-to-image (T2I) models, the performance of A2I remains fundamentally limited by traditional datasets, which often lack both high-fidelity images and precise cross-modal alignment. As a result, existing methods still struggle to achieve high-quality audio-to-image generation through finetuning strong T2I models, thereby constraining practical applications in this area. Motivated by this gap, we introduce A2I-Set, a unified, high-quality tri-modal dataset consisting of 323K paired audio, images, and detailed text captions, specifically designed for audio-visual research, including audio-conditioned image generation. Besides, we developed a new mixed-source test set for the A2I task through human supervision. We further propose an A2I model, AudioCanvas, fine-tuned on our A2I-Set. Experiments show that AudioCanvas achieves more visually expressive as well as cross-modal alignment results that generally outperforming existing approaches. Our dataset and source code are available at https://github.com/gdx012/A2I-Generation.

发表机构

  • University of Science and Technology of China(中国科学技术大学)
  • China Telecom (TeleAI)(中国电信(电信人工智能研究院))
  • Northwest Polytechnical University(西北工业大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑