arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.09866cs.LG

MUNITE:面向任意到任意多模态生成的统一多模态潜在推理

MUNITE: Unified Multimodal Latent Inference for Any-to-Any Multimodal Generation

  • KAIST(韩国科学技术院)

机构由 AI 辅助整理,请以论文原文为准。

Kyeongmin Yeo, Minhyuk Sung

AI总结:

MUNITE提出统一潜在推理框架,将编码与生成视为同一推理问题,通过自蒸馏条件流匹配实现任意到任意多模态生成,在多个基准上取得更优质量与一致性。

AI中文摘要:

我们提出了MUNITE,一个用于灵活任意到任意多模态生成的潜在变量框架,该框架将编码和潜在生成视为在不同数量的观测证据下的同一推理问题。给定任意模态子集,MUNITE对与完整观测相关联的潜在表示的条件分布进行建模。完整观测恢复确定性编码,无观测恢复潜在边缘分布,中间子集定义条件潜在推理,所有这些都在一个单一的条件流模型中完成。共享的潜在样本捕获了在生成目标之间必须保持一致的变化,而模态特定的生成解码器则独立地对剩余的不确定性进行建模。为了从不完整的训练示例中学习这些条件分布,我们通过自蒸馏扩展了条件流匹配:基于更丰富可用观测的预测,在同一中间潜在状态下,监督同一模型在较小子集上的条件预测。当更丰富证据的轨迹遵循精确的条件流时,这提供了与全目标去噪相同的期望学习信号。在PolyMNIST-D-Q、FFHQ64和图像-文本-音频上,MUNITE实现了具有竞争力或更好的生成质量和源-目标对齐,并具有更高的联合生成一致性。特别是,它在所有一对多和无条件图像-文本-音频比较中达到了最高的一致性,展示了统一潜在推理在不同多模态设置中的有效性。

英文摘要:

We introduce MUNITE, a latent-variable framework for flexible any-to-any multimodal generation that treats encoding and latent generation as the same inference problem under different amounts of observed evidence. Given any subset of modalities, MUNITE models the conditional distribution over the latent representation associated with the complete observation. Full observation recovers deterministic encoding, no observation recovers the latent marginal, and intermediate subsets define conditional latent inference, all within a single conditional flow model. A shared latent sample captures variation that must remain consistent across generated targets, while modality-specific generative decoders model the remaining uncertainty independently. To learn these conditional distributions from incomplete training examples, we extend conditional flow matching through self-distillation: predictions conditioned on richer available observations supervise the same model conditioned on smaller subsets at the same intermediate latent state. When the richer-evidence trajectory follows the exact conditional flow, this provides the same expected learning signal as full-target denoising. Across PolyMNIST-D-Q, FFHQ64, and image-text-audio, MUNITE achieves competitive or better generation quality and source-target alignment, with higher joint-generation coherence. In particular, it attains the highest coherence in all one-to-many and unconditional image-text-audio comparisons, showing the effectiveness of unified latent inference across diverse multimodal settings.

补充信息

↑