发表机构
The Chinese University of Hong Kong(香港中文大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出IGA框架,通过熵约束投影实现从模仿到想象的生成,可修复数据多样性并实现可控频谱外推,相关方法适用于DDPM等扩散模型。
AI 中文摘要
生成人工智能模型主要设计用于拟合数据分布,该目标既无法修正学习到的生成器丢失的多样性,也无法定义生成应如何扩展至超出数据本身的多样性。我们提出想象式生成人工智能(Imaginative Generative AI,IGA),这一框架将多样性纳入目标分布设计问题:在与参考分布接近的分布中,IGA选择其频谱多样性达到规定水平的分布。多样性通过固定表示空间中生成分布的核协方差算子的冯·诺依曼熵来衡量,提供一种无参考、由表示引导的度量,用于衡量概率质量在嵌入方向上的分布广度。总体数据分布的频谱熵定义了一道熵墙。在熵墙以下,IGA执行多样性修复,恢复学习到的生成器丢失的变异,同时保持在数据的多样性水平内;在熵墙以上,数据分布本身变得不可行,IGA会刻意偏离数据分布,以产生具有更大表示相关频谱多样性的分布,这是想象式生成的操作概念。这些机制构成了从模仿到想象的单一正则化路径,并在每个规定的多样性水平上定义了一个独立同分布的目标分布。我们发展了这一熵约束投影的理论,证明在与预训练生成器的KL锚定下,最优解满足自洽的指数倾斜关系。这一特性催生了IGA引导,一种无需重新训练的基于分数和扩散模型的推理时方法,包括DDPM和DDIM采样器。在合成数据和视觉基准上的实验表明,IGA在熵墙以下实现了多样性修复,并在熵墙以上实现了可控的频谱外推。
英文摘要
Generative AI models are primarily designed to imitate the data distribution, an objective that neither corrects diversity lost by a learned generator nor defines how generation should extend beyond the diversity of the data itself. We introduce Imaginative Generative AI (IGA), a framework that makes diversity part of the target-distribution design problem: among distributions close to a reference, IGA selects one whose spectral diversity reaches a prescribed level. Diversity is measured by the von Neumann entropy of the generated distribution's kernel covariance operator in a fixed representation space, providing a reference-free representation-guided measure of how broadly probability mass occupies embedding directions. The spectral entropy of the population data distribution defines an Entropy Wall. Below the wall, IGA performs diversity repair, recovering variation that a learned generator has lost while remaining within the diversity level of the data. Beyond the wall, the data distribution itself becomes infeasible, and IGA deliberately departs from it to produce distributions with greater representation-relative spectral diversity, an operational notion of imaginative generation. These regimes form a single regularization path from imitation to imagination and define an i.i.d. target distribution at each prescribed diversity level. We develop the theory of this entropy-constrained projection and show that, under a KL anchor to a pretrained generator, the optimum satisfies a self-consistent exponential-tilt relation. This characterization leads to IGA Guidance, a retraining-free inference-time method for score-based and diffusion models, including DDPM and DDIM samplers. Experiments on synthetic and vision benchmarks demonstrate diversity repair below the Entropy Wall and controlled spectral extrapolation beyond it.