发表机构
Tsinghua University(清华大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对扩散模型正则化忽视条件信号的问题,提出条件感知表示正则化框架CARE,动态调节特征分布,在ImageNet和文本到图像任务中显著降低FID并提升训练效率。
AI 中文摘要
近年来扩散模型的进展凸显了表示正则化在提高样本质量和训练效率方面的重要性。然而,常用的正则化方法往往忽视了直接决定生成目标的内置条件(如标签或文本)。在本工作中,我们展示了条件信号如何影响特征分布,并引入了CARE(条件感知表示正则化)。CARE是一个轻量级的即插即用正则化框架,它根据条件相似性动态调节特征分布。CARE利用内置条件信号来明智地引导表示空间,促进相似条件形成更紧密的特征簇,而无需依赖显式的对齐损失或外部监督。实验上,CARE在类别到图像和文本到图像任务中均持续提高了视觉保真度和收敛稳定性。在ImageNet上,CARE在40万训练步中实现了FID降低19.08%,带来了3.5倍的加速。当应用于文本到图像生成时,CARE在20万次迭代中将FID降低了16.61%,并改善了生成样本与文本提示之间的语义对齐。此外,CARE可以无缝集成到现有正则化方法中,带来额外的性能提升。
英文摘要
Recent advances in diffusion models highlight the importance of representation regularization for improving sample quality and training efficiency. However, commonly used regularization methods often overlook the built-in conditions (such as labels or texts) which directly determine the generation target. In this work, we demonstrate how conditioning signals affect the feature distribution and introduce the CARE (Condition-Aware REpresentation regularization). CARE is a lightweight plug-and-play regularization framework that dynamically modulates feature distribution based on condition similarity. CARE leverages built-in conditioning signals to judiciously guide the representation space, promoting tighter feature clusters for similar conditions without relying on explicit alignment losses or external supervision. Empirically, CARE consistently improves both visual fidelity and convergence stability across both class-to-image and text-to-image tasks. On ImageNet, CARE achieves a 19.08\% reduction in FID in 400k training steps, leading to a 3.5$\times$ speed-up. When applied to text-to-image generation, CARE lowers FID by 16.61\% in 200k iterations and improves semantic alignment between generated samples and text prompts. Moreover, CARE can be seamlessly integrated with existing regularization methods, yielding additional performance gains.