发表机构
Bhabha Atomic Research Centre (B.A.R.C.); Indian Institute of Technology Bombay(巴巴原子研究中心; 印度理工学院孟买分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出语义边界预测器(SBP),通过时间步解耦实现公平人脸生成,无需重训练或外部数据集,在CelebA-HQ上大幅降低人口统计公平差距且保持图像质量,计算开销小易集成。
AI 中文摘要
合成人脸生成中的人口统计不平衡会传播到下游人脸识别系统,因此当扩散模型用于数据生成时,公平性成为重要考量。现有感知公平性的生成方法通常需要模型重训练、架构修改,或在整个反向扩散过程中反复施加引导。本研究提出语义边界预测器(Semantic Boundary Predictor, SBP),这是一种推理时框架,通过反向去噪过程中的单次干预实现人口统计引导。该方法的动机源于观察:不同扩散时间步的隐层表征具有不同语义作用:后期隐层提供更强的人口统计可分性,而早期隐层为语义干预提供更大灵活性。SBP利用这种时间步解耦特性,从后期隐层表征中学习线性语义边界,仅在初始噪声隐层处应用一次,使反向去噪过程的其余部分保持不变。该方法无需对底层隐层扩散模型进行重训练或微调,也无需外部平衡数据集。在CelebA-HQ上的实验表明,该方法在人口统计公平性方面取得显著改进:性别公平差距降低98%,二元种族公平差距降低95%,四类种族公平差距降低15%,同时在各人口统计群体中保持感知图像质量。由于其单次推理策略和模型无关设计,SBP仅引入少量计算开销,可便捷集成到现有预训练隐层扩散模型中。
英文摘要
Demographic imbalance in synthetic face generation can propagate to downstream face recognition systems, making fairness an important consideration when diffusion models are used for data generation. Existing fairness-aware generation approaches often require model retraining, architectural modifications, or repeated guidance throughout the reverse diffusion process. In this work, we introduce Semantic Boundary Predictor (SBP), an inference-time framework that performs demographic guidance through a one-shot intervention during reverse denoising. Our approach is motivated by the observation that latent representations at different diffusion timesteps play distinct semantic roles: late-stage latents provide stronger demographic separability, whereas early-stage latents offer greater flexibility for semantic intervention. SBP leverages this timestep decoupling by learning linear semantic boundaries from late-stage latent representations while applying them only once at the initial noisy latent, allowing the remainder of the reverse denoising process to proceed unchanged. The method requires neither retraining nor fine-tuning of the underlying Latent Diffusion Model and operates without external balanced datasets. Experiments on CelebA-HQ demonstrate substantial improvements in demographic fairness, reducing fairness disparity by 98% for gender, 95% for binary race, and 15% for four-class race, while maintaining perceptual image quality across demographic groups. Owing to its one-shot inference strategy and model-agnostic design, SBP introduces only a small computational overhead and can be readily integrated with existing pre-trained latent diffusion models.