arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MagicMakeup:用于高保真妆容转移的区域可控扩散Transformer

MagicMakeup: A Region-Controllable Diffusion Transformer for High-Fidelity Makeup-Transfer

Ziyi Wang, Siming Zheng, Yang Yang, Shusong Xu, Hao Zhang, Bo Li, Changqing Zou, Peng-Tao Jiang

arXiv 2607.20924首次发表:更新:

发表机构

Zhejiang University; State Key Lab of CAD&CG; vivo Mobile Communication Co., Ltd.; vivo BlueImage Lab(浙江大学; 计算机辅助设计与图形学国家重点实验室; 维沃移动通信有限公司; vivo蓝图像实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究妆容转移中区域可控性等难题,提出MagicMakeup框架,基于空间约束和概念解缠结,用令牌对齐区域门控和跨模态感知指导实现精确编辑与概念澄清,设计数据生成管道并建立基准,提升了多方面性能和鲁棒性。

AI 中文摘要

妆容转移旨在将参考妆容应用于源面部同时保留源身份。尽管基于扩散的方法在全脸编辑方面取得了进展,但强大的区域可控性、妆容保真度和身份保留仍具有挑战性。原因包括像素到注意力的错位、转移/保留概念分离不清晰以及缺乏高分辨率数据集。本文提出MagicMakeup,一个基于扩散Transformer的框架,基于空间约束和概念解缠结实现区域可控和高保真妆容转移。提出令牌对齐区域门控以实现精确的区域特定编辑并保留身份,引入跨模态感知指导以澄清转移和保留的概念。设计了通过特定区域妆容去除生成1024x1024数据对的管道并建立统一基准。实验表明MagicMakeup提高了区域可控性、妆容保真度和身份保留,在多种风格、种族和姿势下具有强大的鲁棒性。

英文摘要

Makeup-transfer applies the reference makeup to the source face while preserving the source identity. Despite advances in full-face editing by diffusion-based methods, strong regional controllability, makeup fidelity, and identity preservation remain challenging. The reasons are (i) pixel-to-attention misalignment that causes spillover into non-target areas and weakens regional control; (ii) unclear transfer/preservation concept separation under two-image conditioning, leading to coupling between makeup attributes and identity; and (iii) the lack of a high-resolution dataset that is identity-consistent and region-labeled for fine-grained supervision. In this paper, we propose MagicMakeup, a diffusion transformer-based framework for region-controllable and high-fidelity makeup transfer, built on spatial constraints and concept disentanglement. To enable precise region-specific editing while preserving identity, we propose Token-Aligned Region Gating, which aligns pixel masks with attention and applies region-specific logit gating. To clarify the concepts of transfer and preservation, we further introduce Cross-Modal Perception Guidance, which aligns text and image features to enhance cross-modal concept perception. We also design a pipeline for the generation of 1024 x 1024 data pairs through region-specific makeup removal and establish a unified benchmark in synthetic and real settings. Extensive quantitative and qualitative experiments show that MagicMakeup improves regional controllability, makeup fidelity, and identity preservation, with strong robustness across styles, races, and poses.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑