发表机构
Beihang University; Pengcheng Laboratory; Shanxi University of Finance and Economics; North China University of Technology; Cardiff University; Beijing Institute of Technology(北京航空航天大学; 鹏城实验室; 山西财经大学; 北方工业大学; 卡迪夫大学; 北京理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究构建了含文本-网格等标注的FaME-G2E数据集,提出RAGMesh框架,含MSRF模块与AdaRAGS约束,在3D人脸生成编辑任务上性能优于现有方法。
AI 中文摘要
长文本驱动的3D人脸生成与编辑仍具挑战性,因为将长文本描述转换为精细的面部几何结构难度较大。现有方法主要将全局文本语义与面部结构对齐,但往往难以捕捉微妙的局部变形,如眉毛张力、脸颊收缩和不对称的嘴部动作,导致几何保真度和编辑精度有限。为促进精细的文本驱动面部建模,我们首先构建了FaME-G2E,这是一个大规模多模态数据集,包含用于统一3D人脸生成与编辑的详细文本-网格标注及配对的文本-混合形状样本。基于该数据集,我们提出了RAGMesh,这是一种检索增强框架,利用与文本相关的几何先验来提升高保真人脸合成与编辑。具体而言,多尺度检索融合(MSRF)模块检索语义一致的全局和区域面部先验,并在混合形状空间中对其进行融合,抑制冲突的局部变形同时保留连贯的变形模式。此外,我们引入了自适应RAG引导监督(AdaRAGS),这是一种区域感知约束,可明确将文本语义与对应面部区域对齐,增强区域可控性和编辑精度。在FaME-G2E上开展的大量实验表明,RAGMesh在局部几何精度、文本引导可控性、区域编辑精度和推理效率方面均优于现有最优方法。视频演示可在此URL获取,源代码和数据集将在论文录用后发布。
英文摘要
Text-driven 3D face generation and editing remains challenging due to the difficulty of translating long-form descriptions into fine-grained facial geometry. Existing methods primarily align global textual semantics with facial structures but often struggle to capture subtle local deformations, such as eyebrow tension, cheek contraction, and asymmetric mouth motions, resulting in limited geometric fidelity and editing precision. To facilitate fine-grained text-driven facial modeling, we first construct FaME-G2E, a large-scale multimodal dataset containing detailed text--mesh annotations and paired text--blendshape samples for unified 3D facial generation and editing. Based on this dataset, we propose RAGMesh, a retrieval-augmented framework that leverages text-correlated geometric priors to improve high-fidelity facial synthesis and editing. Specifically, the Multi-Scale Retrieval Fusion (MSRF) module retrieves semantically consistent global and regional facial priors and fuses them in the blendshape space, suppressing conflicting local deformations while preserving coherent deformation patterns. Furthermore, we introduce Adaptive RAG-guided Supervision (AdaRAGS), a region-aware constraint that explicitly aligns textual semantics with corresponding facial regions, enhancing regional controllability and editing accuracy. Extensive experiments on FaME-G2E demonstrate that RAGMesh achieves superior performance over state-of-the-art methods in local geometric accuracy, text-guided controllability, regional editing precision, and inference efficiency. Video demo is available at https://youtu.be/Yr0_XkpWcNk, and the source code and dataset will be released upon paper acceptance.