发表机构
University of North Carolina at Charlotte(北卡罗来纳大学夏洛特分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对VLM可从3D人脸渲染图像提取敏感属性的隐私挑战,提出3D FaceShell框架,通过可学习高斯壳产生扰动,在保持面部外观的同时引导VLM语义解释,实验证明其能有效提高属性注入和不匹配率。
AI 中文摘要
逼真的3D人脸头像在远程呈现、动画和个性化媒体等应用中越来越多地作为可重复使用的数字资产部署。同时,视觉语言模型(VLM)可以通过开放式语义推理从渲染图像中推断敏感属性而无需任何微调。这带来了新的隐私挑战。现有防御大多在2D图像空间中运行,未解决3D面部表示的身份保留语义操作问题。我们提出3D FaceShell,一个在保留几何保真度和面部身份的同时引导VLM对面部渲染进行解释的框架。它用可学习的高斯壳增强原始3D表示,通过多视图嵌入对齐优化产生微妙的空间分布扰动。实验表明,3D FaceShell显著提高属性注入和不匹配率,同时保持高感知相似度和身份一致性。
英文摘要
Photorealistic 3D face avatars are increasingly deployed as reusable digital assets across applications such as telepresence, animation, and personalized media. At the same time, vision-language models (VLMs) can infer sensitive attributes from rendered images with open-ended semantic reasoning without any fine-tuning. This creates a new privacy challenge: once a 3D face avatar is shared, any of its renderings can be analyzed to extract high-level facial attributes. Existing defenses largely operate in 2D image space and do not address identity-preserving semantic manipulation of 3D facial representations. We propose 3D FaceShell, a framework for steering VLM interpretations of faces rendered from 3D models while preserving geometric fidelity and facial identity. 3D FaceShell augments the original 3D representation with a learnable Gaussian shell that produces subtle, spatially distributed perturbations optimized through multi-view embedding alignment. The perturbations are designed to be visually inconspicuous yet sufficient to redirect VLM-based attribute inference in a view-consistent manner. Extensive experiments on reconstructed celebrity face avatars and multiple black-box VLMs demonstrate that 3D FaceShell significantly increases attribute injection and mismatch rates while maintaining high perceptual similarity and identity consistency. Our results show that it is possible to manipulate VLM-level semantic interpretation of 3D faces without compromising their human-recognizable appearance.
CommentsECCV 2026