arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

冻结视觉基础模型中新兴的区域级面部对应

Emergent Region-Level Facial Correspondence in Frozen Vision Foundation Models

Izaldein Al-Zyoud, Abdulmotaleb El Saddik

arXiv 2607.14423首次发表:更新:

AI 中文总结

研究人脸区域级对应问题,利用冻结的DINOv3特征及FaRL,在真实视频上评估跨身份匹配和时间标签传播,结果表明DINOv3是区域级面部对应强大零样本表示,中间自监督特征对密集面部分析最有用。

AI 中文摘要

冻结的自监督视觉模型可以对齐通用物体的部分,但这种对应关系是否扩展到人脸尚不清楚,人脸的全局布局相同,但特定身份的外观差异很大。我们测试冻结的DINOv3特征是否定义了一个区域级面部坐标系:一个特征空间,其中眼睛、眉毛、鼻子、嘴巴、皮肤和头发在不同人和不同时间都能保持可区分,无需特定于面部的训练。使用DINOv3 ViT-L/16补丁嵌入和FaRL仅作为面部部分标记接口,我们在200个CelebDF-v2真实视频上评估跨身份最近邻匹配和时间标签传播。DINOv3在无约束跨身份匹配下实现了83.0%的区域级语义准确率,而面积加权随机基线为23.0%,在没有学习时间模块的情况下实现了95.5%的时间跟踪准确率。一个无FaRL的对照组降至0.9%,表明FaRL提供语义初始化,而DINOv3提供密集的空间对应。最强的对应出现在中间层:第18块给出了4.93倍的同区域与跨区域辨别率,而最后一块为1.48倍。与CLIP ViT-L/14相比,DINOv3仅显示出小的总体优势,但在解剖区域上有+16.8个百分点的增益,表明图像级对比监督捕获了粗略的面部布局,但没有捕获细粒度的解剖身份。这些结果将冻结的DINOv3确立为区域级面部对应的强大零样本表示,并确定中间的自监督特征是密集面部分析最有用的层。

英文摘要

Frozen self-supervised vision models can align parts of generic objects, but it remains unclear whether this correspondence extends to human faces, where global layout is shared while identity-specific appearance varies sharply. We test whether frozen DINOv3 features define a region-level facial coordinate system: a feature space in which eyes, brows, nose, mouth, skin, and hair remain distinguishable across people and across time without face-specific training. Using DINOv3 ViT-L/16 patch embeddings and FaRL only as a face-part labeling interface, we evaluate cross-identity nearest-neighbor matching and temporal label propagation on 200 CelebDF-v2 real videos. DINOv3 achieves 83.0% region-level semantic accuracy under unconstrained cross-identity matching, compared with a 23.0% area-weighted random baseline, and 95.5% temporal tracking accuracy without a learned temporal module. A no-FaRL control collapses to 0.9%, showing that FaRL supplies semantic initialization while DINOv3 supplies dense spatial correspondence. The strongest correspondence appears at an intermediate layer: block 18 gives a 4.93x same-region versus cross-region discrimination ratio, compared with 1.48x at the final block. Against CLIP ViT-L/14, DINOv3 shows only a small aggregate advantage but a +16.8 pp gain on anatomical regions, indicating that image-level contrastive supervision captures coarse facial layout but not fine-grained anatomical identity. These results establish frozen DINOv3 as a strong zero-shot representation for region-level facial correspondence and identify intermediate self-supervised features as the most useful layer for dense face analysis.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑