FreqFLD:通过频率调制实现一体化面部关键点检测
FreqFLD: Towards All-in-One Facial Landmark Detection via Frequency Modulation
- China Three Gorges University(三峡大学)
- Zhongnan University of Economics and Law(中南财经政法大学)
- National Institute of Natural Hazards, Ministry of Emergency Management of China(中国应急管理部国家自然灾害防治研究院)
- Yunnan University(云南大学)
- Wuhan University(武汉大学)
- Nanyang Technological University(南洋理工大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对现有面部关键点检测方法忽视频率信息导致跨数据集泛化受限的问题,提出频率调制框架FreqFLD,通过频率调制模块和混合专家机制实现一体化检测,在流行数据集上取得竞争力性能。
AI中文摘要:
深度学习的最新进展显著推进了面部关键点检测。然而,大多数现有方法在数据集特定的训练范式下以空间域方式处理特征,这忽略了面部关键点检测本质上是由几何驱动且对频率变化敏感的事实,从而限制了复杂场景下的跨数据集泛化能力,并阻碍了面部关键点检测模型的发展。为解决这一问题,我们提出了FreqFLD,一种面向一体化面部关键点检测的频率调制框架。具体而言,FreqFLD引入了一个频率调制模块(FreqMoM),通过解耦和调制低频与高频分量来显式引入频率先验,随后将其注入后续特征建模中,以实现全局面部结构和局部关键点细节的平衡建模。此外,FreqFLD采用了一种频率调制混合专家(FreqMoE),专家选择自适应地以频率调制先验为条件,从而能够在多样且具有挑战性的场景下灵活建模异构面部关键点模式。为了在一体化范式下规范频率一致性建模,我们进一步引入了一种频率一致性路由(FreqCR)损失,该损失约束频率感知专家的路由和分配,以促进跨不同面部场景的专家均衡利用,从而实现稳定的专家特化并获得鲁棒的面部关键点检测。大量实验表明,所提出的FreqFLD在流行数据集上取得了具有竞争力的性能。代码可在以下网址获取:此https URL。
英文摘要:
Recent progress in deep learning has significantly advanced facial landmark detection. However, most existing methods process features in a spatial-domain manner under a dataset-specific training paradigm, which overlooks the fact that facial landmark detection is inherently geometry-driven and sensitive to frequency variations, thereby limiting cross-dataset generalization under complex scenarios and hindering the development of a facial landmark detection model. To address this issue, we propose \textbf{FreqFLD}, a \textbf{freq}uency-modulated framework towards All-in-One \textbf{f}acial \textbf{l}andmark \textbf{d}etection. Specifically, FreqFLD introduces a Frequency Modulation Module (FreqMoM) to explicitly induce the frequency prior by decoupling and modulating low- and high-frequency components, which is then injected into subsequent feature modeling to enable balanced modeling of global facial structure and local landmark details. Furthermore, FreqFLD employs a Frequency-Modulated Mixture-of-Experts (FreqMoE), with expert selection adaptively conditioned on frequency-modulated priors, enabling flexible modeling of heterogeneous facial landmark patterns under diverse and challenging scenarios. To regularize frequency-consistent modeling under the All-in-One paradigm, we further introduce a Frequency-Consistent Routing (FreqCR) loss, which constrains the routing and assignment of frequency-aware experts to promote balanced expert utilization across diverse facial scenarios, thereby enabling stable expert specialization and achieving robust facial landmark detection. Extensive experiments demonstrate that the proposed FreqFLD achieves comparable performance on popular datasets. The code is available at: https://github.com/jkj1059657014/FreqFLD.