发表机构
Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, School of Informatics, Xiamen University; Fujian Key Laboratory of Pattern Recognition and Image Understanding, Xiamen University of Technology; School of Computer and Information Engineering, Xiamen University of Technology(中国教育部多媒体可信感知与高效计算重点实验室,厦门大学信息学院; 厦门理工学院福建省模式识别与图像理解重点实验室; 厦门理工学院计算机与信息工程学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对跨域少样本面部表情识别问题,提出AUCH-Net,通过新的AFL和VFL模块,在关系一致性等损失指导下学习特征,有效建模AUs与表情类别联系,实验证明该方法优于现有方法,凸显建模AUs关系的潜力。
AI 中文摘要
近年来,跨域少样本面部表情识别(CF-FER)备受关注。然而,由于在大域差异下可转移特征学习能力不足以及目标样本有限,现有CF-FER方法性能仍不尽人意。幸运的是,指示不同面部肌肉运动的动作单元(AUs)为描述域内和跨域表情提供了一致的概念语义。受此启发,我们提出了一种新颖的基于动作单元的一致性感知超图网络(AUCH-Net)用于CF-FER。具体而言,AUCH-Net提出了新的AU特征学习(AFL)模块和新的视觉特征学习(VFL)模块。AFL模块在新的关系一致性损失和AU正则化损失的指导下学习AU特征,而VFL模块在关系一致性损失和分类损失的监督下学习视觉特征。通过学习一致的AU特征,AUCH-Net有效地对AUs与表情类别之间的联系进行建模。实验表明,我们的方法始终优于几种现有方法。结果表明,在跨域少样本场景下,对AUs之间的关系进行建模在FER中具有巨大潜力。
英文摘要
Recently, cross-domain few-shot facial expression recognition (CF-FER) has received considerable attention. However, the performance of existing CF-FER methods is still unsatisfactory due to inferior transferable feature learning under large domain discrepancy and limited target samples. Fortunately, the action units (AUs), which indicate the movements of different facial muscles, provide consistent conceptual semantics for describing expressions within and across domains. Inspired by this, we propose a novel Action Unit-based Consistency-aware Hypergraph Network (AUCH-Net), which constructs consistency-aware hypergraphs on AUs, for CF-FER. Specifically, AUCH-Net presents a new AU feature learning (AFL) module and a new visual feature learning (VFL) module. The AFL module learns AU features under the guidance of a novel relation consistency loss and an AU regularization loss, while the VFL module learns visual features supervised by a relation consistency loss and a classification loss. By learning consistent AU features, AUCH-Net effectively models the connections between AUs and expression categories. As a result, we can bridge the gap between fine-grained facial variations and high-level expression categories, greatly facilitating the learning of transferable feature representations.Extensive experiments on both in-the-lab and in-the-wild datasets show that our method consistently outperforms several state-of-the-art methods. Our results clearly show that modeling the relationships among AUs holds significant potential for FER under cross-domain few-shot scenarios.