AI 中文总结
针对现有视觉Transformer面部表情识别方法复杂度高难以边缘部署的问题,提出SAE模型,通过丢弃90%无价值图像令牌实现轻量FER,在RAF-DB数据集上达最优结果且复杂度降90%。
AI 中文摘要
面部表情识别(Facial Emotion Recognition, FER)是一项重要任务,在生物识别、医疗健康、人机交互等多个领域具有重要意义。当前基于视觉Transformer的方法呈现出二次复杂度$\boldsymbol{\textit{O}}(N^2)$,其中$N$为输入序列长度,这使得它们难以在边缘设备上部署。在本文中,我们假设FER任务并不需要所有面部信息来正确解读情感状态,因为眼睛、嘴巴以及部分脸颊等特定区域携带的判别性信息足以识别情感。基于此,我们提出了面向情感的稀疏注意力(Sparse Attention to Emotion, SAE)模型,该模型会丢弃对情感上下文无附加价值的图像令牌,同时保持良好的准确率并显著降低计算成本。令人惊讶的是,即使在抑制90%的图像令牌后,我们的模型仍达到了与现有最优方法相当的准确率,且成本大幅降低,提供了一种轻量级的面部表情识别方法。实验结果表明,SAE在RAF-DB数据集上取得了新的最优结果,同时将计算复杂度降低了多达90%。
英文摘要
Facial Emotion Recognition (FER) is an important task that has significant implications across various fields such as biometrics, health, and human-computer interaction. Current Vision Transformer-based approaches display quadratic complexity $\mathcal{O}(N^2)$, with N being the input sequence length, making them cumbersome to deploy at the edge. In this paper, we hypothesize that the FER task does not necessarily require all facial information to correctly interpret emotional states, as specific regions such as the eyes, the mouth, and parts of the cheeks carry discriminative information that can be sufficient to recognize emotions. Based on this, we propose Sparse Attention to Emotion (SAE), a model that discards image tokens that have no added value to the emotional context, while preserving good accuracy and achieving a significant gain in computational cost. Surprisingly, even after suppressing 90\% of the image tokens, our model achieves competitive accuracy to state of the art methods at much lower cost, providing a lightweight Facial Emotion Recognition approach. Experimental results demonstrate that SAE achieves new state of the art results on the RAF-DB dataset while reducing the computational complexity by up to 90\%.
Comments6 pages, 2 figures, published at ICIP 2026