arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SoftRerank:面向长尾微动作识别的候选标签重排序分层软融合方法

SoftRerank: Hierarchical Soft Fusion with Candidate-Label Reranking for Long-Tailed Micro-Action Recognition

Yichi Zhang, Zhichao Xia, Yanjun Chi, Lingsi Zhu, Yuefeng Zou, Jun Yu, Qingsong Liu, Jianqing Sun, Shengping Liu

arXiv 2609.08221首次发表:更新:

发表机构

University of Science and Technology of China; Unisound AI Technology Co., Ltd.(中国科学技术大学; 云知声人工智能科技有限公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出SoftRerank方法,结合InternVideo2.5全参数微调、分层软融合和候选标签重排序,解决长尾微动作识别难题,在MA-52上取得79.99%的F1均值并获ACM Multimedia 2026挑战赛第一名。

AI 中文摘要

微动作是细微、低强度的非语言行为,为包括情绪和意图在内的细粒度人类状态提供线索。由于微动作持续时间短、视觉变化微弱,且不同类别之间常呈现相似的运动模式,识别它们仍然困难。本文针对这些挑战提出了一种细粒度微动作识别方法,该方法结合了InternVideo2.5的全参数微调、分层软融合以及轻量级候选标签重排序器。针对MA-52中的长尾标签分布,我们使用类别平衡采样和逆频率重加权来减少训练过程中频繁出现类别的影响。我们对InternVideo2.5进行端到端微调,并在共享视频表示上附加粗粒度头和组条件细粒度分类头,以提高粗粒度与细粒度预测之间的一致性。对于模糊样本,候选标签重排序器利用困难样本和视频-标签匹配来聚焦于易混淆的细粒度动作。实验验证了所提方法的有效性,其在MA-52上取得了79.99%的F1均值,并在ACM Multimedia 2026第三届微动作分析大挑战赛中排名第一。

英文摘要

Micro-actions are subtle, low-intensity non-verbal behaviors that provide cues to fine-grained human states, including emotions and intentions. Recognizing them remains difficult because they are brief, contain weak visual changes, and often exhibit similar motion patterns across categories. This paper addresses these challenges with a fine-grained micro-action recognition method that combines full fine-tuning of InternVideo2.5, hierarchical soft fusion, and a lightweight candidate-label reranker. For the long-tailed label distribution in MA-52, we use class-balanced sampling and inverse-frequency reweighting to reduce the effect of frequent classes during training. We fine-tune InternVideo2.5 end to end and attach coarse and group-conditional fine-grained classification heads to the shared video representation, improving the consistency between coarse and fine predictions. For ambiguous samples, the candidate-label reranker uses hard samples and video-label matching to focus on easily confused fine-grained actions. Experiments validate the proposed method, which achieves a 79.99% F1-mean on MA-52 and ranks first in the 3rd Micro-Action Analysis Grand Challenge at ACM Multimedia 2026.

Comments7 pages, 2 figures, 3 tables. Accepted to the 34th ACM International Conference on Multimedia (MM '26). Ranked 1st in the 3rd Micro-Action Analysis Grand Challenge at ACM MM 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑