arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.02510cs.CVcs.HCcs.LG

表演者无关的人体运动情感识别的正交集成与经检验的解释

Orthogonal Ensembles and Tested Explanations for Performer-Independent Body-Motion Emotion Recognition

Naoto Nishida, Yoshio Ishiguro

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对留表演者评估的人体运动情感识别,提出结合11种正交误差模式模型的集成方法提升性能,还开发事后解释工具揭示决策依赖身体区域证据,与拉班动作分析对齐度远高于经典运动学。

中文摘要 AI 辅助

我们研究仅利用骨骼运动进行的12类表演情感分类,采用留表演者(LPO)评估这一困难且欠定的设置:随机猜测准确率为8.3%,按协议复现的STGCN++基线仅达到25.73±4.03%的Macro-F1。我们证明可靠的性能提升并非来自新架构,而是来自11种具有正交误差模式的模型的组合:在标记训练表演者上的10折LPO交叉验证下,等权重logit平均集成达到每折36.80±4.00%的Macro-F1,相比同划分复现基线提升了11.07个百分点(相对提升43%)。我们的核心贡献是一套经检验的解释工具:针对一个强集成成员,通过部分掩码和反事实编辑,可展示(而非断言)其决策依赖于基于运动的身体区域证据,且该区域显著性与基于规则的拉班动作分析(LMA)属性的对齐程度远高于与经典运动学的对齐:区域级显著性-LMA的斯皮尔曼相关系数ρ为+0.500,而与经典运动学仅为+0.033,相差约15倍,且提交的11路集成本身也保持这种对齐,ρ为+0.517;该审计为事后分析,无需重新训练。同一工具还如实报告了一个负面结果:窗口内时间显著性是弥散的,而非局部化的。

英文摘要

We study body-only, 12-class acted-emotion classification from skeleton motion under leave-performer-out (LPO) evaluation, a hard, underdetermined setting: chance is 8.3%, and a protocol-matched reproduced STGCN++ baseline reaches only 25.73 +/- 4.03% Macro-F1. We show that reliable gains come not from a new architecture but from combining eleven models with orthogonal error modes: under 10-fold LPO cross-validation on the labeled training performers, an equal-weight logit-mean ensemble reaches 36.80 +/- 4.00% per-fold Macro-F1, a protocol-matched +11.07 pp (+43% relative) over the same-split reproduced baseline. Our central contribution is a tested explanation suite: for a strong ensemble member, part-masking and counterfactual edits show (rather than assert) that its decisions depend on motion-grounded body-region evidence, and this region saliency aligns with rule-based Laban Movement Analysis (LMA) attributes far more than with classical kinematics: region-level saliency-LMA Spearman rho = +0.500 versus +0.033, roughly 15x, and the alignment holds for the submitted 11-way ensemble itself at rho = +0.517; the audit is post hoc and needs no retraining. The same suite faithfully reports a negative: within-window temporal saliency is diffuse rather than localized. On the hidden challenge test set the submitted ensemble scored 37.23 % Macro-F1 and received the Best Performance Award of the MMAC Challenge 2026 (Human score is 39 %). Code is available at https://github.com/nawta/diema-challenge and the presentation at https://nawta.github.io/mmac2026/.

发表机构

  • The University of Tokyo(东京大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑