发表机构
eHealth Center, Faculty of Computer Science, Multimedia and Telecommunicactions, Universitat Oberta de Catalunya; MIND/IN2UB, Department of Electronic and Biomedical Engineeering, Universitat de Barcelona; Institute of Semiconductor Technology (IHT) & Laboratory for Emerging Nanometrology (LENA), Technische Universität Braunschweig(加泰罗尼亚开放大学计算机科学、多媒体与电信学院电子健康中心; 巴塞罗那大学电子与生物医学工程系MIND/IN2UB; 布伦瑞克工业大学半导体技术研究所(IHT)和新兴纳米计量实验室(LENA))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对BAH数据集上的ABAW A/H挑战,提出校准多模态集成系统,由三个融合模型组成。在公共测试集上宏F1达0.7358,介绍五次私密测试提交的配置等,首次提交在私密测试中宏F1为0.7361,验证系统能推广到未见参与者。
AI 中文摘要
矛盾和犹豫(A/H)会破坏数字行为改变干预措施,从视频中自动识别它们是BAH数据集上ABAW A/H挑战的目标。我们描述了针对第11版挑战的系统:一个基于冻结的面部、音频、文本和姿势嵌入的三个融合模型的校准等权重集成,在公共测试集上达到0.7358的宏F1。今年的私密测试在30名新参与者的不相交集上进行,基于五次允许的提交进行评分;我们报告了五次提交中每次的配置和基本原理,以及已获得的私密测试分数。我们的首次提交是仅在公共验证上调整的校准集成的精确副本,在私密测试中获得了0.7361的宏F1,几乎与我们的公共测试估计完全匹配,并证实了该管道可以推广到未见参与者而无泄漏。
英文摘要
Ambivalence and hesitancy (A/H) undermine digital behaviour-change interventions, and recognizing them automatically from video is the goal of the ABAW A/H challenge on the BAH dataset. We describe HEDGE (Hesitancy/Ambivalence Estimation via Distribution-aware, Generalized Ensembling), our system for the 11th edition of the challenge: a calibrated, equal-weight ensemble of three fusion models over frozen face, audio, text, and pose embeddings, which reaches 0.7358 macro-F1 on the public test set. We also submitted four variants of this system to this year's private test (30 new participants): a fixed-threshold version, a single individual model, and an ensemble with cache-personalization test-time adaptation (TTA). The plain calibrated ensemble scored 0.7361 macro-F1 on the private test, closely matching our public-test estimate, and the TTA variant scored highest of all five at 0.7367 macro-F1, our official challenge result (team AIWELL, rank 5 of 12), even though TTA showed no benefit on the public test. The single individual model dropped to 0.6759, far more than any ensemble variant. We explain both results: TTA only has distribution shift to correct on the private test, which the public test lacks, and a text-only linear probe reaches 0.716 macro-F1 (within noise of the full system, correlated at 0.91 in its errors), so the ensemble's robustness to new participants comes from the same modality redundancy that makes single, less-diversified models comparatively brittle. We additionally report a systematic study of more than 60 further controlled experiments (modality, backbone, loss, and adaptation ablations) that did not improve on this system, and an explainability analysis showing the transcript's delivery style dominates the signal while the extractable non-verbal ceiling saturates near 0.60 macro-F1.
Comments8 pages, 1 figure, ECCV workshops