发表机构
Pontifical Catholic University of Paraná(巴拉那天主教大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对第11届ABAW竞赛的矛盾/犹豫视频识别挑战,提出音频-文本系统,结合多种特征,通过时间交叉注意力融合音频与文本,在门控多实例学习前注入支持特征,五个独立模型预测平均后,在公共开发集上取得较好成绩。
AI 中文摘要
我们为第11届ABAW竞赛的矛盾/犹豫视频识别挑战赛提出了一个音频-文本系统。该方法排除视觉帧,将每个视频表示为与转录时间戳对齐的重叠5秒窗口。每个窗口结合了一个320维的韵律音频描述符、一个768维的面向情感的RoBERTa嵌入以及74个捕捉不确定性、模糊性和态度冲突的手工制作特征。音频和文本通过时间交叉注意力融合,同时在门控多实例学习(MIL)池化之前注入支持特征以调节窗口的重要性。对五个独立初始化的模型的预测进行平均。在有标签的公共开发集上,集成模型实现了0.875的平均精度和0.72的宏F1。我们的源代码可在这个https URL上公开获取。
英文摘要
We present a frame-independent audio-text system for the 3rd Ambivalence/Hesitancy Video Recognition Challenge at the 11th Affective & Behavior Analysis in-the-Wild (ABAW) Workshop. Videos are divided into overlapping 5-s windows aligned with transcript timestamps. Each window combines prosodic audio descriptors, emotion-oriented RoBERTa embeddings, and 74 psycholinguistic features representing uncertainty, hedging, and attitudinal conflict. Temporal cross-attention fuses audio and text, while the support features condition gated Multiple Instance Learning (MIL) pooling. A five-seed ensemble achieves an average precision of 0.875 and a macro-F1 of 0.722 on the 525-video labeled public-test split. Notably, our submission ranked third overall on the official challenge leaderboard, with a macro-F1 of 0.7455. Source code is available at https://github.com/Liga-de-IA-PUCPR/abaw-11-ah-challenge/.
CommentsAccepted for presentation at the 2026 European Conference on Computer Vision (ECCV) - ABAW Workshop