arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于矛盾/犹豫识别的具有心理语言学支持特征的音频-文本交叉注意力

Audio-Text Cross-Attention with Psycholinguistic Support Features for Ambivalence/Hesitancy Recognition

Luiz F. B. F. Martins, Rodrigo W. Pisaia, Matheus M. Girardi, Isabella V. Berkembrock, João A. Almeida, Andre G. Hochuli, Rayson Laroca, Alceu S. Britto

arXiv 2607.13345首次发表:更新:

发表机构

Pontifical Catholic University of Paraná(巴拉那天主教大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对第11届ABAW竞赛的矛盾/犹豫视频识别挑战,提出音频-文本系统,结合多种特征,通过时间交叉注意力融合音频与文本,在门控多实例学习前注入支持特征,五个独立模型预测平均后,在公共开发集上取得较好成绩。

AI 中文摘要

我们为第11届ABAW竞赛的矛盾/犹豫视频识别挑战赛提出了一个音频-文本系统。该方法排除视觉帧,将每个视频表示为与转录时间戳对齐的重叠5秒窗口。每个窗口结合了一个320维的韵律音频描述符、一个768维的面向情感的RoBERTa嵌入以及74个捕捉不确定性、模糊性和态度冲突的手工制作特征。音频和文本通过时间交叉注意力融合,同时在门控多实例学习(MIL)池化之前注入支持特征以调节窗口的重要性。对五个独立初始化的模型的预测进行平均。在有标签的公共开发集上,集成模型实现了0.875的平均精度和0.72的宏F1。我们的源代码可在这个https URL上公开获取。

英文摘要

We present a frame-independent audio-text system for the 3rd Ambivalence/Hesitancy Video Recognition Challenge at the 11th Affective & Behavior Analysis in-the-Wild (ABAW) Workshop. Videos are divided into overlapping 5-s windows aligned with transcript timestamps. Each window combines prosodic audio descriptors, emotion-oriented RoBERTa embeddings, and 74 psycholinguistic features representing uncertainty, hedging, and attitudinal conflict. Temporal cross-attention fuses audio and text, while the support features condition gated Multiple Instance Learning (MIL) pooling. A five-seed ensemble achieves an average precision of 0.875 and a macro-F1 of 0.722 on the 525-video labeled public-test split. Notably, our submission ranked third overall on the official challenge leaderboard, with a macro-F1 of 0.7455. Source code is available at https://github.com/Liga-de-IA-PUCPR/abaw-11-ah-challenge/.

CommentsAccepted for presentation at the 2026 European Conference on Computer Vision (ECCV) - ABAW Workshop

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑