Open Your Model's Eyes: Video and Context-Aware Multimodal Backchannel Prediction
睁开模型之眼:视频与上下文感知的多模态反馈预测
机构 * Korea University(韩国大学) ; Kyung Hee University(庆熙大学)
AI总结 研究视频与上下文感知的多模态反馈预测问题,提出CAMA - BC框架,利用多层多模态对齐,分上下文对齐和反馈对齐两阶段,显著优于现有方法及简单多模态基线,尤其在识别复杂反馈上效果突出。
Comments 18 pages, Accepted at ACL 2026
Journal ref Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Association for Computational Linguistics, pp. 3738-3755, 2026