TellTale:融合多实例LoRA文本编码器和零样本大语言模型判断器用于视频中的矛盾/犹豫识别
TellTale: Blending Multi-Instance LoRA Text Encoders and a Zero-Shot LLM Judge for Ambivalence/Hesitancy Recognition in Videos
浏览论文内容
中文总结 AI 辅助
TellTale旨在解决视频中矛盾/犹豫识别问题,它融合多实例LoRA文本编码器和零样本大语言模型判断器,结合三个概率流,在BAH数据集上取得优异成绩,相比官方基于视觉的基线有显著提升。
中文摘要 AI 辅助
我们提出了TellTale,一种用于访谈视频中矛盾/犹豫(A/H)识别的纯文本方法,并在作为第三届A/H视频识别挑战赛(第11届ABAW研讨会,ECCV 2026)一部分的BAH数据集上进行评估。尽管数据集提供了视频、音频、面部裁剪和转录本,但TellTale仅依赖转录本并结合三个概率流。两个文本编码器在多实例学习(MIL)目标下用参数高效的LoRA适配器进行微调,转录本块单独评分并通过平滑最大值合并,仅需视频级标签进行监督。第三个流无需训练,通过零样本提示量化的14B指令大语言模型对每个转录本进行A/H评分。三个概率通过加权平均和单个决策阈值组合,两者均在按参与者分组的交叉验证预测上选择。在来自未见过的参与者的152个视频的组织者评分的私人测试集上,TellTale实现了0.7364的宏F1和0.7940的平均精度,而官方基于视觉的基线宏F1为0.2827。
英文摘要
We present TellTale, a text-only approach to ambivalence/hesitancy (A/H) recognition in interview videos, evaluated on the BAH dataset as part of the 3rd A/H Video Recognition Challenge (11th ABAW Workshop, ECCV 2026). Although the dataset provides video, audio, facial crops, and transcripts, TellTale relies on the transcript alone and combines three probability streams. Two text encoders, multilingual-e5-large and mDeBERTa-v3-base, are fine-tuned with parameter-efficient LoRA adapters under a multiple-instance learning (MIL) objective, in which transcript chunks are scored individually and pooled with a smooth maximum so that only the video-level label is needed for supervision. The third stream requires no training: a quantized 14B instruction LLM is prompted, zero-shot, to rate each transcript for A/H. The three probabilities are combined by a weighted average and a single decision threshold, both selected on participant-grouped cross-validated predictions. On the organizer-scored private test set of 152 videos from unseen participants, TellTale achieves a Macro-F1 of 0.7364 and an average precision of 0.7940, compared with 0.2827 Macro-F1 for the official vision-based baseline.
发表机构
- Sheffield Hallam University(谢菲尔德哈勒姆大学)
- Yarmouk University(亚尔穆克大学)
- The University of Jordan(约旦大学)
机构由 AI 辅助整理,请以论文原文为准。