arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越匿名字幕:在视频字幕生成与问答中锚定角色身份

Beyond Anonymous Captions: Grounding Character Identity in Video Captioning and Question Answering

Anas Filali Razzouki, Killian Steunou, Khalil Guetari, Thomas Kling, Mounîm El-Yacoubi, Yannis Tevissen

arXiv 2610.10163首次发表:更新:

发表机构

Télécom SudParis, Institut Polytechnique de Paris; Moments Lab Research(巴黎电信学院,巴黎综合理工学院; Moments Lab 研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出身份感知视频字幕与问答框架,通过自动角色识别与空间锚定,结合LoRA微调,显著提升身份理解,BAC-8B达93.20%准确率。

AI 中文摘要

将人物的外貌和动作与角色身份关联起来,对于理解视频叙事至关重要。我们提出了一个用于身份感知视频字幕生成和以人为中心的问答的框架,该框架结合了自动角色识别、显式空间锚定和任务特定自适应。从LSMDC v2电影片段出发,我们的流程将检测到的人脸与演员参考图像进行匹配,跨帧跟踪角色,并构建带有身份关联边界框的输入。一个强大的视觉语言模型生成身份感知的字幕和问题,经过人工验证和筛选,创建了一个包含750个带字幕片段和3000个以人为中心的问题的基准。我们研究了五种锚定策略,结合文本坐标与视觉人脸框或估计的人物框,应用于约2B、4B和8B参数规模的Video-MLLM系列模型以及更大的前沿模型。将视觉人脸框与文本坐标相结合,在不同规模上产生了最一致的表现,并显著优于仅使用坐标的整体性能。较小的模型在被查询人物未被锚定时倾向于过度分配已知身份,而较大的模型能更好地识别此类未识别案例。我们通过LoRA微调Qwen模型在约32K个身份感知字幕片段上,在2B、4B和8B规模上引入了BAC。在所有规模上,BAC都优于每个其他评估的同等规模模型系列。BAC-8B达到了93.20%的整体问答准确率,在我们研究评估的前沿模型中仅次于GPT-5.6 Sol。总体而言,显式传达谁在哪里,加上轻量级的任务特定自适应,在不改变底层架构的情况下显著提升了身份感知视频理解。我们在以下https URL发布了基准、训练数据、代码和BAC检查点。

英文摘要

Linking people's appearance and actions to character identities is essential for understanding video narratives. We present a framework for identity-aware video captioning and person-centric question answering that combines automatic character identification, explicit spatial grounding, and task-specific adaptation. Starting from LSMDC v2 movie clips, our pipeline matches detected faces to actor reference images, tracks characters across frames, and builds inputs with identity-linked bounding boxes. A strong vision-language model generates identity-aware captions and questions, which are manually verified and filtered to create a benchmark of 750 captioned clips and 3,000 person-centric questions. We study five grounding strategies combining textual coordinates with visual face or estimated person boxes across Video-MLLM families at roughly 2B, 4B, and 8B parameters and larger frontier models. Combining visual face boxes with textual coordinates yields the most consistent performance across scales and significantly improves overall performance over coordinates alone. Smaller models tend to over-assign known identities when the queried person is not grounded, while larger models better recognize such UNIDENTIFIED cases. We introduce BAC by LoRA fine-tuning Qwen models at 2B, 4B, and 8B scales on about 32K identity-aware captioned clips. Across all scales, BAC outperforms every other evaluated model family of comparable size. BAC-8B reaches 93.20\% overall QA accuracy, ranking behind only GPT-5.6 Sol among the frontier models evaluated in our study. Overall, explicitly communicating who is where, together with lightweight task-specific adaptation, substantially improves identity-aware video understanding without changing the underlying architecture. We release the benchmark, training data, code, and BAC checkpoints at https://github.com/momentslab/beyond-anonymous-captions.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑