arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

视频语言模型真的在观看吗?诊断长视频中的角色跟踪失败

Do Video-LLMs Actually Watch? Diagnosing Character-Tracking Failures in Long-Form Video

Mohammad Al-Ratrout, Shayla Sharmin, Aditya Raikwar, Roghayeh Leila Barmaki

arXiv 2607.11078首次发表:更新:

AI 中文总结

研究视频语言模型在长视频中跟踪角色的能力,通过九条件诊断协议测试三个开源模型及Gemini,发现模型准确性非来自角色跟踪,存在性别线索利用问题,还发布诊断工具包揭示基准分数衡量的实际情况。

AI 中文摘要

视频大语言模型(Video-LLM)能否在长视频中跟踪一个人,准确记录其身份,并按顺序报告其在完整电视剧集中的服装变化?基准测试越来越多地对这类任务进行评分,最强的开源7-8B模型在InfiniBench的全局外观任务上达到了37-38%。但这个分数是来自跟踪指定角色,还是更简单的方法?我们用九条件诊断协议测试了三个架构不同的开源Video-LLM,以Gemini~2.5~Flash作为前沿参考,发现准确性并非来自角色跟踪。当改变问题中指定的角色时,模型仅在4-31%的时间里改变答案,它们很大程度上忽略了问题所问的是谁。按交换名字的性别的测试表明,当名字换成不同性别的角色时,模型反应更大,能捕捉到粗略的性别线索,但无法区分同性别的个体。当去掉多项选择选项并开放式提问时,开源模型的准确率下降18-25分,151个答案无一完全正确,而Gemini下降12分。进一步检查排除了添加字幕、使用信息最丰富的帧或增加帧数等明显无害的解释,瓶颈不在于模型看到多少视频,而在于如何将视频与问题中提到的人联系起来。我们发布了一个诊断工具包,用于审核此类基准分数实际衡量的内容。

英文摘要

Can a Video Large Language Model (Video-LLM) follow one person through a long video, keeping track of who they are well enough to report, in order, how their outfit changes across a full TV episode? Benchmarks increasingly score this kind of task, and the strongest open-source 7--8B models now reach 37--38% on InfiniBench's global appearance task, which asks exactly that. But does that score come from tracking the named character, or from something easier? We test this with a nine-condition diagnostic protocol applied to three architecturally distinct open-source Video-LLMs, with Gemini~2.5~Flash as a frontier reference, and find the accuracy does not come from character tracking. When we change the character named in the question to a different cast member, leaving the video and answer options untouched, the models change their answer only 4--31% of the time, so they are largely ignoring who the question asks about. Breaking that test down by the gender of the swapped name shows why: the models react more when the name is changed to a different-gender character than to a same-gender one (a 13--28 point gap), picking up coarse gender cues but unable to tell same-gender individuals apart. This shallow processing surfaces again when we drop the multiple-choice options and ask the same questions open-endedly: open-source accuracy drops 18--25 points, with none of 151 answers fully correct, versus a 12-point drop for Gemini. Further checks rule out the obvious innocent explanations, adding subtitles, using the most informative frames, or doubling the number of frames all leave character tracking unimproved, so the bottleneck is not how much video the model sees but how it ties that video to the person the question names. We release a diagnostic toolkit for auditing what such benchmark scores actually measure.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑