arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.32280cs.CV

HeroFrame-Bench:基于评分标准-排名协同演化的参考锚定电影英雄帧选择评估基准

HeroFrame-Bench: Reference-Anchored Evaluation via Rubric--Ranking Co-Evolution for Movie Hero Frame Selection

Weitai Kang, Hanieh Deilamsalehy, Yumo Xu, Dewang Sultania, Serdar Cellat, Yan Yan

首次发表
浏览论文内容

中文总结 AI 辅助

针对英雄帧选择任务,提出HeroFrame-Bench基准,采用VLM评判与参考锚定百分位数评估,并通过评分标准-排名协同演化提升一致性,实验显示人机一致性从77.56%提升至83.78%。

中文摘要 AI 辅助

英雄帧是电影中的静态画面,用作影院海报、流媒体封面、电影数据库条目及其他宣传位置的源图像。作为首个视觉入口,它们塑造了观众对电影的第一印象以及随后的观看意愿。选择这些帧(我们称之为英雄帧选择任务)需要在内容相关性与美学吸引力之间取得平衡。一个相关的任务是关键帧选择,但其基准优先考虑相关性而非美学,要么使用排除有效替代方案的有限标注,要么使用将选择质量与下游模型能力纠缠在一起的视频问答。因此,我们引入了HeroFrame-Bench,通过可扩展的VLM作为评判者的框架构建。我们从多样化的元数据中构建多模态上下文,以支持一个VLM评判者,该评判者使用我们的参考锚定百分位数对所选帧进行评分。该百分位数通过将每帧插入可重复使用的、预排名的参考链中获得,从而实现直接、可扩展且可靠的评估。为了减少这些主观判断中的歧义并提高一致性,我们进一步提出了评分标准-排名协同演化,它生成特定于电影的评分标准来调节VLM评判者,并联合优化评分标准与最终排名。在此过程中,我们引入了几个可验证的信号,最显著的是倒置评分标准攻击,以选择稳健的评分标准。最后,HeroFrame-Bench在204部电影上实例化,包含2,031条参考链和1,970个学习到的评分标准。我们构建了一个用于人类一致性研究的标注界面,研究表明我们的构建设计将VLM与人类的一致性从77.56%提高到83.78%。对多种方法的评估表明,英雄帧选择仍然具有挑战性。

英文摘要

Hero frames are in-film stills used as source imagery for theatrical posters, streaming cover art, film database listings, and other promotional placements. As the first visual entry point, they shape audiences' initial impressions of the movie and their subsequent willingness to watch it. Selecting these frames, a task we term hero frame selection, requires balancing content relevance with aesthetic appeal. A related task is keyframe selection, yet its benchmarks prioritize relevance over aesthetics, using either finite annotations that exclude valid alternatives or VideoQA that entangles selection quality with downstream model capability. We therefore introduce HeroFrame-Bench, built through a scalable VLM-as-a-Judge framework. We construct multimodal contexts from diverse metadata to ground a VLM judge that scores selected frames using our Reference-anchored Percentile. The percentile is obtained by inserting each frame into reusable, pre-ranked reference chains, enabling direct, extensible, and reliable evaluation. To reduce ambiguity and improve consistency in these subjective judgements, we further propose Rubric-Ranking Co-Evolution, which generates movie-specific rubrics to condition the VLM judge and refines rubrics jointly with the resulting rankings. Within this process, we introduce several verifiable signals, most notably the Inverted Rubric Attack, to select robust rubrics. Finally, HeroFrame-Bench is instantiated over 204 movies with 2,031 reference chains and 1,970 learned rubrics. We build an annotation interface for human-alignment studies which show that our construction design improves VLM agreement with human from 77.56% to 83.78%. Evaluation on multiple methods show that hero frame selection remains challenging.

补充信息

↑