arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

顺序很重要:LVLMs作为图像序列时间推理的评判者

Order Matters: LVLMs as Judges for Temporal Reasoning in Image Sequences

Martina Ianaro, Guilherme Fernandes, Maurizio Gabbrielli, Joao Magalhaes

arXiv 2608.10908首次发表:更新:

发表机构

University of Bologna; NOVA School of Science and Technology; NOVA Laboratory for Computer Science and Informatics(博洛尼亚大学; NOVA科技学院; NOVA计算机科学与信息实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究发现大视觉语言模型(LVLMs)作为多模态评判者存在时间顺序判别缺陷,受首因、近因等结构性偏差影响,呼吁构建时间感知的视觉序列评估范式。

AI 中文摘要

随着生成式多媒体从静态图像合成发展到复杂的、交织的视觉叙事,一个基础性瓶颈已然出现:评判危机。人类感知能自然整合故事的时间与逻辑脉络,但自动评估系统在很大程度上对序列连续性“视而不见”,往往无法区分连贯叙事与语义混乱或矛盾的序列。本研究指出当前多模态评估范式存在关键结构性缺口,认为依赖大视觉语言模型(LVLMs)作为评判者的做法受限于架构偏差,存在根本性局限。我们的分析揭示了显著的性能二分现象:模型在孤立的逐点评分中看似具备能力,但在需要执行时间顺序的成对判别时会出现灾难性崩溃。我们证明这不仅是数据稀缺问题,更是结构性问题。通过一系列诊断探测,我们发现了系统性的位置不对称性,即首因效应与近因效应:模型对故事的判断会显著受帧的位置影响,其影响程度往往超过帧的语义一致性。这些偏差可能源于因果掩码与旋转位置嵌入,表明当前基于Transformer的评判者天生不适合长形式视觉推理。通过揭示这些盲点,我们呼吁多媒体领域超越以快照为中心的指标,开创时间感知评估范式,将视觉序列视为统一的逻辑结构,而非无序的帧集合。

英文摘要

As generative multimedia evolves from static image synthesis to complex, interleaved visual narratives, a foundational bottleneck has emerged: the judgment crisis. While human perception naturally synthesizes the temporal and logical flow of a story, automated evaluation systems remain largely "blind" to sequential continuity, often failing to distinguish between a coherent narrative and a semantically shuffled or contradictory sequence. This work identifies a critical structural gap in current multimodal evaluation paradigms, arguing that the reliance on Large Vision-Language Models (LVLMs) as judges is fundamentally limited by architectural biases. Our analysis reveals a profound performance dichotomy: while models may appear competent in isolated pointwise scoring, they suffer a catastrophic collapse when required to perform pairwise discrimination of temporal order. We demonstrate that this is not merely a data-scarcity issue but a structural one. Through a series of diagnostic probes, we uncover systematic positional asymmetries, specifically primacy and recency effects, where a model's judgment of a story is significantly influenced by the placement of a frame, often more than by its semantic consistency. These biases, potentially rooted in causal masking and rotary embeddings, suggest that current transformer-based judges are inherently ill-equipped for long-form visual reasoning. By exposing these blind spots, we challenge the multimedia community to move beyond snapshot-centric metrics and instead pioneer Temporally-Aware Evaluation paradigms that treat visual sequences as unified logical structures rather than unordered collections of frames.

Comments34 pages, camera-ready

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑