arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

EgoSafe:用于视觉安全理解的第一人称移动采集基准

EgoSafe: A First-Person Mobile-Captured Benchmark for Visual Safety Understanding

Yuyun Chen, Tianao Li, TianQuan Feng, Cen Chen, Huiping Zhuang, Hao Peng, Ziqian Zeng

arXiv 2607.26518首次发表:更新:

发表机构

South China University of Technology; Beihang University(华南理工大学; 北京航空航天大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究推出第一人称移动采集的视觉安全理解基准 EgoSafe-Bench,含12000个样本,用分层推理评估(HRE)协议测试,发现现有大视觉语言模型存在感知-推理解耦问题,为逻辑鲁棒视频理解系统提供评估框架

AI 中文摘要

真实场景中可靠的视觉安全理解不仅需要物体识别,还需要在认知不确定性下进行因果推理。尽管大视觉语言模型(LVLMs)在标准基准上展现出出色的语义对齐能力,但当基于第一人称体验的动态、部分可观测特性进行 grounding 时,它们往往难以区分表面相关性与真正的取证逻辑。现有评估多以第三人称监控视频和二分类指标为主,无法暴露这种认知差距。为解决该问题,我们推出 EgoSafe-Bench,一个专门用于探测自我中心安全场景中取证推理的基准。它包含12000个独特评估样本,通过将3000个视频片段中的每一个与我们提出的分层推理评估(HRE)协议控制的问答链配对生成。与标准基准不同,HRE 要求从初始特征锚定到盲区推断和意图推理的严格推理轨迹,从而强制逻辑一致性并惩罚基于捷径的评估。我们对最先进的 LVLMs(如 Qwen3-VL、Gemini、VideoLLaMA 3)的评估显示存在显著的感知-推理解耦:模型常获得较高描述性分数,但在因果推理和逻辑闭合上表现出明显脆弱性。我们的工作提供了一个具有挑战性的数据集和一个系统的评估框架,以促进逻辑鲁棒视频理解系统的开发。

英文摘要

Reliable visual safety understanding in real-world scenarios demands more than just object recognition; it requires causal reasoning under epistemic uncertainty. While Large Vision-Language Models (LVLMs) demonstrate impressive semantic alignment on standard benchmarks, they often struggle to distinguish between superficial correlation and genuine forensic logic when grounded in the dynamic, partially observable nature of first-person experiences. Existing evaluations, dominated by third-person surveillance footage and binary classification metrics, fail to expose this cognitive gap. To address this, we introduce EgoSafe-Bench, a benchmark specifically designed to probe forensic reasoning in egocentric safety scenarios. It comprises 12,000 unique evaluation samples, generated by pairing each of the 3,000 video clips with a QA chain governed by our proposed Hierarchical Reasoning Evaluation (HRE) protocol. Unlike standard benchmarks, HRE mandates a rigorous reasoning trajectory from initial feature anchoring to blind-spot deduction and intent inference, thereby enforcing logical consistency and penalizing shortcut-based predictions. Extensive evaluations of state-of-the-art LVLMs (e.g., Qwen3-VL, Gemini, VideoLLaMA 3) reveal a significant perception-reasoning decoupling: models often achieve high descriptive scores but exhibit notable fragility in causal reasoning and logical closure. Our work provides both a challenging dataset and a systematic evaluation framework to foster the development of logically robust video understanding systems.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑