arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.31823cs.CVcs.AI

SynDORBench:物理约束可见性条件下LVLM感知鲁棒性评估

SynDORBench: Evaluating LVLM Perceptual Robustness Under Physically Constrained Visibility Conditions

Jeremy Stephen Gabriel Yee, Zhengkui Wang, Zhiyuan Zhang, Avinash Anand, Timothy Liu, Benedict Chan, Aik Beng Ng, Simon See

首次发表
浏览论文内容

中文总结 AI 辅助

SynDORBench是首个物理基础基准,通过DORI校准和54k问答对评估LVLM在像素密度等物理约束下的感知鲁棒性,发现紧凑开源模型优于商业基线。

中文摘要 AI 辅助

大型视觉-语言模型(LVLMs)在多模态推理基准上表现出显著性能,但其在物理约束成像条件下的感知可靠性仍未被充分理解。现有评估主要假设理想视觉输入,因此无法刻画相机距离、光照、视角和像素密度如何从根本上影响语义可恢复性。我们引入SynDORBench,这是首个在DORI校准条件下、与人类视觉能力标准对齐的、用于评估LVLM感知鲁棒性的物理基础基准。SynDORBench包含超过54,000个问答对,通过可控合成流程生成,该流程根据物理可解释的像素密度机制系统变化观看距离、光照、相机几何和动作姿态。为支持可扩展的低可见性监督,我们进一步提出一个可辨性标注框架,利用掩码条件统计特征和集成学习传播人类感知标签。我们在逐渐退化的可见性条件下,评估了16个开源LVLM、一个商业LVLM基线以及YOLO11x在人类存在分类和动作识别任务上的表现。我们的结果显示,LVLM的感知失败主要由像素密度和物理成像约束决定,而非仅由模型规模决定。令人惊讶的是,几个紧凑的开源LVLM在远距离和低光照条件下优于更大的商业基线,并显著超过YOLO11x的鲁棒性。SynDORBench为物理基础的多模态评估建立了新的基准范式,能够在真实世界感知约束下系统分析LVLM可靠性,并与人类可见性阈值进行直接比较。

英文摘要

Large vision-language models (LVLMs) have demonstrated remarkable performance on multimodal reasoning benchmarks, yet their perceptual reliability under physically constrained imaging conditions remains poorly understood. Existing evaluations predominantly assume ideal visual inputs and therefore fail to characterize how camera distance, illumination, viewpoint, and pixel density fundamentally affect semantic recoverability. We introduce SynDORBench, the first physically grounded benchmark for evaluating LVLM perceptual robustness under DORI-calibrated conditions aligned with human visual capability standards. SynDORBench comprises over 54k question--answer pairs generated through a controllable synthetic pipeline that systematically varies viewing distance, lighting, camera geometry, and action pose according to physically interpretable pixel-density regimes. To support scalable low-visibility supervision, we further propose a discernibility annotation framework that propagates human perceptual labels using mask-conditioned statistical features and ensemble learning. We evaluate 16 open-source LVLMs, a commercial LVLM baseline, and YOLO11x across human-presence classification and action recognition tasks under progressively degraded visibility conditions. Our results reveal that perceptual failure in LVLMs is strongly governed by pixel density and physical imaging constraints rather than model scale alone. Surprisingly, several compact open-source LVLMs outperform larger commercial baselines and substantially exceed YOLO11x robustness under long-range and low-light conditions. SynDORBench establishes a new benchmark paradigm for physically grounded multimodal evaluation, enabling systematic analysis of LVLM reliability under real-world perceptual constraints and direct comparison against human visibility thresholds.

↑