arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LogiScope-VQA:面向工业场景物流危险识别的视觉语言模型基准测试

LogiScope-VQA: Benchmarking Vision-Language Models for Logistics Hazard Identification in Industrial Scenarios

Hanjing Zhou, Mingze Yin, Ying Lian, Jun Ma, Chang-Yu Hsieh, Yanbing Zhou

arXiv 2609.09790首次发表:更新:

发表机构

Cainiao Group, Alibaba Group; Zhejiang University(菜鸟集团,阿里巴巴集团; 浙江大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

LogiScope-VQA构建了包含2,476张图像、2,918个视频和10,274个VQA的物流危险识别基准,通过39个子任务评估主流LMMs,发现即使GPT-5.5等专有模型与人类表现仍有显著差距,并揭示了安全偏差问题。

AI 中文摘要

大型多模态模型(LMMs)在工业仓储环境中的大规模部署,特别要求模型具备人类专家级别的面向危险的感知、理解和推理能力。然而,真实工业数据的稀缺性及其与商业条款的紧密耦合,严重阻碍了进一步的进展。为弥补这一差距,我们构建了LogiScope-VQA,以研究主流LMMs在真实物流运营中的实际适用性。LogiScope-VQA包含2,476张图像和2,918个视频,主要来源于真实物流园区,以及由人工标注者精心策划和验证的10,274个视觉问答(VQA)对。基于18个核心物体和20种风险类型,我们设计了39个子任务,与三个主要主题对齐:工业元素感知、仓储知识理解和潜在风险推理。此外,我们融入了动态思考预算配置和双维度风险偏差分析,以阐明LMMs的特性。大量实验揭示,即使是强大的专有模型,包括GPT-5.5、Gemini-3.1-Pro和Claude-Opus-4.7,与人类表现相比也存在显著差距。将感知、理解和推理联合集成以进行危险识别的独特挑战,为在LogiScope-VQA上的进一步改进提供了巨大空间。我们还揭示了普遍存在的安全偏差问题,该问题阻碍了LLMs在实际场景中的部署。该工业数据集在CC BY-NC-SA 4.0许可下公开可用。

英文摘要

Large Multimodal Models (LMMs) large-scale deployment in industrial warehouse settings specifically necessitates that models exhibit human-expert-level hazard-oriented perception, understanding, and reasoning capabilities. However, the scarcity of real industrial data, tightly coupled to commercial terms, significantly hampers further advancement. To bridge this gap, we curate LogiScope-VQA to investigate the practical applicability of mainstream LMMs in real-world logistics operations. LogiScope-VQA comprises 2,476 images and 2,918 videos primarily sourced from real-world logistics parks, along with 10,274 VQAs meticulously curated and validated by human annotators. Grounded in 18 core objects and 20 risk types, we devise 39 subtasks aligned with three principal themes: industrial element perception, warehouse knowledge understanding, and potential risk reasoning. Furthermore, we incorporate dynamic thinking-budget configurations and dual-dimensional risk bias analyses to elucidate the properties of LMMs. Extensive experiments unveil that even powerful proprietary models, including GPT-5.5, Gemini-3.1-Pro, and Claude-Opus-4.7, exhibit a significant gap relative to human performance. The unique challenge of jointly integrating perception, understanding, and reasoning for hazard identification poses substantial headroom for further improvement on LogiScope-VQA. We additionally reveal the pervasive security bias issue that impedes LLMs' practical deployment in real-world settings. The industrial dataset is publicly available under the CC BY-NC-SA 4.0 license.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑