发表机构
School of Mechanical Engineering, University of Science and Technology Beijing; The Laboratory for Computational Sensing and Robotics, Johns Hopkins University(北京科技大学机械工程学院; 约翰霍普金斯大学计算传感与机器人实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出MonitorVLM-v2框架,将VLM安全评估转化为有限符号空间概率推理,结合SymPO算法与熵驱动分类机制,在地下矿部署中实现19.45倍提速,确认违规量达人工检查的2.78倍,满足实时可审计工业监控需求。
AI 中文摘要
大型视觉-语言模型(VLMs)可对复杂视觉场景进行逐步推理,但这种开放式、自回归思维链(CoT)方法并不适用于安全关键、规则驱动的场景,如工业监控,此类场景要求决策有界、确定且低延迟。由于CoT推理成本随推理长度和并发流数量共同增长,它会产生吞吐量瓶颈,无法满足工业问责所需的实时多流监控需求。本文提出MonitorVLM-v2,这是一种面向部署的框架,将基于VLM的安全评估重新定义为有限监管决策空间上的概率推理,将多模态推理压缩为单步规则ID预测,将解码从可变长度序列简化为单个token。我们引入符号策略优化(SymPO),一种新型对比策略优化算法,用于锐化此有限符号空间内的决策边界,同时引入熵驱动分类机制,将不确定预测路由至人工审核员进行专家确认。在一个运营中的地下采矿设施开展的为期四个月、覆盖10个并发摄像头流的前瞻性部署中,MonitorVLM-v2实现了19.45倍的推理速度提升,识别出的已确认违规数量是该场地常规人工检查工作流程的2.78倍,证明了压缩符号决策在实时、可审计工业监控中的实用价值。
英文摘要
Large vision--language models (VLMs) can reason step by step about complex visual scenes, but this open-ended, autoregressive chain-of-thought (CoT) approach is poorly suited to safety-critical, rule-governed settings such as industrial surveillance, where decisions must be bounded, deterministic, and low-latency. Because CoT inference cost scales jointly with reasoning length and the number of concurrent streams, it creates a throughput bottleneck that precludes the real-time, multistream monitoring required for industrial accountability. Here we present MonitorVLM-v2, a deployment-oriented framework that recasts VLM-based safety assessment as probabilistic inference over a finite regulatory decision space, compressing multimodal reasoning into single-step rule-ID predictions and reducing decoding from a variable-length sequence to a single token. We introduce symbolic policy optimization (SymPO), a novel contrastive policy optimization algorithm that sharpens decision boundaries within this finite symbolic space, together with an entropy-driven triage mechanism that routes uncertain predictions to human reviewers for expert confirmation. In a four-month prospective deployment across 10 concurrent camera feeds in an operational underground mining facility, MonitorVLM-v2 achieved a 19.45-fold increase in inference speed and identified 2.78 times as many confirmed violations as the site's routine manual inspection workflow, demonstrating the practical value of compressed symbolic decision-making for real-time, auditable industrial monitoring.