发表机构
Rensselaer Polytechnic Institute(伦斯勒理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究旨在实现零样本公交视频分析,提出GHR-VLM框架,利用边缘-云设计,通过轻量级边缘监视器跟踪门状态、分割乘客片段,后端VLM经两阶段细化识别乘客与支付行为,减少云推理,评估显示其在公交支付分析中有潜力且面临视频条件挑战。
AI 中文摘要
公交视频理解能提供传统乘客计数器和票务系统无法获取的有价值细粒度数据。然而,监督视频模型需要特定任务注释,直接将视觉语言模型(VLM)应用于长车载视频既不可靠又成本高。为利用两种方法的互补优势,我们提出GHR-VLM,一种用于零样本公交视频分析的视觉基础混合推理框架。它通过将长监控流转换为紧凑的、以乘客为中心的时空证据,明确的视觉基础可改善VLM推理。具体而言,我们提出一种边缘-云设计,其中轻量级边缘监视器持续跟踪门状态并分割乘客片段。后端VLM随后通过时空证据的两阶段粗到精细化识别上车乘客并对支付行为进行分类。通过仅在基础乘客片段和联系表上调用VLM,GHR-VLM减少了云推理,避免了特定支付训练数据,并提供了VLM难以识别的局部证据。对486分钟真实世界公交监控视频的评估证明了基于视觉基础的边缘-云推理在乘客级支付分析中的潜力,同时突出了视频条件退化带来的挑战。
英文摘要
Transit video understanding can provide valuable fine-grained data that conventional passenger counters and fare systems cannot capture. However, supervised video models require task-specific annotations, while applying vision-language models (VLMs) directly to long onboard videos is unreliable and costly. To leverage the complementary strengths of both approaches, we propose GHR-VLM, a visual grounded hybrid reasoning framework for zero-shot transit-bus video analytics. It is motivated by the observation that explicit visual grounding can improve VLM reasoning by converting long surveillance streams into compact, passenger-centered spatiotemporal evidence. Specifically, we propose an edge-cloud design in which a lightweight edge-based monitor continuously tracks door status and segments passenger clips. A backend VLM then identifies boarding passengers and classifies payment behavior through a two-stage coarse-to-fine refinement of spatiotemporal evidence. By invoking the VLM only on grounded passenger clips and contact sheets, GHR-VLM reduces cloud inference, avoids payment-specific training data, and supplies the localized evidence that VLMs otherwise struggle to identify. Evaluation on 486 minutes of real-world bus surveillance video demonstrates the potential of grounded edge-cloud reasoning for passenger-level payment analytics while highlighting the challenges posed by degraded video conditions.
CommentsAccepted by 2026 IEEE/ACM Symposium on Edge Computing (SEC)