arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.19515eess.IVcs.CV

BLUE:用于高效视觉语言监控分析的语义保留视频压缩

BLUE: Semantics-Preserving Video Compression for Efficient Vision-Language Surveillance Analytics

  • KGraph AI Solutions Pvt. Ltd.(KGraph AI解决方案私人有限公司)

机构由 AI 辅助整理,请以论文原文为准。

Shubham Baid, Akash James, Sahil Chachra, Nishant Sinha, Kunal Kislay

AI总结:

研究持续监控视频给企业视频分析系统带来的负担问题,提出BLUE固定摄像头监控压缩方法,在VIRAT和CHAD数据集上实验,结果表明该方法能在保留VLM语义性能时降低带宽和推理成本。

AI中文摘要:

持续的监控视频给企业视频分析系统带来了日益增长的存储、传输和推理负担。虽然H.265等现代编解码器可降低供人观看视频的比特率,但过度压缩会降低下游计算机视觉性能,且不一定减少语义视频理解所需的视觉语言模型(VLM)推理调用次数。本文评估了BLUE(一种固定摄像头监控压缩方法)对基于VLM的事件和异常理解的影响,该方法在保留前景活动的同时抑制静态背景冗余。我们在两个监控数据集上比较了原始H.265和经BLUE压缩的H.265视频:VIRAT(包含来自106个片段的227对事件样本)和CHAD(包含54个人类活动异常片段)。对于每一对,使用VLM字幕管道评估相同的帧索引,并使用盲判协议根据注释得出的地面真值对输出进行评分。结果表明语义推理质量没有可测量的下降。在VIRAT上,原始H.265和BLUE之间的平均VLM分数基本保持不变,在0-10的尺度上平均差异约为-0.01。在CHAD上,原始H.265和BLUE分别获得了接近等效的平均分数4.31和4.26。压缩节省也与VIRAT上的VLM分数变化无关(r = 0.004),这表明更高的BLUE压缩不会预测语义质量损失。除了减少存储外,BLUE还将CHAD上重跳过P帧的比例从1.4%提高到53.2%,通过基于数据包大小的帧跳过估计可减少53%的VLM调用。这些发现表明,BLUE作为监控视频的以机器为中心的压缩层,在保留VLM语义性能的同时降低了带宽和推理成本。

英文摘要:

Continuous surveillance video creates a growing storage, transmission, and inference burden for enterprise video analytics systems. While modern codecs such as H.265 reduce bitrate for human-viewable video, aggressive compression can degrade downstream computer-vision performance and does not necessarily reduce the number of vision-language model (VLM) inference calls required for semantic video understanding. This paper evaluates BLUE, a fixed-camera surveillance compression approach that suppresses static-background redundancy while preserving foreground activity, for its effect on VLM-based event and anomaly understanding. We compare raw H.265 and BLUE-compressed H.265 video on two surveillance datasets: VIRAT, comprising 227 paired event samples from 106 clips, and CHAD, comprising 54 human-activity anomaly clips. For each pair, the same frame index is evaluated using a VLM captioning pipeline, and outputs are scored against annotation-derived ground truth using a blind judging protocol. The results show no measurable degradation in semantic inference quality. On VIRAT, the mean VLM score remains effectively unchanged between raw H.265 and BLUE, with a mean difference of approximately -0.01 on a 0-10 scale. On CHAD, raw H.265 and BLUE obtain near-equivalent mean scores of 4.31 and 4.26, respectively. Compression saving is also uncorrelated with VLM score change on VIRAT (r = 0.004), indicating that higher BLUE compression does not predict semantic quality loss. Beyond storage reduction, BLUE increases the share of skip-heavy P-frames on CHAD from 1.4% to 53.2%, enabling an estimated 53% reduction in VLM calls through packet-size-based frame skipping. These findings suggest that BLUE functions as a machine-centric compression layer for surveillance video, reducing bandwidth and inference cost while preserving VLM semantic performance.

补充信息

↑