AI 中文总结
针对比特流损坏的恶劣视觉理解问题,提出BLMSP框架,构建CHP数据集,通过VBBM提取比特流语义先验注入视觉模型,显著提升视频恢复、字幕及姿态估计性能。
AI 中文摘要
比特流损坏的恶劣视觉理解(BcHVU)旨在理解在现实世界多媒体通信中从严重损坏的比特流解码出的严重退化视频。BcHVU的不适定性对现有视觉模型构成重大挑战,因为即使是细微的比特流损坏也会导致不可逆的像素失真和显著的语义损失。为解决BcHVU中的这些挑战,我们提出比特流语言建模作为鲁棒语义先验(BLMSP),这是一种用于学习和注入比特流原生语义线索的框架。我们提出的BLMSP框架通过比特流语言建模学习提取比特流原生语义线索,并将其作为先验注入BcHVU任务的现成视觉模型中。具体而言,我们提出视频比特流字节模型(VBBM),该模型集成了字节级建模和跨编解码器语义蒸馏,使其能够从多种损坏比特流格式的字节序列中解释鲁棒语义。学习到的比特流语义被用作鲁棒先验,并融合到BcHVU模型骨干中,以提高视频恢复、图像字幕和人体姿态估计的质量。为训练BLMSP,我们构建了大规模多源损坏比特流恶劣视频配对(CHP)数据集,包含60.7万个损坏比特流片段和28.7万个配对恶劣视频片段。大量实验结果表明,学习到的比特流先验分别使视频恢复的PSNR平均提高2.51 dB、图像字幕的CIDEr平均提高0.20、人体姿态估计的PCK@0.2平均提高0.18。这些结果证明,损坏的比特流可作为鲁棒语义先验,用于解决BcHVU中的像素失真和语义损失问题。
英文摘要
Bitstream-corrupted Harsh Visual Understanding (BcHVU) aims to understand harshly degraded videos originally decoded from a severely corrupted bitstream in real-world multimedia communication. The ill-posed nature of BcHVU poses a major challenge for existing vision models, as even subtle bitstream corruption can lead to irreversible pixel distortion and significant semantic loss. To address these challenges in BcHVU, we propose Bitstream Language Modeling as Robust Semantic Priors (BLMSP), a framework for learning and injecting bitstream-native semantic cues. Our proposed BLMSP framework learns to extract bitstream-native semantic cues by bitstream language modeling, and leverages them as priors by injecting into off-the-shelf vision models of BcHVU tasks. Specifically, we present a Video Bitstream Byte Model (VBBM) that integrates byte-level modeling and cross-codec semantic distillation, enabling it to interpret robust semantics from byte sequences in multiple corrupted bitstream formats. The learned bitstream semantics are leveraged as robust priors and fused into BcHVU model backbones for improving the quality of video restoration, captioning, and human pose estimation. To train BLMSP, we construct a large-scale multi-source Corrupted-bitstream Harsh-video Paired (CHP) dataset containing 607k corrupted bitstream segments and 287k paired harsh video clips. Extensive experimental results show that the learned bitstream priors improve video restoration, captioning, and human pose estimation by 2.51 dB in PSNR, 0.20 in CIDEr, and 0.18 in PCK@0.2 on average, respectively. These results demonstrate that corrupted bitstream can serve as robust semantic priors in solving pixel distortion and semantic loss in BcHVU.
Comments9 pages, 5 figures, 4 tables