现代化压缩域视频字幕生成器:SigLIP2与GPT-2替换的受控研究
Modernising the Compressed-Domain Video Captioner: A Controlled Study of SigLIP2 and GPT-2 Substitutions
浏览论文内容
中文总结 AI 辅助
本研究在压缩域视频字幕生成中受控替换编码器与解码器,发现SigLIP2提升全部指标(+4.5 CIDEr),而GPT-2因过拟合削弱增益,并报告了各配置的推理延迟。
中文摘要 AI 辅助
压缩域视频字幕生成通过直接操作I帧、运动向量和残差来避免完整的视频解码,以少量精度换取推理速度的大幅提升。CoCap使用CLIP视觉编码器和浅层BERT风格多模态解码器建立了这一流程。这两个组件都早于明显更强大的替代方案。我们提出了一个狭窄且受控的问题:CoCap的精度在多大程度上受限于这两个组件,其中哪一个才是约束瓶颈?我们将CLIP I帧编码器替换为SigLIP2,将BERT风格解码器替换为GPT-2,并在相同的数据、采样预算和优化计划下评估三种配置(原始配对、仅编码器替换、两者同时替换)。所有比较均基于我们自己对CoCap的复现而非其已发表的数据,因为我们在VATEX的4,999个片段子集上以降低的采样预算进行训练;因此绝对值与原始工作不可比。我们的复现与已发表结果紧密吻合:CIDEr和METEOR略高于原文(54.9对52.7;23.4对23.2),BLEU-4和ROUGE-L略低于原文(29.7对31.4;48.9对49.4)。我们将差异归因于我们的评估子集而非任何方向的改进。我们发现这两个替换方向相反:仅SigLIP2就提升了所有指标(+4.5 CIDEr),而在其之上添加GPT-2则削弱了这一增益,因为预训练解码器在两个周期内过拟合了4,999个片段。我们还报告了每种配置的推理延迟,因为速度是推动压缩域字幕生成的初衷,以延迟成本换取的精度提升应如实报告。
英文摘要
Compressed-domain video captioning avoids full video decoding by operating directly on I-frames, motion vectors and residuals, trading a small amount of accuracy for a large gain in inference speed. CoCap established this pipeline using a CLIP vision encoder and a shallow BERT-style multimodal decoder. Both components predate substantially stronger alternatives. We ask a narrow, controlled question: how much of CoCap's accuracy is limited by these two components, and which of the two is the binding constraint? We replace the CLIP I-frame encoder with SigLIP2 and the BERT-style decoder with GPT-2, and evaluate three configurations (the original pairing, the encoder substitution alone, and both substitutions together) under identical data, sampling budget and optimisation schedule. All comparisons are made against our own reproduction of CoCap rather than its published numbers, because we train on a 4,999-clip subset of VATEX at a reduced sampling budget; absolute values are therefore not comparable with the original work. Our reproduction tracks the published result closely: CIDEr and METEOR run slightly above it (54.9 against 52.7; 23.4 against 23.2), BLEU-4 and ROUGE-L slightly below (29.7 against 31.4; 48.9 against 49.4). We attribute the differences to our evaluation subset rather than to any improvement in either direction. We find the two substitutions pull in opposite directions: SigLIP2 alone improves every metric (+4.5 CIDEr), while adding GPT-2 on top erodes that gain, because a pretrained decoder overfits 4,999 clips within two epochs. We additionally report inference latency for each configuration, since speed is the property that motivates compressed-domain captioning in the first place, and an accuracy gain purchased at a latency cost should be reported as such.