发表机构
School of Software, Northwestern Polytechnical University; School of Information Science and Technology, Beijing University of Technology; School of Computer Science and Technology, Xi’an Jiaotong University; School of Computer Science, The University of Sydney(西北工业大学软件学院; 北京工业大学信息科学与技术学院; 西安交通大学计算机科学与技术学院; 悉尼大学计算机科学学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出WSV零样本视频字幕生成框架,通过文本到视频生成模型、抛光器、提示器弥合跨模态差距,在三个公开数据集上取得了B@4 52、CIDEr 95.7的成绩。
AI 中文摘要
仅文本训练是零样本视频字幕生成领域的主流范式,该范式中模型训练阶段无法获取视频分布,导致训练阶段(仅文本)与推理阶段(仅视频)之间存在跨模态差距。现有研究尝试通过简单线性变换弥合该差距,但文本与视频间的固有差距使得跨模态表征空间对齐不足,生成的句子不准确。为解决此问题,本文提出一种新型零样本视频字幕生成框架WSV,包含两个训练阶段:首先通过预训练的文本到视频生成模型生成对应的合成视频隐表征;为增强隐表征的保真度,本文提出一种抛光器(polisher),用于弥合真实与合成视频分布间的差距;随后设计一个提示器(prompter),在第二训练阶段将GPT-2与抛光后的隐表征关联以生成字幕。推理阶段,输入视频经预训练的3D因果变分自编码器(3D Causal VAE)编码后直接输入提示器,提示器引导GPT-2生成最终字幕。在MSVD、MSR-VTT、VATEX数据集上的实验结果表明,本文方法在B@4指标上得分为52,在CIDEr指标上得分为95.7。
英文摘要
Text-only training is a popular paradigm in zero-shot video captioning, where the video distribution is not available to the model during training, leading to a cross-modal gap between the training (text-only) and the inference (video-only). Previous works attempt to bridge the gap through simple linear transformations. However, the inherent gap between text and video makes cross-modal representation space alignment insufficient, resulting in inaccurate sentences. To address this issue, we propose a novel zero-shot video captioning framework (WSV) consisting of two training stages, which first generates corresponding synthetic video latent representations via a pretrained text-to-video generation model. To strengthen the fidelity of the latent representations, we propose a polisher capable of bridging the gap between real and synthetic video distributions. Subsequently, we design a prompter that conditions GPT-2 on the polished latent representations to generate the captions in the second training stage. During inference, an input video is encoded by a pretrained 3D Causal VAE and then fed directly into the prompter, which in turn guides GPT-2 to produce the final caption. Experimental results conducted on MSVD, MSR-VTT, and VATEX datasets demonstrate that our proposed method achieves scores of 52 and 95.7 on the B@4 and CIDEr metrics, respectively.