arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

观看合成视频:针对零样本视频字幕生成的视觉合成跨模态表征对齐

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning

Liangyu Fu, Junbo Wang, Yuke Li, Ya Jing, Xuecheng Wu, Zhiyong Wang

arXiv 2608.11013首次发表:更新:

发表机构

School of Software, Northwestern Polytechnical University; School of Information Science and Technology, Beijing University of Technology; School of Computer Science and Technology, Xi’an Jiaotong University; School of Computer Science, The University of Sydney(西北工业大学软件学院; 北京工业大学信息科学与技术学院; 西安交通大学计算机科学与技术学院; 悉尼大学计算机科学学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出WSV零样本视频字幕生成框架,通过文本到视频生成模型、抛光器、提示器弥合跨模态差距,在三个公开数据集上取得了B@4 52、CIDEr 95.7的成绩。

AI 中文摘要

仅文本训练是零样本视频字幕生成领域的主流范式,该范式中模型训练阶段无法获取视频分布,导致训练阶段(仅文本)与推理阶段(仅视频)之间存在跨模态差距。现有研究尝试通过简单线性变换弥合该差距,但文本与视频间的固有差距使得跨模态表征空间对齐不足,生成的句子不准确。为解决此问题,本文提出一种新型零样本视频字幕生成框架WSV,包含两个训练阶段:首先通过预训练的文本到视频生成模型生成对应的合成视频隐表征;为增强隐表征的保真度,本文提出一种抛光器(polisher),用于弥合真实与合成视频分布间的差距;随后设计一个提示器(prompter),在第二训练阶段将GPT-2与抛光后的隐表征关联以生成字幕。推理阶段,输入视频经预训练的3D因果变分自编码器(3D Causal VAE)编码后直接输入提示器,提示器引导GPT-2生成最终字幕。在MSVD、MSR-VTT、VATEX数据集上的实验结果表明,本文方法在B@4指标上得分为52,在CIDEr指标上得分为95.7。

英文摘要

Text-only training is a popular paradigm in zero-shot video captioning, where the video distribution is not available to the model during training, leading to a cross-modal gap between the training (text-only) and the inference (video-only). Previous works attempt to bridge the gap through simple linear transformations. However, the inherent gap between text and video makes cross-modal representation space alignment insufficient, resulting in inaccurate sentences. To address this issue, we propose a novel zero-shot video captioning framework (WSV) consisting of two training stages, which first generates corresponding synthetic video latent representations via a pretrained text-to-video generation model. To strengthen the fidelity of the latent representations, we propose a polisher capable of bridging the gap between real and synthetic video distributions. Subsequently, we design a prompter that conditions GPT-2 on the polished latent representations to generate the captions in the second training stage. During inference, an input video is encoded by a pretrained 3D Causal VAE and then fed directly into the prompter, which in turn guides GPT-2 to produce the final caption. Experimental results conducted on MSVD, MSR-VTT, and VATEX datasets demonstrate that our proposed method achieves scores of 52 and 95.7 on the B@4 and CIDEr metrics, respectively.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑