arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CAPE-T2V:面向文本到视频生成中双侧条件对齐的以字幕生成器为锚点的提示增强方法

CAPE-T2V: Captioner-Anchored Prompt Enhancement toward Two-Sided Conditioning Alignment in Text-to-Video Generation

Yizhuo Jia, Jingyun Hua, Yuanxing Zhang

arXiv 2608.03046首次发表:更新:

AI 中文总结

该研究针对文本到视频生成中存在的PE-字幕 gap,提出CAPE-T2V框架,通过两步微调PE与DiT,在多个基准数据集上实现了更高的生成综合得分,有效缩小了训练与推理间的条件不匹配。

AI 中文摘要

文本到视频(T2V)扩散Transformer(DiTs)是基于详细的视频字幕进行训练的,而推理阶段通常依赖于由提示增强器(PE)改写的用户提示。现有工作通过优化PE、DiT或同时优化两者来提升生成效果,部分方法也尝试通过共享模式缩小训练-推理之间的不匹配。然而,即使在共享模式下,推理阶段的PE输出与DiT训练所用字幕仍可能在细节选择、信息组织、描述粒度及措辞上存在差异,我们将这种剩余不匹配称为PE-字幕 gap。本文提出CAPE-T2V,这是一种面向T2V生成中双侧条件对齐的两步式以字幕生成器为锚点的提示增强框架。首先,CAPE-T2V构建三类PE训练示例:将字幕生成器生成的目标与简洁源字幕、详细源字幕或从这些目标衍生的伪用户提示配对,随后微调PE以将每个输入映射到其配对的目标。其次,CAPE-T2V使用由锚定PE改写的视频衍生字幕对DiT进行微调;在推理阶段,同一PE会改写用户提示。相对于使用相同字幕模式的基线方法,CAPE-T2V在Wan2.2和LTX-2.3上的StoryEval、VBench-2.0和T2V-CompBench上均实现了更高的综合得分。此外,CAPE-T2V的PE-字幕 gap小于基线方法:在固定嵌入空间中通过平方最大均值差异测量,其用于DiT微调的字幕在分布上更接近推理阶段的PE输出。总体而言,这些结果表明CAPE-T2V是缓解PE-字幕 gap的有效方法。该项目可在指定网址获取。

英文摘要

Text-to-video (T2V) diffusion transformers (DiTs) are trained with detailed video captions, whereas inference often relies on user prompts rewritten by a prompt enhancer (PE). Prior work has improved generation by optimizing the PE, the DiT, or both; some methods have also sought to narrow the training-inference mismatch through shared schemas. Yet even within a shared schema, inference-time PE outputs and DiT training captions may still differ in detail selection, information organization, descriptive granularity, and phrasing. We refer to this residual mismatch as the PE-Caption gap and introduce CAPE-T2V, a two-step Captioner-Anchored Prompt Enhancement framework toward two-sided conditioning alignment in T2V generation. First, CAPE-T2V constructs three types of PE training examples, pairing captioner-generated targets with concise source captions, detailed source captions, or pseudo user prompts derived from those targets. It then fine-tunes the PE to map each input to its paired target. Second, CAPE-T2V fine-tunes the DiT on video-derived captions rewritten by the Anchored PE; the same PE rewrites user prompts at inference. Relative to a baseline using the same caption schema, CAPE-T2V achieves higher aggregate scores on StoryEval, VBench-2.0, and T2V-CompBench across Wan2.2 and LTX-2.3. Further, CAPE-T2V exhibits a smaller PE-Caption gap than the baseline: its DiT fine-tuning captions are closer in distribution to inference-time PE outputs, as measured by squared maximum mean discrepancy in a fixed embedding space. Overall, these results support CAPE-T2V as an effective approach to mitigating the PE-Caption gap. The project is available at https://github.com/yizzz927/CAPE-T2V.

CommentsIncludes appendix; 11 figures. Project page: https://github.com/yizzz927/CAPE-T2V

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑