arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

NoisEasier:用于文本到视频生成的测试时噪声优化

NoisEasier: Test-Time Noise Optimization for Text-to-Video Generation

Yujiang Pu, Yu Kong

arXiv 2608.30194首次发表:更新:

发表机构

Michigan State University(密歇根州立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出NoisEasier框架,通过测试时噪声优化提升文本到视频生成的组合对齐效果,在多数据集实验中取得超10%平均增益,为可控文本到视频生成提供有效新范式。

AI 中文摘要

扩散模型近期在文本到视频(T2V)生成领域取得了进展,但仍难以实现细粒度的组合对齐,例如属性绑定、空间关系和对象交互。基于奖励的微调虽能改善对齐情况,但易受奖励黑客攻击,且对新提示分布的适应性较差。本研究提出了NoisEasier,这是一种测试时缩放框架,无需修改底层模型,即可通过可微分奖励引导的噪声优化来改进T2V生成。通过将高效的短步生成器与多目标奖励公式相结合,NoisEasier能在实际推理预算下实现稳定且实用的测试时优化。核心见解在于,与仅优化初始潜在变量相比,联合优化整个随机轨迹可加速奖励收敛并提升组合对齐效果,且额外计算和时间成本可忽略不计。在VBench和T2V-CompBench上的实验表明,其在多个主干模型上实现了一致的改进,在属性绑定、对象交互和数值计算等具有挑战性的维度上取得了超过10%的平均增益。总体而言,NoisEasier可作为基于奖励的微调的灵活替代方案和补充增强手段,确立了测试时缩放作为可控文本到视频生成的有效范式。

英文摘要

Diffusion models have recently advanced text-to-video (T2V) generation, yet they still struggle with fine-grained compositional alignment, such as attribute binding, spatial relations, and object interactions. While reward-based fine-tuning improves alignment, it is susceptible to reward hacking and adapts poorly to new prompt distributions. In this work, we propose NoisEasier, a test-time scaling framework that improves T2V generation through differentiable reward-guided noise optimization without modifying the underlying model. By combining efficient short-step generators with a multi-objective reward formulation, NoisEasier enables stable and practical test-time optimization under realistic inference budgets. Our key insight is that jointly optimizing the entire stochastic trajectory accelerates reward convergence and improves compositional alignment over optimizing only the initial latent, with negligible additional computational and time cost. Experiments on VBench and T2V-CompBench demonstrate consistent improvements across multiple backbones, achieving over 10% average gains on challenging dimensions such as attribute binding, object interaction, and numeracy. Overall, NoisEasier serves as both a flexible alternative and a complementary enhancement to reward-based fine-tuning, establishing test-time scaling as an effective paradigm for controllable text-to-video generation.

CommentsAccepted to ECCV 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑