发表机构
The University of Texas at Austin(德克萨斯大学奥斯汀分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究视频扩散测试时搜索中缓存对候选者排名的影响,提出CachedSearch方法,通过积极缓存探索候选者,仅全计算时重生成获胜者,在多模型多系列中有效提升效率,以低成本获高增益,且免训练、与验证器无关,可用于测试时扩展。
AI 中文摘要
测试时搜索能让小型视频扩散模型媲美大型模型,但成本高出2至10倍。所有候选者都被完全去噪,尽管大多数会被丢弃。免训练缓存使每次展开速度加快2至3倍且质量接近无损。只有当有损缓存保留验证器排名时,合成才是安全的。我们首次研究了缓存是否会在视频测试时搜索中破坏候选者排名。在Wan2.1 - T2V - 1.3B上使用自适应缓存包装器(每个候选者加速约2倍),ImageReward对种子匹配的缓存和完整展开进行评分。在VBench套件上,每个提示的中位数Spearman等级相关性为0.905,前1名一致性为72%。VBench - 2.0在更难的套件上复制了这一结果。以全计算重新计算缓存的获胜者保留了全搜索增益的90 - 94%。错误聚集在接近平局的候选者中,使破坏具有自我限制性。这一发现促成了CachedSearch。它通过积极缓存探索每个候选者,然后仅在全计算时重新生成获胜者。在N = 8时,它以63%的成本获得了最佳N中94.7%的增益。捕获率随宽度增加。在匹配预算下,它搜索宽度翻倍,增益增加38%。该结果在六个模型和四个系列(Wan、LTX、CogVideoX和浑元)中从1.3B到14B都成立。Wan2.1 - 14B与1.3B模型的保真度匹配。轨迹中值修剪将探索节省倍数提高到3.11倍,捕获率为88.6%。移植到其他模型系列只需重新校准单个参数,表明保真度与架构相关而非参数数量。CachedSearch免训练、与验证器无关且与搜索算法正交,是测试时扩展的插件乘数。
英文摘要
Test-time search lets small video diffusion models rival larger ones, but costs 2-10x more. All candidates are fully denoised, although most are discarded. Training-free caching makes each rollout 2-3x faster at near-lossless quality. Composition is safe only if lossy caching preserves verifier rankings. We present the first study of whether caching corrupts candidate ranking in video test-time search. On Wan2.1-T2V-1.3B with an adaptive caching wrapper (~2x per-candidate speedup), ImageReward scores seed-matched cached and full rollouts. Median per-prompt Spearman rank correlation is 0.905, with 72% top-1 agreement on the VBench suite. VBench-2.0 replicates this result on a harder suite. Recomputing the cached winner at full compute retains 90-94% of the full-search gain. Errors cluster among near-tied candidates, making corruption self-limiting. This finding leads to CachedSearch. It explores every candidate with aggressive caching, then re-generates only the winner at full compute. At N=8, it captures 94.7% of best-of-N's gain at 63% of the cost. Capture rises with width. At matched budget, it searches twice as wide for 38% more gain. The result holds from 1.3B-14B across six models and four families: Wan, LTX, CogVideoX, and Hunyuan. Wan2.1-14B matches the 1.3B model's fidelity. Mid-trajectory pruning multiplies the exploration saving to 3.11x at 88.6% capture. Ports to other model families require recalibrating a single parameter, showing that fidelity tracks architecture rather than parameter count. CachedSearch is training-free, verifier-agnostic, and orthogonal to the search algorithm, making it a plug-in multiplier for test-time scaling.