发表机构
Peking University; Alibaba Group(北京大学; 阿里巴巴集团)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对视频扩散模型测试时缩放的“生成-丢弃”范式缺陷,提出无需训练的 GEARS 框架,通过阶段感知调度器与候选回收器提升轻量模型性能,在 VBench 上使 13 亿参数模型得分接近 140 亿参数模型。
AI 中文摘要
近期视频扩散模型已实现出色的生成质量,但高保真结果仍在很大程度上依赖闭源系统或昂贵的大规模基础设施。测试时缩放(Test-Time Scaling, TTS)提供了一种无需训练的方法,通过投入额外的推理计算来提升轻量生成器的性能,不过现有方法大多仍局限于噪声搜索范式:它们采样、选择或扰动去噪轨迹,并在昂贵的生成后丢弃低分候选。这种“生成-丢弃”过程不仅浪费计算资源,还浪费了可恢复样本中已编码的部分运动、布局或外观结构。我们提出 GEARS(Guided Editing for Adaptive Recycling Search,自适应回收搜索的引导编辑),这是一种无需训练的框架,通过“生成-评估-编辑”循环将此类候选转化为可编辑先验,从而将诊断引导候选回收引入视频 TTS。GEARS 包含两个协同组件:阶段感知调度器(Stage-Aware Scheduler)确定要修复的内容、修复时机,以及应保留、回收或丢弃哪些候选;候选回收器(Candidate Recycler)从关键帧和多维奖励反馈中诊断可恢复的失败,推导候选特定的修复提示,并通过流形感知潜在 SDEdit 修复对应候选。修复后的候选被回收到搜索池中,形成超出标准噪声扰动的优化路径,同时保留有用结构。在匹配的 NFE 预算下,GEARS 在 VBench 上始终优于现有视频 TTS 方法,使 13 亿参数模型的总得分可与 140 亿参数模型相当, ablation 实验验证了自适应调度、诊断条件编辑和流形感知重去噪的必要性。代码可在 GitHub 上获取。
英文摘要
Recent video diffusion models have achieved remarkable generation quality, but high-fidelity results still largely depend on closed-source systems or costly large-scale infrastructure. Test-time scaling (TTS) offers a training-free way to improve lightweight generators by spending additional inference compute, yet existing methods mostly remain within a noise-search paradigm: they sample, select, or perturb denoising trajectories and discard low-scoring candidates after expensive generation. This generate-and-discard process wastes not only computation but also the partial motion, layout, or appearance structure already encoded in recoverable samples. We present \textbf{GEARS} (\textbf{G}uided \textbf{E}diting for \textbf{A}daptive \textbf{R}ecycling \textbf{S}earch), a training-free framework that introduces {diagnosis-guided candidate recycling} into video TTS by turning such candidates into editable priors through a generation-evaluation-editing loop. GEARS consists of two collaborative components. The \textbf{Stage-Aware Scheduler} determines what to repair, when to repair it, and which candidates should be preserved, recycled, or discarded. The \textbf{Candidate Recycler} diagnoses recoverable failures from keyframes and multi-dimensional reward feedback, derives candidate-specific repair prompts, and repairs the corresponding candidates through manifold-aware latent SDEdit. The repaired candidates are recycled into the search pool, creating refinement paths beyond standard noise perturbation while preserving useful structure. Under matched NFE budgets, GEARS consistently outperforms existing video TTS methods on VBench, bringing a 1.3B model to a total score comparable to a 14B counterpart, and ablations verify the necessity of adaptive scheduling, diagnosis-conditioned editing, and manifold-aware re-denoising. Code is available on GitHub.
CommentsAccepted by ACM TOG (SIGGRAPH Asia 2026)
DOI:10.1145/3842526