AI 中文总结
ChordVideo扩展低能量原理到视频时间,通过共享噪声等技术解决ChordEdit的视频编辑时间问题,在TGVE/DAVIS上多指标提升,步数远少于多步编辑器
AI 中文摘要
一步式文本到图像模型支持仅需1-2次网络函数评估(NFE)的无训练、无逆编辑,ChordEdit通过沿采样时间的低能量平滑稳定此类编辑。但将其独立应用于视频帧时,会产生时间闪烁和编辑强度漂移。我们提出ChordVideo,将相同的低能量原理扩展到视频时间,方法包括共享噪声、逐帧Chord场的运动对齐因果聚合,以及可选的时间平滑近端校正。我们推导了扭曲误差边界,可区分运动偏差与随机闪烁,并预测时间窗口增大时收益递减。在TGVE/DAVIS数据集上,采用两个一步式骨干网络,ChordVideo将扭曲误差降低78%、闪烁降低49%,CLIP帧一致性提升9-10个点,背景PSNR提高约1.5dB,同时保持每帧2次NFE。与7种多步编辑器相比,它每片段使用10-60倍更少的模型步数,实现了具有竞争力的时间一致性和源保留效果。
英文摘要
One-step text-to-image models enable training-free, inversion-free editing with only 1--2 network function evaluations (NFE), while ChordEdit stabilizes such edits through low-energy smoothing along sampling time. Applied independently to video frames, however, it produces temporal flicker and edit-strength drift. We introduce \textbf{ChordVideo}, which extends the same low-energy principle to video time through shared noise, motion-aligned causal aggregation of per-frame Chord fields, and an optional temporally smoothed proximal correction. We derive a warping-error bound that separates motion bias from stochastic flicker and predicts diminishing returns with larger temporal windows. On TGVE/DAVIS with two one-step backbones, ChordVideo reduces warping error by \textbf{78\%} and flicker by \textbf{49\%}, improves CLIP frame consistency by \textbf{9--10 points}, and increases background PSNR by about \textbf{1.5,dB}, while retaining \textbf{2 NFE/frame}. Compared with seven multi-step editors, it achieves competitive temporal consistency and source preservation using \textbf{10--60$\times$ fewer model steps per clip