对象感知的背景受控编辑:基于加权速度引导
Object-Aware Background-Controlled Editing via Weighted Velocity Guidance
- Carnegie Mellon University(卡内基梅隆大学)
- Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所)
- University of Chinese Academy of Sciences(中国科学院大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出免训练的对象感知速度控制(OAVC)框架,通过背景锚定参考和对象局部安全注入抑制背景漂移,在图像和视频编辑基准上提升背景保持、结构保真与边界稳定性。
AI中文摘要:
免训练图像编辑通过在推理时修改提示条件去噪速度,引导扩散或流匹配生成模型。现有的基于速度的编辑方法通常在整个潜在空间上全局应用提示引起的残差,并依赖模型隐式地定位语义变化。对于以对象为中心的编辑,这些残差在目标对象之外很少为零,因此小的非目标分量在多步积分过程中会累积,导致背景漂移和对象边界不稳定。我们提出对象感知速度控制(OAVC),一个免训练框架,将对象级控制引入速度积分过程。OAVC将语义残差允许作用的位置与它们注入动力学的方式解耦。它在源提示下构建背景锚定参考界面,然后在目标提示下执行对象局部安全语义注入。约束注入算子抑制引起漂移的速度分量,而时间自适应空间加权稳定对象边界附近的过渡。OAVC不需要训练或修改预训练模型参数。在图像和视频矫正流骨干的对象中心图像和视频基准上的实验表明,改进的背景保持、结构保真度、边界稳定性和时间一致性,同时保留有效的局部可编辑性。
英文摘要:
Training-free image editing steers diffusion or flow-matching generative models at inference time by modifying prompt-conditioned denoising velocities. Existing velocity-based editors often apply prompt-induced residuals globally over the latent space and rely on the model to localize semantic changes implicitly. For object-centric edits, these residuals are rarely zero outside the target object, so small non-target components can accumulate during multi-step integration, causing background drift and unstable object boundaries. We propose Object-Aware Velocity Control (OAVC), a training-free framework that introduces object-level control into the velocity-integration process. OAVC decouples where semantic residuals are allowed to act from how they are injected into the dynamics. It constructs a background-anchored reference interface under the source prompt and then performs object-localized safe semantic injection under the target prompt. A constrained injection operator suppresses drift-inducing velocity components, while time-adaptive spatial weighting stabilizes the transition near object boundaries. OAVC requires no training or modification of pretrained model parameters. Experiments on object-centric image and video benchmarks with image and video rectified-flow backbones show improved background preservation, structural fidelity, boundary stability, and temporal consistency while retaining effective localized editability.