PredErase:基于预测潜在引导的无训练物体与效果移除方法
PredErase: Training-Free Object-and-Effect Removal with Predictive Latent Guidance
- The University of Hong Kong(香港大学)
- Sun Yat-sen University(中山大学)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
PredErase是基于冻结FLUX.2和I-JEPA的无训练方法,通过分离可重写像素位置与孔洞结构,改进了原生FLUX.2在物体移除任务上的表现,且不替代有监督配对数据擦除器。
中文摘要 AI 辅助
移除物体与填充其掩码并不相同,投射的阴影和接触阴影通常位于用户提供的实例掩码M_obj之外,因此仅编辑该掩码的冻结Fill模型会在附近表面留下物体的光度足迹。有监督移除模型通过配对的干净图像学习这种联合擦除,无训练编辑模型冻结预训练权重,但大多数仍将M_obj视为整个可编辑支持区域,并使用CLIP或DINO能量引导采样,而这些能量无法预测被遮挡的场景。我们提出PredErase,一种基于冻结FLUX.2和I-JEPA的无训练推理过程,该方法将Fill可重写像素的位置与应占据孔洞的结构分离;M_obj的接触带扩展M_flux暴露了支撑平面上的局部残差,预训练用于掩码令牌预测的I-JEPA在表示空间中提供了上下文条件的孔洞目标,稀疏投影梯度使解码的Fill补全在实例内部与该目标对齐,而打包支撑区域外的坐标保持锁定。在RemovalBench、RORD-Val和DEFACTO-Val的仅实例掩码设置下,PredErase改进了原生FLUX.2骨干模型;有监督移除模型在若干全图像外观指标上仍表现更强,本文支持的主张是对冻结Fill进行无训练的物体与效果编辑,而非替代配对数据擦除器。
英文摘要
Removing an object is not the same as filling its mask. Cast shadows and contact shading usually lie outside the user-provided instance mask M_obj, so a frozen Fill model that edits only that mask leaves the object's photometric footprint on nearby surfaces. Supervised removers learn this joint erasure from paired clean plates. Training-free editors freeze pretrained weights, yet most still treat M_obj as the entire editable support and steer sampling with CLIP or DINO energies that do not predict the occluded scene. We present PredErase, a training-free inference procedure on frozen FLUX.2 and I-JEPA. The method separates where Fill may rewrite pixels from what structure should occupy the hole. A contact-band expansion M_flux of M_obj exposes local residuals on the supporting plane. I-JEPA, pretrained for masked token prediction, supplies a context-conditioned hole target in representation space; sparse projected gradients align decoded Fill completions with that target inside the instance, while coordinates outside the packed support stay locked. Under instance-only masks on RemovalBench, RORD-Val, and DEFACTO-Val, PredErase improves the native FLUX.2 backbone. Supervised removers remain stronger on several full-image appearance metrics; the supported claim is training-free object-and-effect editing of frozen Fill, not replacement of paired-data erasers.