arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

语义对齐的梯度驱动上下文保持图像编辑

Semantically Aligned Gradient-Driven Context-Preserving Image Editing

Chiranjeev Chiranjeev, Muskan Dosi, Mayank Vatsa, Richa Singh

arXiv 2609.12691首次发表:更新:

发表机构

Indian Institute of Technology Jodhpur(印度理工学院焦特布尔分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出IABEdit,通过可微语义验证训练实现上下文保持的图像编辑,在多个基准上提升结构保真度和指令遵循,尤其改善遮挡下的身份保持。

AI 中文摘要

指令引导的图像编辑存在训练时的盲区。生成式编辑器从未被要求语义验证其输出是否真正满足指令。监督在重建和输入文本级条件处停止。这导致编辑不完整、空间溢出和定位不佳。我们提出IABEdit,一个模型无关的框架,将可微的语义验证嵌入训练中。一个冻结的视觉-语言模型从真实编辑中提取空间感知描述符。然后,一个可训练的校准器从生成输出中重现这些描述符。两者之间的残差成为梯度,教导生成器编辑什么以及在哪里编辑,且推理时无VLM成本。IABEdit兼容多种骨干网络,包括U-Net(Stable Diffusion)和MMDiT(FLUX),且不改变其推理流程。在MagicBrush上,它比最佳扩散基线提高了+3.49 DINO-I的结构保真度,比最佳整体基线提高了+1.26,同时在指令对齐方面保持竞争力。它还在RealEdit和EMU Edit基准上基于嵌入指标实现了最先进的指令遵循性能。最重要的是,在D-LORD监控基准上,它在重度遮挡下超越了专有的Gemini智能体,DINO-P提高了+5.13,而身份保持最为困难。这表明梯度对齐的VLM蒸馏在类似真实世界的监控和遮挡条件下依然有效。人类和GPT-4o评估确认了感知上精确、定位良好的编辑。

英文摘要

Instruction-guided image editing has a training-time blind spot. Generative editors are never required to semantically verify whether their outputs actually satisfy the instruction. Supervision stops at reconstruction and input textual-level conditioning. This produces incomplete edits, spatial spillover, and poor localization. We present IABEdit, a model-agnostic framework that embeds differentiable semantic verification into training. A frozen vision-language model extracts spatially-aware descriptors from the ground-truth edit. A trainable aligner then reproduces them from the generated output. The residual between the two becomes a gradient that teaches the generator both what to edit and where, with no inference-time VLM cost. IABEdit is compatible with diverse backbones, including U-Net (Stable Diffusion) and MMDiT (FLUX), without altering their inference pipelines. On MagicBrush, it improves structural fidelity by +3.49 DINO-I over the best diffusion baseline and +1.26 over the best overall baseline, while remaining competitive on instruction alignment. It also achieves state-of-the-art instruction adherence performance on RealEdit and EMU Edit benchmarks based on embedding-based metrics. Most consequentially, on the D-LORD surveillance benchmark, it surpasses the proprietary Gemini agent by +5.13 DINO-P under heavy occlusion, where preserving identity is hardest. This shows that gradient-aligned VLM distillation holds up under real-world-like surveillance and occlusion conditions. Human and GPT-4o evaluations confirm perceptually precise, well-localized edits.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑