发表机构
Southeast University; Purple Mountain Laboratories; Singapore Management University(东南大学; 紫金山实验室; 新加坡管理大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
MoRe-Drag通过将像素扭曲作为运动证据注入潜在重组,实现高精度拖拽编辑,在DragBench基准上优于现有方法。
AI 中文摘要
现代图像编辑器在语义操作和视觉合成方面表现出色,但在精确空间控制方面仍存在局限,这推动了拖拽编辑的发展。然而,现有的拖拽方法往往难以在拖拽精度与自然、合理且符合意图的生成之间取得平衡。我们提出了MoRe-Drag,一种基于运动锚定的拖拽编辑方法。我们的关键洞察是将像素空间扭曲视为粗略的运动证据,并将此证据注入生成采样轨迹中。具体而言,MoRe-Drag在细化、修复和锚定区域上执行区域感知的潜在重组,并结合阶段自适应条件,该条件逐步从基于运动的结构形成过渡到语义细化。我们进一步通过适配基于MLLM的文本编码器进行拖拽感知的指令推断,支持无需指令的界面。在DragBench-SR和DragBench-DR上的实验表明,MoRe-Drag在强基座编辑器上显著提高了拖拽精度,并在最先进的拖拽方法中实现了优越的拖拽准确性,同时提供了强大的语义一致性和视觉上真实的结果。代码和数据集将公开发布。
英文摘要
Modern image editors excel at semantic manipulation and visual synthesis, yet remain limited in precise spatial control, motivating the development of drag-based editing. However, existing drag-based methods often struggle to balance drag accuracy with natural, plausible, and intent-aligned generation. We propose MoRe-Drag, a motion-grounded drag-based editing method. Our key insight is to treat pixel-space warping as coarse motion evidence, and to inject this evidence into the generative sampling trajectory. Specifically, MoRe-Drag performs region-aware latent recomposition over refinement, inpainting, and anchor regions, coupled with stage-adaptive conditioning that progressively shifts from motion-grounded structure formation to semantic refinement. We further support an instruction-free interface by adapting the MLLM-based text encoder for drag-aware instruction inference. Experiments on DragBench-SR and DragBench-DR show that MoRe-Drag substantially improves drag precision over strong base editors and achieves superior drag accuracy among SOTA drag-based methods, while delivering strong semantic consistency and visually realistic results. Code and dataset will be publicly released.