DDA-Thinker: 解耦双原子强化学习用于推理驱动的图像编辑
DDA-Thinker: Decoupled Dual-Atomic Reinforcement Learning for Reasoning-Driven Image Editing
浏览论文内容
中文总结 AI 辅助
本文提出DDA-Thinker框架,通过解耦的规划模块与固定生成模型,提升图像编辑中的推理能力,采用双原子强化学习分解反馈,结合可验证检查清单提升训练效果。
中文摘要 AI 辅助
最近的图像编辑模型在视觉保真度上表现优异,但往往在需要复杂推理的任务上表现不佳。为了研究和增强图像编辑中的推理驱动规划,我们提出了DDA-Thinker,一个以Thinker为中心的框架,旨在对规划模块(Thinker)和固定生成模型(Editor)进行独立优化。这种解耦的Thinker中心范式促进了对规划模块的受控分析,并使其在固定Editor下的贡献更容易评估。为了有效引导此Thinker,我们引入了双原子强化学习框架。该框架将反馈分解为两个不同的原子奖励,通过可验证的检查清单实现:一个认知原子奖励直接评估Thinker的可执行计划质量,作为Thinker推理的行动结果;另一个视觉原子奖励评估最终图像质量。为了提高检查清单质量,我们的检查清单合成不仅基于源图像和用户指令,还基于理想编辑场景的理性参考描述。为了支持此训练,我们进一步开发了一个两阶段数据整理管道,首先合成多样化且以推理为导向的数据集,然后应用难度感知的细化来整理强化学习的有效训练课程。在推理驱动的图像编辑基准上进行的广泛实验,包括RISE-Bench和KRIS-Bench,证明了我们的方法显著提高了整体性能。我们的方法使社区模型能够实现与强专有模型竞争的结果,突显了在固定编辑器设置下Thinker中心优化的实用潜力。
英文摘要
Recent image editing models have achieved strong visual fidelity but often struggle with tasks requiring complex reasoning. To investigate and enhance the reasoning-grounded planning for image editing, we propose DDA-Thinker, a Thinker-centric framework designed for the independent optimization of a planning module (Thinker) over a fixed generative model (Editor). This decoupled Thinker-centric paradigm facilitates a controlled analysis of the planning module and makes its contribution under a fixed Editor easier to assess. To effectively guide this Thinker, we introduce a dual-atomic reinforcement learning framework. This framework decomposes feedback into two distinct atomic rewards implemented through verifiable checklists: a cognitive-atomic reward to directly assess the quality of the Thinker's executable plan, which serves as the actionable outcome of the Thinker's reasoning, and a visual-atomic reward to assess the final image quality. To improve checklist quality, our checklist synthesis is grounded not only in the source image and user instruction but also in a rational reference description of the ideal post-edit scene. To support this training, we further develop a two-stage data curation pipeline that first synthesizes a diverse and reasoning-focused dataset, then applies difficulty-aware refinement to curate an effective training curriculum for reinforcement learning. Extensive experiments on reasoning-driven image editing benchmarks, including RISE-Bench and KRIS-Bench, demonstrate that our approach substantially improves overall performance. Our method enables a community model to achieve results competitive with strong proprietary models, highlighting the practical potential of Thinker-centric optimization under a fixed-editor setting.
发表机构
- Alibaba Group(阿里巴巴集团)
- Zhejiang University(浙江大学)
机构由 AI 辅助整理,请以论文原文为准。