发表机构
South China Normal University; KlingAI Research(华南师范大学; 可灵AI研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出DARS框架,解决基于指令的图像编辑系统仅用最终奖励训练效率低的问题,通过双层级信用分配实现更好性能,在推理密集型编辑上增益显著。
AI 中文摘要
基于指令的图像编辑采用规划器-渲染器流水线:视觉语言模型(VLM)先将指令转换为编辑计划,扩散模型再执行该计划。仅用最终图像奖励训练此类系统效率低下,因为糟糕的编辑无法表明额外优化应更侧重规划器还是渲染器,且即使是规划器主导的情况,也难以在自由形式的推理轨迹中定位。我们提出DARS,这是一种针对该两阶段设置的双层级信用分配强化学习框架。跨模块时,多计划多渲染的rollout估计计划间和计划内的奖励变异性,用于软模块路由,而rollout平均奖励为自适应课程提供难度估计。在规划器内部,四字段结构化推理输出支持前缀门控奖励和token级优势重加权,将结果级反馈转化为局部监督。在五个基准上的实验表明,DARS在相同骨干网络、数据、奖励模型和rollout预算下,优于联合强化学习基线,在推理密集型编辑上的增益最大。
英文摘要
Instruction-based image editing uses a planner-renderer pipeline: a vision-language model (VLM) first converts the instruction into an edit plan, and a diffusion model then executes that plan. Training such systems with only final-image rewards is inefficient because a poor edit does not reveal whether additional optimization should place more emphasis on the planner or the renderer, and even planner-dominant cases remain difficult to localize within a free-form reasoning trace. We present DARS, a reinforcement learning framework for dual-level credit assignment in this two-stage setting. Across modules, multi-plan multi-render rollouts estimate between-plan and within-plan reward variability for soft module routing, while rollout mean rewards provide hardness estimates for an adaptive curriculum. Within the planner, a four-field structured reasoning output enables a prefix-gated reward and token-level advantage reweighting, turning outcome-level feedback into localized supervision. Experiments on five benchmarks show that DARS outperforms a Joint~RL baseline with the same backbone, data, reward model, and rollout budget, with the largest gains on reasoning-intensive edits.