arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FASTER:奖励引导图像编辑中端点细化的快速伴随随机输运

FASTER: Fast Adjoint Stochastic Transport for Endpoint Refinement in Reward-Guided Image Editing

Yimiao Zhou, Zejia Zhong, Jingya Wang, Ye Shi

arXiv 2610.04538首次发表:更新:

发表机构

ShanghaiTech University(上海科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

FASTER通过为每个源和目标训练小型网络,在避免预训练主干反向传播的同时,实现奖励引导图像编辑的快速端点细化,在SD3和SD1.5上分别获得高达6.91倍和24.14倍的编辑加速。

AI 中文摘要

测试时的奖励引导图像编辑旨在提高指定奖励,同时保留源内容和视觉合理性。许多现有方法通过预训练生成过程优化候选,使得重复调整依赖于昂贵的大模型执行,在某些情况下还依赖于主干网络的反向传播。我们开发了一个理论框架,该框架联合考虑了奖励、源内容保留和预训练先验偏好,允许期望的输出分布与用于实现该分布的动力学分开指定。基于此框架,我们提出了FASTER,它为每个源和目标训练一个小型网络以执行廉价的编辑,而预训练模型和奖励模型为候选输出提供反馈。通过在多次小型网络更新中重用每个候选及其反馈,FASTER减少了重复采样和监督查询,而无需将预训练生成主干置于内部优化循环中。在SD3上,FASTER在所评估的方法中引领了所有四个目标指标和几个验证指标。与沿预训练生成轨迹优化控制的评估基线相比,FASTER在Stable Diffusion 3上实现了高达6.91倍的编辑时间加速,在Stable Diffusion 1.5上实现了高达24.14倍的加速。

英文摘要

Reward-guided image editing at test time seeks to improve a specified reward while preserving source content and visual plausibility. Many existing approaches optimize candidates through pretrained generation processes, making repeated adjustment depend on costly large-model execution and, in some cases, backbone backpropagation. We develop a theoretical framework that jointly accounts for reward, source preservation, and pretrained-prior preferences, allowing the desired output distribution to be specified separately from the dynamics used to realize it. Based on this framework, we introduce FASTER, which trains a small network for each source and objective to perform inexpensive editing, while pretrained and reward models provide feedback on candidate outputs. By reusing each candidate and its feedback across multiple small-network updates, FASTER reduces repeated sampling and supervision queries without placing the pretrained generative backbone inside the inner optimization loop. On SD3, FASTER leads all four target metrics and several validation metrics among the evaluated methods. Compared with the evaluated baseline that optimizes controls along pretrained generation trajectories, FASTER achieves editing-time speedups of up to \({6.91\times}\) on Stable Diffusion 3 and \({24.14\times}\) on Stable Diffusion 1.5.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑