arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

GDB-Reward:从评估指标到平面设计的训练奖励

GDB-Reward: From Evaluation Metrics to Training Rewards for Graphic Design

Adrienne Deganutti, Purvanshi Mehta, Simon Hadfield, Andrew Gilbert

arXiv 2609.02813首次发表:更新:

发表机构

Lica World; University of Surrey(利卡世界; 萨里大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出GDB-Reward框架,将异构平面设计评估指标转化为强化学习奖励,在冻结图像生成器的情况下,提升了平面设计任务中对设计规范的遵守度,为不可微监督领域的优化提供了新方案。

AI 中文摘要

文本到图像模型在自然图像合成方面表现出色,但在平面设计领域却存在不足,该领域的成功取决于对排版、布局、颜色和视觉传达的精确约束。虽然提示优化是替代成本高昂的扩散模型微调的有吸引力的方案,但针对冻结图像生成器学习提示需要信息丰富的奖励函数,尽管生成过程完全不可微。强化学习不需要可微目标,仅需能够对候选输出进行排名的标量奖励。这就提出了一个简单的问题:设计评估指标本身能否成为强化学习的奖励?我们的核心贡献是GDB-Reward,这是一个将异构平面设计评估指标系统地转换为统一强化学习奖励的框架。实验表明,GDB-Reward提供了有效的优化目标,在保持图像生成器完全冻结的同时,大幅提高了对设计规范的遵守程度,包括感知质量、渲染保真度和空间准确性。更广泛地说,我们的结果表明,异构、不可微的评估指标可以超越被动基准测试,在无法获得可微监督的领域成为强化学习的有效优化目标。

英文摘要

Text-to-image models excel at natural image synthesis but struggle with graphic design, where success depends on satisfying precise constraints on typography, layout, color, and visual communication. While prompt optimization offers an attractive alternative to expensive diffusion model fine-tuning, learning prompts for frozen image generators requires informative reward functions despite the entirely non-differentiable generation process. Reinforcement learning does not require differentiable objectives; it requires only scalar rewards capable of ranking candidate outputs. This raises a simple question: can design evaluation metrics themselves become reinforcement learning rewards? Our central contribution is GDB-Reward, a framework that systematically transforms heterogeneous graphic design evaluation metrics into a unified reinforcement learning reward. Experiments demonstrate that GDB-Reward provides an effective optimization objective, substantially improving adherence to the design specification in perceptual quality, rendering fidelity, and spatial accuracy while keeping the image generator entirely frozen. More broadly, our results demonstrate that heterogeneous, non-differentiable evaluation metrics can move beyond passive benchmarking to become effective optimization objectives for reinforcement learning in domains where differentiable supervision is unavailable.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑