面向电商图像生成的元素感知组学习
Element-Aware Group Learning for E-Commerce Image Generation
浏览论文内容
中文总结 AI 辅助
针对电商图像生成中GRPO仅在完整提示层面分配信用的缺陷,提出EAGLE-GRPO,通过分解奖励并推导闭式解实现元素级信用分配,提升提示质量与生成图像性能。
中文摘要 AI 辅助
近期图像生成与编辑领域的进展已使提示质量成为电商创意的关键瓶颈。视觉语言模型(VLMs)可从产品图像和元数据生成图像编辑提示,但要进一步提升其提示撰写能力,需利用生成图像的反馈进行后训练。组相对策略优化(GRPO)是这类结果级奖励优化的自然框架,不过它仅在完整提示层面分配信用,而图像质量往往取决于构图、背景、卖点呈现等特定设计元素。现有细粒度信用分配方法通常需要步骤级监督或学习型评判器。为解决该问题,我们提出EAGLE-GRPO(面向电商图像生成的元素感知组学习),它将以组为中心的奖励分解到预定义元素上。我们将元素级信用分配转化为核岭回归问题,并推导得到闭式解,无需额外rollout或独立的信用分配模型。这可得到可解释的元素级优势值,实现更精准的策略更新。实验表明,EAGLE-GRPO在更多训练步骤中维持性能提升后才趋于平稳,且生成的提示能产出比主流VLM提示撰写基线质量更高的电商图像。
英文摘要
Recent advances in image generation and editing have made prompt quality a key bottleneck for e-commerce creatives. Vision-language models (VLMs) can generate image-editing prompts from product images and metadata, but further improving their prompt-writing capabilities requires post-training with feedback from the generated images. Group Relative Policy Optimization (GRPO) is a natural framework for such outcome-level reward optimization. However, it assigns credit only at the full-prompt level, even though image quality often depends on specific design elements such as composition, background, and the presentation of selling points. Existing fine-grained credit assignment methods typically require step-level supervision or learned critics. To address this, we propose EAGLE-GRPO (Element-Aware Group Learning for E-Commerce Image Generation), which decomposes the group-centered reward over predefined elements. We cast element-level credit assignment as a kernel ridge regression problem and derive a closed-form solution, without additional rollouts or separate credit-assignment models. This yields interpretable per-element advantages and more precise policy updates. Experiments show that EAGLE-GRPO sustains performance gains over more training steps before plateauing and generates prompts that produce higher-quality e-commerce images than competitive VLM prompt-writing baselines.