CommerceVibe:通过双反馈强化学习学习将电商创意设计为可执行视觉代码
CommerceVibe: Learning to Design E-Commerce Creatives as Executable Visual Code via Dual-Feedback Reinforcement Learning
浏览论文内容
中文总结 AI 辅助
该研究提出CommerceVibe,将电商创意转化为可执行视觉代码,通过双反馈强化学习优化生成效果,在基准测试中优于仅监督微调的变体及外部模型,实现可控可编辑的规模化电商创意生产。
中文摘要 AI 辅助
高质量电商创意对于展示产品和传达营销信息至关重要。近期的扩散模型能够实现创意的规模化生成,并产出视觉上引人注目的图像,但其扁平化的光栅输出常包含扭曲的文本和不一致的产品细节,部署前需优化。此外,缺乏明确结构的生成创意难以编辑和复用,而复杂的设计需求仍难以编码为可验证的训练信号。为应对这些挑战,我们提出CommerceVibe,将创意表示为可执行视觉代码,并将生成过程建模为条件HTML/CSS程序合成。输入产品图像、设计需求和产品信息后,它可生成可渲染、可编辑、可复用的创意。我们进一步引入双反馈强化学习:基于规则的反馈评估渲染后程序的文本可读性、产品可见性和布局有效性;来自视觉语言模型(VLM)的视觉反馈则从六个感知和商业维度评估渲染创意与输入规格的匹配度。这些互补的反馈信号共同提升了约束满足度和依赖感知的质量。我们在超过28,000个电商示例上对Qwen3.5-9B进行监督微调(SFT),随后开展双反馈强化学习。在1,300个案例的基准测试中,优化后的CommerceVibe模型加权得分为94.0/100,仅监督微调变体得分为87.3,且其性能优于强大的外部模型。五名电商设计专家的盲评进一步验证了这些改进。CommerceVibe支持可控、可编辑且规模化的电商创意生产。
英文摘要
High-quality e-commerce creatives are essential for presenting products and conveying marketing messages. Recent diffusion models enable scalable creative generation and produce visually compelling images, but their flattened raster outputs often contain distorted text and inconsistent product details, requiring refinement before deployment. Moreover, without explicit structure, the resulting creatives are difficult to edit and reuse, while complex design requirements remain challenging to encode as verifiable training signals. To address these challenges, we present CommerceVibe, which represents creatives as executable visual code and formulates generation as conditional HTML/CSS program synthesis. Given product images, design requirements, and product information, it produces renderable, editable, and reusable creatives. We further introduce dual-feedback reinforcement learning, in which rule-based feedback evaluates rendered programs for text readability, product visibility, and layout validity, while visual feedback from a vision-language model (VLM) assesses rendered creatives against input specifications across six perceptual and commercial dimensions. Together, these complementary feedback signals improve both constraint satisfaction and perception-dependent quality. We perform supervised fine-tuning (SFT) of Qwen3.5-9B on over 28,000 e-commerce examples, followed by dual-feedback reinforcement learning. On a 1,300-case benchmark, the optimized CommerceVibe model achieves a weighted score of 94.0/100, compared with 87.3 for the SFT-only variant, and outperforms strong external models. Blind evaluations by five e-commerce design experts further validate these improvements. CommerceVibe supports controllable, editable, and scalable e-commerce creative production.
发表机构
- Tongji University(同济大学)
- Alibaba Group(阿里巴巴集团)
机构由 AI 辅助整理,请以论文原文为准。