arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RAGDiffusion++:面向服装生成的从宏观检索到微观保真度对齐

RAGDiffusion++: From Macro-Retrieval to Micro-Fidelity Alignment for Garment Generation

Yuhan Li, Xianfeng Tan, Fangao Zeng, Wenxiang Shang, Pipei Huang, Hao Zhou, Zhiyu Jin, Wenjun Zhang, Bingbing Ni

arXiv 2608.29280首次发表:更新:

发表机构

Shanghai Jiao Tong University; Alibaba Group(上海交通大学; 阿里巴巴集团)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对服装生成的高频轨迹坍塌问题,提出RAGDiffusion++模型,通过AR-GRPO策略结合STGarment-Plus数据集、FLUX架构与Garment-RM奖励模型,实现服装图像的宏观结构准确与微观纹理保真。

AI 中文摘要

标准服装资产生成任务是从各类真实场景中重建正面平铺的服装图像,具有巨大商业价值,但同时需要宏观拓扑结构准确性与微观物理保真度。尽管我们此前的工作RAGDiffusion通过检索增强的宏观约束有效消除了大规模结构幻觉,但达到工业级的微观纹理真实感仍是未解决的瓶颈。我们将此局限正式定义为高频轨迹坍塌:监督微调(SFT)收敛到训练分布的条件均值,该均值由平滑的低频纹理主导,导致高频图案(如织物纹理、复杂标志)几乎无法采样。在训练后直接应用强化学习(RL)会进一步触发伪影作弊,即模型通过生成欺骗性棋盘噪声来利用通用奖励模型中的语义偏差。我们的核心见解是,RL可在准确的奖励引导下从根本上重塑流模型的采样分布,提升高保真轨迹的概率,同时对抗性正则化可防止模型利用奖励盲点。实现这一原理需要三个前提条件:(i)固有能力,通过包含27725对高复杂度服装的数据集STGarment-Plus和双图像流FLUX架构升级确立;(ii)感知奖励,由在50万张图像上通过细粒度对比学习训练的新型属性感知奖励模型Garment-RM提供,达到84.67%的人类偏好准确率;(iii)作弊预防,由我们的对抗性正则化GRPO(AR-GRPO)策略实施,该策略将动态判别器集成到RL采样轨迹中,以惩罚伪影并丰富真实高频细节。

英文摘要

Standard clothing asset generation---restoring forward-facing flat-lay garment images from diverse real-world contexts---holds immense commercial value yet demands both macroscopic topological accuracy and microscopic physical fidelity. Although our previous work RAGDiffusion effectively eradicated large-scale structural hallucinations via retrieval-augmented macro-constraints, achieving industrial-grade micro-texture realism remains an unsolved bottleneck. We formally identify this limitation as High-Frequency Trajectory Collapse: supervised fine-tuning (SFT) converges to the conditional mean of the training distribution, which is dominated by smooth, low-frequency textures, causing high-frequency patterns (e.g., fabric weaves, intricate logos) to become nearly un-sampleable. Naively applying Reinforcement Learning (RL) post-training further triggers Artifact Hacking, where models exploit semantic biases in generic reward models by generating deceptive checkerboard noise. Our key insight is that RL can fundamentally reshape the sampling distribution of flow models---elevating the probability of high-fidelity trajectories under accurate reward guidance---while adversarial regularization prevents exploitation of reward blind spots. Realizing this principle requires three prerequisites: (i)inherent capacity, established through a 27,725-pair high-complexity garment dataset (STGarment-Plus) and a Dual-Image-Stream FLUX architecture upgrade; (ii)perceptive reward, provided by a novel attribute-aware reward model (Garment-RM) trained on 500K images via fine-grained contrastive learning, achieving 84.67% human preference accuracy; and (iii)hacking prevention, enforced by our Adversarial-Regularized GRPO (AR-GRPO) strategy that integrates a dynamic discriminator into the RL sampling trajectory to penalize artifacts while enriching authentic high-frequency details.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑