TryOnReward:学习用于虚拟试穿强化微调的中央凹一致性
TryOnReward: Learning Foveated Consistency for Reinforcement Fine-Tuning of Virtual Try-On
浏览论文内容
中文总结 AI 辅助
提出TryOnReward,一种基于视觉语言模型和中央凹校准的细粒度奖励模型,用于虚拟试穿强化微调,通过联合优化偏好与维度评分,在多个基准上显著提升人类偏好对齐并缓解奖励黑客问题。
中文摘要 AI 辅助
虚拟试穿(VTON)旨在用参考服装为人物着装,生成符合人类偏好的视觉合理结果。将这一以偏好为导向的目标转化为可操作的目标,依赖于与人类品味对齐的评分函数。然而,经典的保真度指标与人类判断的相关性较弱,而通用视觉语言模型(VLM)无法提供试穿质量评估所需的判别性粒度,试穿质量评估关键在于忠实保留服装和人物细节。这一缺陷在强化微调(RFT)优化中被进一步放大,导致严重的奖励黑客问题。为此,我们提出了TryOnReward,一个为VTON量身定制的细粒度奖励模型。基于视觉语言骨干网络,它采用中央凹校准目标,将每个质量维度锚定在相关区域,以避免全局捷径学习。同时,TryOnReward通过边际感知监督联合优化成对偏好和逐维度质量分数,利用相对和绝对质量信号。为了模型训练和评估,我们构建了TryOnReward-100K,一个带有人工标注的逐维度评分数据集,以及TryOn-Bench和TryOnRewardBench两个覆盖多样真实场景的基准。大量实验证实,TryOnReward在人类偏好对齐方面显著优于通用评判器,并且当作为RFT奖励函数时,它在多个基线上持续产生符合人类偏好的试穿结果。
英文摘要
Virtual Try-On (VTON) aims to dress a person with the reference garment, producing visually reasonable results aligned with human preferences. Turning this preference-oriented goal into an actionable objective relies on a scoring function aligned with human taste. However, classic fidelity metrics exhibit weak correlation with human judgments, and generic VLMs fail to provide the discriminative granularity demanded by try-on quality evaluation, which hinges on faithfully preserving garment and person details. This shortcoming is further exacerbated in the reinforcement fine-tuning (RFT) optimization and leads to severe reward hacking. To this end, we present TryOnReward, a fine-grained reward model tailored for VTON. Built on a vision-language backbone, it adopts a foveation calibration objective that grounds each quality dimension in the relevant region to avoid global shortcut learning. Meanwhile, TryOnReward jointly optimizes pairwise preferences and per-dimension quality scores via margin-aware supervision, leveraging both relative and absolute quality signals. For model training and evaluation, we build TryOnReward-100K, a human-annotated per-dimension rating dataset, alongside TryOn-Bench and TryOnRewardBench, two benchmarks covering diverse real scenarios. Extensive experiments confirm that TryOnReward significantly outperforms generic judges in human preference alignment, and when serving as the RFT reward function, it consistently yields human-preferred try-on results across multiple baselines.
发表机构
- University of Science and Technology of China(中国科学技术大学)
- South China University of Technology(华南理工大学)
机构由 AI 辅助整理,请以论文原文为准。