发表机构
Oxygen AIGC Group; Joy Future Academy(氧气AIGC集团; 京东探索研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究提出时尚原生的Oxygen-TryOn基础模型,用于任意物品虚拟试穿。通过专用数据引擎和特定训练构建,重新定义试穿任务,设计三阶段训练方法,在单物品和多物品试穿中取得先进成果,超越多个系统。
AI 中文摘要
我们提出了Oxygen-TryOn,一个用于任意物品虚拟试穿的统一基础模型。它并非 repurposing 通用图像编辑器,而是时尚原生的,通过专用数据引擎和特定于试穿的训练构建。给定一个或多个参考物品及单个目标主体图像,能合成主体穿着几乎任何时尚类别的物品的逼真图像。先前系统有局限,而Oxygen-TryOn支持多样物品和场景,将试穿重新定义为多参考、理解驱动的生成任务。构建了数据引擎并设计了三阶段训练方法,在公共基准测试和内部基准测试中取得了最先进的单物品试穿一致性和真实感以及多物品试穿领先地位。
英文摘要
We present Oxygen-TryOn, a unified foundation model for any-item virtual try-on. Rather than repurposing a general-purpose image editor, Oxygen-TryOn is fashion-native, built for try-on through a dedicated data engine and try-on-specific training. Given one or more reference items (clean product shots or in-the-wild worn-on photos) and a single target subject image, it synthesizes a photorealistic image of the subject wearing the items across virtually any fashion category. Prior systems handle a single garment category in a studio setting, and recent multi-reference methods remain garment-centric; in contrast, Oxygen-TryOn supports diverse items and scenarios, including full- and half-body views, a variable number of references, and free multi-item composition, while faithfully preserving both subject identity and item appearance. Instead of mask-based inpainting, we reformulate try-on as a multi-reference, understanding-driven generation task. We build a data engine that collects, manufactures, annotates, and filters high-quality try-on data at scale, and design a three-stage recipe of continued pre-training (CPT), supervised fine-tuning (SFT), and reinforcement learning (RL). The RL stage uses a hybrid reward combining an in-house try-on reward model with a proprietary, rubric-guided general-purpose model, jointly supervising fine-grained consistency and instruction-level quality. It also follows general editing instructions (e.g., pose changes) in the same pass. Across public benchmarks and our in-house Oxygen-TryOn Bench, it achieves state-of-the-art consistency and realism on single-item try-on and leads on multi-item try-on, matching or surpassing both leading proprietary systems (Nano Banana Pro, GPT-Image-2, Seedream5 Lite) and open-source models (FLUX.2).