发表机构
NXN Labs; KAIST(NXN实验室; 韩国科学技术院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究旨在解决虚拟试穿中用户对服装穿着方式控制不足的问题。提出通过VIP - SAM解决视觉实例提示分割,引入CtrlVTON可控框架,将试穿转为图像编辑问题并添加分割掩码控制布局。二者在各自任务达先进水平,CtrlVTON能更忠实地遵循用户布局且保证服装逼真度。
AI 中文摘要
虚拟试穿(VTO)在将服装逼真地转移到目标人物身上方面取得了重大进展。然而,大多数系统让用户几乎无法控制服装的穿着方式,包括尺寸(宽松或合身)、款式(如塞进或不塞进、敞开或闭合)以及在身体上的空间位置。我们通过两个互补的贡献来解决这一差距。首先,我们通过VIP - SAM定义并解决视觉实例提示分割:给定一件服装的平铺图像,在穿着该服装的人的照片中分割出特定实例。这是一个实例级任务,不同于通常研究的类别级分割。其次,我们引入CtrlVTON,一个可控的VTO框架,将试穿重新定义为图像编辑问题,并添加分割掩码作为对服装布局的像素级控制,包括款式、尺寸和在身体上的空间位置。VIP - SAM和CtrlVTON在各自任务上均取得了领先成果。特别是,CtrlVTON生成的图像比最强的专有编辑系统更忠实地遵循用户提供的布局,同时在服装逼真度上与之匹配。
英文摘要
Virtual try-on (VTO) has made significant progress in realistically transferring garments onto a target person. Yet most systems give the user little control over how a garment should be worn -- its size (loose or fitted), style (e.g., tucked in or untucked, open or closed), and spatial placement on the body. We address this gap with two complementary contributions. First, we define and solve Visual-Instance-Prompt Segmentation via VIP-SAM: given a flatlay image of a garment, segment that specific instance in a photograph of a person wearing it. This is an instance-level task, distinct from the typically studied category-level segmentation. Second, we introduce CtrlVTON, a controllable VTO framework that recasts try-on as an image editing problem and adds segmentation masks as pixel-level control over garment layout, including style, size, and spatial placement on the body. VIP-SAM and CtrlVTON each achieve state-of-the-art results on their respective tasks. In particular, CtrlVTON generates images that follow user-provided layouts far more faithfully than the strongest proprietary editing systems while matching them on garment fidelity.
Comments13 + 17 pages, 20 figures