arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于统一序列组合模型的时尚穿搭生成

Autonomous Fashion Outfit Composition via Unified Aesthetic Foresight Model

Kaicheng Pang, Xingxing Zou, Ruohan Xu, Waikeung Wong

arXiv 2608.13888首次发表:更新:

发表机构

Laboratory for Artificial Intelligence in Design; Hong Kong Polytechnic University; The University of Queensland(人工智能设计实验室; 香港理工大学; 昆士兰大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文将时尚穿搭生成形式化为约束集成生成,提出统一序列组合模型与潜在扩展蒙特卡洛树搜索机制,在多数据集上实现了时尚穿搭生成任务的最优性能。

AI 中文摘要

从海量单品库中生成风格一致的时尚穿搭(即时尚穿搭生成)仍是一项非平凡的挑战,主要源于审美适配性的非单调与隐式特性,以及指数级庞大的组合搜索空间。本文将该任务形式化为约束集成生成(Constrained Ensemble Generation, CEG),并将其建模为有限时域确定性马尔可夫决策过程。为解决时尚领域的CEG问题,我们提出统一序列组合模型(Unified Sequential Composition Model, USCM),该模型联合建模集合级适配性与潜在组合意图。在USCM学习到的先验引导下,我们提出潜在扩展蒙特卡洛树搜索(Latent Expansion Monte Carlo Tree Search, LE-MCTS)机制以处理组合过程中的单品检索,平衡局部审美协同与全局结构平衡。在Polyvore Outfits数据集上的大量实验,以及在iFashion和PolyvoreU数据集上的零样本评估表明,我们的框架在独立人类偏好评估、自动化审美代理指标、以及约束时尚穿搭生成的结构有效性指标上均达到了当前最优性能。

英文摘要

Fashion Outfit Composition (FOC) requires sequentially assembling fashion items into a stylistically cohesive ensemble. Existing works struggle to model this step-by-step process effectively, primarily because they fail to jointly optimize the two critical capabilities required for FOC: intermediate outfit value evaluation and complementary item prediction. This structural disconnect, compounded by the severe sparsity of step-wise aesthetic signals, leaves them without a mechanism to autonomously determine when to stop the composition process. To address these challenges, we introduce the Unified Aesthetic Foresight Model (UAFM), which seamlessly unifies both capabilities into a dual-head architecture over a shared backbone. Crucially, we formulate FOC as a deterministic Markov Decision Process and optimize a value head via a post-decision state temporal difference (TD) objective to recursively backpropagate sparse terminal aesthetic rewards to intermediate states. This enables the model to accurately estimate the potential of a partial outfit evolving into a compatible ensemble. Consequently, UAFM can compute step-wise marginal aesthetic gains for dynamic termination. Furthermore, this unified design strictly aligns aesthetic evaluation and complementary item prediction within a shared representation manifold, intrinsically regularizing the combinatorial search space. Extensive experiments on the Polyvore-Outfits dataset demonstrate that UAFM establishes a new state-of-the-art across both FOC and conventional fashion tasks. Extensive ablation studies further confirm that our post-decision state TD formulation provides the aesthetic value prediction necessary for autonomous dynamic termination, while validating the synergistic benefits of our unified dual-head architecture.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑