发表机构
Shenzhen University; Zhejiang University(深圳大学; 浙江大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出SewFusion统一自回归框架,结合下一令牌预测与流匹配分别生成离散拓扑和裁片级连续几何,并引入Panel Geometry VAE与Panel-Forcing,在SewFactory和GCD-MM上超越现有方法。
AI 中文摘要
从图像和文本生成缝纫纸样需要对由离散拓扑和连续几何组成的异构表示进行建模。现有方法主要遵循两种范式:基于扩散的方法通过将整个纸样转换为连续表示来实现整体几何生成,但削弱了离散拓扑建模;相比之下,自回归方法通过下一令牌预测保留离散拓扑,但将连续几何回归绑定到令牌级隐藏状态,且裁片级上下文有限。为弥合这一差距,我们提出SewFusion,一个统一的自回归框架,采用针对离散拓扑和裁片级连续几何的定制化生成机制,前者使用下一令牌预测,后者使用流匹配。为支持裁片级连续几何生成,我们引入裁片几何VAE,学习一个固定大小的潜空间以表示变长裁片几何,并配套裁片几何流用于潜空间生成。我们进一步提出Panel-Forcing以减少拓扑上下文中的训练-推理不匹配,并提高对拓扑预测错误的鲁棒性。在SewFactory和GCD-MM上的大量实验表明,SewFusion在各种设置下持续优于先前最先进的方法,在基于图像-文本的生成设置中实现了+6.36%的裁片准确率、+11.30%的缝线准确率和-1.90的顶点L2误差。
英文摘要
Generating sewing patterns from images and text requires modeling a heterogeneous representation composed of discrete topology and continuous geometry. Existing methods mainly follow two paradigms: diffusion-based methods enable holistic geometry generation by converting the entire pattern into a continuous representation, but weaken discrete topology modeling; in contrast, autoregressive methods preserve discrete topology through next-token prediction, but tie continuous geometry regression to token-level hidden states with limited panel-level context. To bridge this gap, we propose SewFusion, a unified autoregressive framework that adopts tailored generation mechanisms for discrete topology and panel-level continuous geometry, using next-token prediction for the former and flow matching for the latter. To support panel-level continuous geometry generation, we introduce a Panel Geometry VAE that learns a fixed-size latent space for variable-length panel geometry, together with Panel Geometry Flow for latent generation. We further propose Panel-Forcing to reduce the training--inference mismatch in topology context and improve robustness to topology prediction errors. Extensive experiments on SewFactory and GCD-MM demonstrate that SewFusion consistently outperforms previous state-of-the-art methods across various settings, achieving +6.36% Panel Accuracy, +11.30% Stitch Accuracy, and -1.90 Vertex L2 error in the image-text-based generation setting.