发表机构
Institut Teknologi Del(德尔理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对巴塔克Ulos纹样传统设计局限,提出融合微调Stable Diffusion XL与LLaMA的多模态生成框架,通过四种条件控制机制实现可控生成,实验验证了最佳组合方案并获得织工与公众的显著认可。
AI 中文摘要
传统巴塔克Ulos编织产业因传统手工设计方法的局限性,在生成多样化、创新纹样方面面临日益严峻的挑战。本研究提出一种多模态生成框架,将微调的潜在扩散模型(通过LoRA微调的Stable Diffusion XL v1.0)与多模态大语言模型(LLaMA 1.5-7B)相结合,以实现可控且文化保真的Ulos纹样生成。四种互补的条件控制机制:文本、图像、表征和语义图(通过ControlNet)共同引导生成过程,每种机制分别控制从语义意图到空间布局的不同方面。在三个场景(形状变换、颜色变化和高复杂度输入)中进行的五级消融研究表明,条件控制的有效性与组合机制的数量不成正比:文本+图像+语义图组合取得了最佳FID(270)和CLIP分数(0.65-0.70),但SSIM最弱(0.65);而文本+图像+表征组合提供了最佳的整体平衡,SSIM稳定(0.84)且FID具有竞争力(280)。四种机制全部组合时FID最差(330),表明存在冲突的优化信号。九位织工和三十位公众参与者的定性评估确认了统计上显著的积极接受度(Wilcoxon检验,p=0.007和p<0.001)。此外,还开发了一个支持文本到图像和图像到图像生成的基于Web的原型系统,为文化遗产保护提供了实用的数字设计工具。
英文摘要
The traditional Batak Ulos weaving industry faces growing challenges in producing diverse, innovative motifs due to limitations in conventional, manually driven design methods. This study proposes a multimodal generative framework integrating a fine-tuned Latent Diffusion Model (Stable Diffusion XL v1.0 via LoRA) with a Multimodal Large Language Model (LLaMA 1.5-7B) to enable controllable, culturally faithful Ulos motif generation. Four complementary conditioning mechanisms: text, image, representation, and semantic map (via ControlNet) jointly guide the generation process, each governing a distinct aspect from semantic intent to spatial layout. A five level ablation study across three scenarios (shape transformation, colour variation, and high-complexity input) shows that conditioning effectiveness is not proportional to the number of mechanisms combined: Text + Image + Semantic Map achieved the best FID (270) and CLIP Score (0.65 - 0.70) but the weakest SSIM (0.65), while Text + Image + Representation offered the best overall balance, with stable SSIM (0.84) and competitive FID (280). Combining all four mechanisms yielded the weakest FID (330), indicating conflicting optimization signals. Qualitative evaluation by nine weavers and thirty public participants confirmed statistically significant positive acceptance (Wilcoxon, p=0.007 and p<0.001, respectively). A web-based prototype supporting text-to-image and image-to-image generation was also developed, offering a practical digital design tool for cultural heritage preservation.
CommentsC. Anutariya, M.M. Bonsangue, M.N. Mahrin (eds.) "Proceedings of the 4th International Conference on Data Science and Artificial Intelligence (DSAI 2026)", Kuala Lumpur, Malaysia, November 12-13, 2026, in volume 3251 of Communications in Computer and Information Science, Springer, November 2026