ViSTA: Visual Storytelling using Multi-modal Adapters for Text-to-Image Diffusion Models
ViSTA: 基于多模态适配器的文本到图像扩散模型用于视觉叙事
机构 * Department of Computer Science, Georgetown University(计算机科学系,乔治城大学)
专题命中 文生图 :text-to-image(title,abstract);diffusion(title,abstract);分类 cs.CV
AI总结 ViSTA通过多模态历史适配器提升文本到图像扩散模型在视觉叙事中的生成一致性与文本对齐能力。
Comments Accepted to WACV 2026