发表机构
Canva Research(Canva研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Giraffe 是一种轻量映射架构,以单 [IMG] token 实现从文本隐藏表示到视觉模型嵌入空间的转换,在平面设计生成任务中表现优异。
AI 中文摘要
多模态大语言模型(MLLMs)在理解和解释多媒体内容方面已取得显著进展,但其生成媒体的能力仍有限。近期研究尝试通过将 token 序列的隐藏表示转换为视觉模型的嵌入空间,或直接转换为原始图像数据来缩小这一差距。然而,这些方法通常用多个专用 token 表示单张图像,大幅增加输入长度,这对平面设计生成等任务构成重大限制——此类任务的输出通常需无缝融合文本、多幅图像及布局信息的数千个 token。为应对该挑战,本文提出一种新型架构,将隐藏 token 表示映射至视觉模型(如 CLIP ViT-L/14)的嵌入空间,每张图像仅用一个 [IMG] token。该架构采用两个浅层 MLP 块,每个块含独立压缩模块及共享扩展模块,通过六种不同损失函数训练;其中一个块在训练期间辅助另一个,推理阶段则省略,最终形成轻量解决方案。该方法在图像到设计及文本到设计生成任务中均展现出优异性能。
英文摘要
Multimodal large language models (MLLMs) have made significant progress in understanding and interpreting mul- timedia content. However, their ability to generate me- dia remains limited. Recent approaches have attempted to bridge this gap by translating the hidden representations of token sequences into the embedding space of visual models or directly into raw image data. However, these methods often represent each image using multiple specialised to- kens which significantly increases the input length. This be- comes a major limitation for tasks such as graphic design generation where the output typically involves a seamless blend of thousands of tokens across text, multiple images, and layout information. To address this challenge, a novel architecture is proposed that maps hidden token represen- tations to the embedding space of visual models, such as CLIP ViT-L/14, using a single [IMG] token per image. The architecture employs two shallow MLP blocks, each with a separate compression module followed by a shared expan- sion module, trained with six distinct loss functions. One block aids the other during training and is omitted during inference, resulting in a lightweight solution. Strong perfor- mance is demonstrated in both image-to-design and text-to- design generation tasks.