发表机构
School of Electronic Engineering and Computer Science, Queen Mary University of London(伦敦玛丽女王大学电子工程与计算机科学学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出DSV系统,通过大型语言模型拆分梦境描述为四部分,结合文本到图像模型生成连贯的四面板图像序列,经DreamBank数据集50次可视化评估,用CLIP等模型指标验证了其质量、保真度与连贯性。
AI 中文摘要
梦境情感强烈但难以传达。我们提出了Dream Scene Visualiser(DSV)系统,可将书面梦境描述转化为可视化梦境的四面板图像时间序列。该系统先通过大型语言模型将梦境描述拆分为四个时间顺序部分,再由文本到图像模型为各部分生成图像,同时保持序列间的视觉连贯性,DSV还会重新生成与文本不匹配的图像。我们在DreamBank的梦境描述上对DSV的50次可视化结果进行评估,采用CLIP、DINOv2和Qwen2-VL视觉语言模型的客观指标,报告了质量、保真度和连贯性结果。
英文摘要
Dreams can be emotionally intense but difficult to communicate. We describe the Dream Scene Visualiser (DSV) system which turns written dream descriptions into a temporal sequence of four panel images visualising the dream. This starts with a large language model prompted to split a dream description into four chronological parts. Then a text-to-image model produces images for each part with visual coherence maintained across the sequence, and DSV regenerates any image not suitably matching the text. We evaluate DSV over 50 visualisations from dream descriptions in DreamBank, and report quality, fidelity and coherence results via objective measures employing the CLIP, DINOv2 and Qwen2-VL vision-language models.
Commentsshort paper accepted at ICCC 2026