发表机构
Tufts University(塔夫茨大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出一个基于51个系统的生成式图像模型交互式视觉界面设计空间,通过UI、可控对象和映射函数三部分分解系统,并构建语料库浏览器以支持比较分析,助力发现设计机会。
AI 中文摘要
交互式视觉界面已成为控制生成式图像模型的重要手段,使用户能够通过提示词、直接操作以及一系列交互方式来操控生成过程。然而,现有技术通常以独立系统的形式呈现,这使得人们难以理解它们之间的关联、比较其交互机制,或识别新界面设计的机会。我们提出了一个针对生成式图像模型交互式视觉界面的设计空间,该设计空间源自对51个研究系统和实践者工具的归纳。该框架将每个系统分解为三个互补的组成部分:用户界面(UI)、可控模型对象(Z)以及将用户交互转换为模型操作的映射函数($\phi$)。这种分解为分析跨模型家族的异构交互技术提供了统一的表示方法,揭示了重复出现的设计模式以及设计空间中尚未充分探索的区域。我们还展示了一个支持比较分析的交互式语料库浏览器,并讨论了在教育场景以及HCI/AI实践者识别研究和设计机会方面的使用场景。
英文摘要
Interactive visual interfaces have become an important means of controlling generative image models, enabling users to manipulate generation through prompts, direct manipulation, and a range of interactions. However, existing techniques are typically presented as independent systems, making it difficult to understand how they relate, compare their interaction mechanisms, or identify opportunities for new interface designs. We introduce a design space for interactive visual interfaces for generative image models derived from 51 research systems and practitioner tools. The framework decomposes each system into three complementary components: the user interface (U), the controllable model objects (Z), and the mapping function ($ϕ$) that translates user interaction into model operations. This decomposition provides a common representation for analyzing heterogeneous interaction techniques across model families, revealing recurring design patterns and underexplored regions of the design space. We further present an interactive corpus explorer that support comparative analysis, and discuss usage scenarios for both educational settings and HCI/AI practitioners identifying research and design opportunities.