发表机构
KAIST; Texas A&M University; Adobe Research(韩国科学技术院; 德克萨斯农工大学; Adobe研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过收集1600个草图构建AnnoSketch数据集,分析草图输入对MLLM图表标注的帮助时机,并定性评估其有用性,为交互式标注提供经验指导。
AI 中文摘要
随着多模态大语言模型(MLLMs)支持越来越多的输入模态,越来越多的研究探索如何结合粗略草图来传达用户意图。对于带标注的图表生成,目前尚不清楚人们会提供什么样的标注草图,以及这种视觉输入何时能帮助MLLMs生成更有用的标注。在本研究中,我们考察了在不同图表类型和标题类型下,草图输入对MLLM生成的图表标注何时有用。此外,我们定性分析了参与者对其输出偏好的解释,以描述使生成标注更有用或更不有用的因素。为了进一步记录参与者的标注草图,我们提出了AnnoSketch,包含从160个图表-标题对中收集的1,600个标注草图,这些草图来自草图指导被证明最有益的条件,同时还包括参与者的标注意图、感知理解难度和自我报告的表达限制。我们还为这些草图标注了结构化元数据,描述每个草图与其标题的关系以及参与者如何通过视觉标记表达标注。总之,我们的研究和AnnoSketch有助于确定何时征求草图输入,并为人们如何绘制图表标注以支持标题提供经验来源。数据集和补充材料可在我们的OSF存储库中获取。
英文摘要
As multimodal large language models (MLLMs) support a growing range of input modalities, increasing work explores how to incorporate rough sketches to convey user intent. For annotated chart generation, it remains unclear what annotation sketches people provide and when such visual input helps MLLMs generate more useful annotations. In this study, we examine when sketch input is useful for MLLM-generated chart annotations across variation in chart type and caption type. In addition, we qualitatively analyze participants' explanations of their output preferences to characterize what made generated annotations more or less helpful. To further document participants' annotation sketches, we present AnnoSketch, comprising 1,600 annotation sketches collected across 160 chart-caption pairs from the conditions in which sketch guidance proved most beneficial, together with participants' annotation intents, perceived comprehension difficulty, and self-reported expressive limitations. We also label these sketches with structured metadata describing how each sketch relates to its caption and how participants express annotations through visual marks. Together, our study and AnnoSketch help determine when to solicit sketch input and provide empirical source for how people sketch chart annotations to support captions. The dataset and supplemental materials are available in our OSF repository.