发表机构
University of Southern California; Microsoft; Microsoft Research(南加州大学; 微软; 微软研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过分析微软Copilot和ChatGPT的大量图像上传对话,提出十能力层级框架,刻画多模态LLM的真实使用任务分布,发现其任务空间更广且现有基准覆盖不均,为基准设计提供依据。
AI 中文摘要
多模态大语言模型(LLMs)日益整合视觉与文本,但人们在自然环境中如何使用它们仍未被充分探索。我们试图回答一个问题:当用户上传图像时,他们试图完成什么任务?通过分析来自微软Copilot的超过40,000次去标识化的图像上传对话,我们通过一个包含十种能力的层级框架(涵盖感知、认知和生成)来刻画真实世界的多模态使用。首先,我们刻画了这些能力的分布和组成,发现大多数图像上传任务涉及多种能力。其次,我们发现多模态使用所覆盖的任务空间比纯文本交互更广泛且更多样,具有不对称的覆盖范围以及依赖跨模态基础的任务类别。第三,将观察到的能力需求映射到253个现有基准上,揭示了基准覆盖与真实世界使用之间的不均衡对齐:基准集中于针对固定答案的感知和推理,而涉及文本、代码和数据生成的常见工作流则相对测试不足。我们在一个独立的ChatGPT数据集上验证了我们的分类法和发现。我们的结果提供了关于用户试图通过多模态LLM完成什么的大规模实证刻画,并强调了基于观察到的用户需求进行基准设计的机会。
英文摘要
Multimodal large language models (LLMs) increasingly integrate vision and text, yet how people use them in natural settings remains underexplored. We seek to answer the question: when users upload images, what tasks are they trying to accomplish? Analyzing over 40,000 de-identified image-upload conversations from Microsoft Copilot, we characterize real-world multimodal use through a hierarchical framework of ten capabilities, spanning perception, cognition, and generation. First, we characterize the distribution and composition of these capabilities, finding that the majority of image-upload tasks involve multiple capabilities. Second, we find that multimodal use spans a broader and more diverse task space than text-only interactions, with asymmetric coverage and task classes that rely on cross-modal grounding. Third, mapping observed capability demand onto 253 existing benchmarks reveals uneven alignment between benchmark coverage and real-world use: benchmarks concentrate on perception and reasoning toward fixed answers, while common workflows involving text, code, and data generation remain comparatively undertested. We validate our taxonomy and findings on an independent ChatGPT dataset. Our results provide a large-scale empirical characterization of what users seek to accomplish with multimodal LLMs and highlight opportunities for benchmark design grounded in observed user demand.