arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2412.08687cs.CV

VisionArena:23万真实世界用户与VLM对话及偏好标签数据集

VisionArena: 230K Real World User-VLM Conversations with Preference Labels

  • Stanford(斯坦福大学)
  • UC Berkeley(加州大学伯克利分校)

机构由 AI 辅助整理,请以论文原文为准。

Christopher Chou, Lisa Dunlap, Koki Mashita, Krishna Mandal, Trevor Darrell, Ion Stoica, Joseph E. Gonzalez, Wei-Lin Chiang

更新

AI总结:

VisionArena构建了23万条真实用户与VLM的对话及偏好标签数据集,涵盖三个子集,并发现开放式任务依赖风格、VLM在空间推理上不足,微调后性能显著提升。

AI中文摘要:

随着视觉语言模型(VLM)的日益普及和能力的提升,需要能够捕捉真实用户与VLM交互的基准测试。为此,我们创建了VisionArena,一个包含23万条用户与VLM之间真实世界对话的数据集。该数据集收集自Chatbot Arena——一个用户与VLM互动并提交偏好投票的开源平台——涵盖了7.3万名独立用户、45个VLM和138种语言。我们的数据集包含三个子集:VisionArena-Chat,包含20万条用户与VLM之间的单轮和多轮对话;VisionArena-Battle,包含3万条比较两个匿名VLM并附有用户偏好投票的对话;以及VisionArena-Bench,一个包含500个多样化用户提示的自动基准,能够高效地近似实时Chatbot Arena模型排名。此外,我们重点分析了用户提出的问题类型、回答风格对偏好的影响以及模型经常失败的领域。我们发现开放式任务(如图像描述和幽默)高度依赖于风格,而当前VLM在空间推理和规划任务上表现不佳。最后,我们展示了在VisionArena-Chat上微调相同基础模型优于Llava-Instruct-158K,在MMMU上获得17个百分点的提升,在WildVision基准上获得46个百分点的提升。数据集可在https://huggingface.co/lmarena-ai获取。

英文摘要:

With the growing adoption and capabilities of vision-language models (VLMs) comes the need for benchmarks that capture authentic user-VLM interactions. In response, we create VisionArena, a dataset of 230K real-world conversations between users and VLMs. Collected from Chatbot Arena - an open-source platform where users interact with VLMs and submit preference votes - VisionArena spans 73K unique users, 45 VLMs, and 138 languages. Our dataset contains three subsets: VisionArena-Chat, 200k single and multi-turn conversations between a user and a VLM; VisionArena-Battle, 30K conversations comparing two anonymous VLMs with user preference votes; and VisionArena-Bench, an automatic benchmark of 500 diverse user prompts that efficiently approximate the live Chatbot Arena model rankings. Additionally, we highlight the types of question asked by users, the influence of response style on preference, and areas where models often fail. We find open-ended tasks like captioning and humor are highly style-dependent, and current VLMs struggle with spatial reasoning and planning tasks. Lastly, we show finetuning the same base model on VisionArena-Chat outperforms Llava-Instruct-158K, with a 17-point gain on MMMU and a 46-point gain on the WildVision benchmark. Dataset at https://huggingface.co/lmarena-ai

补充信息

↑