arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.25294cs.CVcs.AIcs.CLcs.LG

CLBench-V:评估从基础到知识获取的多模态上下文学习

CLBench-V: Evaluating Multimodal Context Learning from Grounding to Knowledge Acquisition

Lai Wei, Chengqi Li, Jiapeng Li, Ruina Hu, Yue Wang, Weiran Huang

首次发表
浏览论文内容

中文总结 AI 辅助

研究多模态上下文学习问题,引入CLBench-V基准,围绕上下文基础、新信息应用和新知识学习三个维度组织任务,结合公共基准与新数据集,经自动化程序构建,测试六个模型,分析相关因素,揭示多模态上下文学习远未饱和及各模型表现情况。

中文摘要 AI 辅助

现实世界任务常需模型从特定任务上下文学习,而非仅依赖预训练知识。虽近期工作强调此为上下文学习,但现有评估主要聚焦文本上下文。在许多实际场景中,待学习的上下文是多模态的。我们引入CLBench-V,一个多模态上下文学习基准,通过围绕上下文基础、新信息应用和新知识学习三个维度组织任务来解决定位上下文使用故障点的难题。它结合了转换后的公共基准和新构建的数据集。通过自动化构建和过滤程序降低构建特定领域上下文学习任务的成本。在3443个实例和六个多模态模型上测试,最佳总体分数仅0.2847,表明多模态上下文学习远未饱和。此外,InternVL3.5-30B-A3B在上下文基础和新知识学习方面表现最佳,Qwen3.5-Plus在新信息应用方面表现最佳。还进一步分析了判断可靠性、上下文长度、图像数量和代表性失败案例。代码可在指定链接获取。

英文摘要

Real-world tasks often require models to learn from task-specific context rather than relying only on pre-trained knowledge. While recent work has highlighted this capability as context learning, existing evaluations mainly focus on textual contexts. In many practical settings, however, the context to be learned from is multimodal: scientific findings are conveyed through figures and tables, financial indicators are scattered across converted reports, and spatial decisions depend on maps, scenes, or web pages. We introduce CLBench-V, a benchmark for multimodal context learning that addresses the difficulty of localizing where context use breaks down by organizing tasks around three dimensions: context grounding, new information application, and new knowledge learning. CLBench-V combines converted public benchmarks with newly constructed datasets spanning domains such as science, finance, long-document understanding, spatial reasoning, and web-based visual question answering. To reduce the cost of constructing domain-specific context-learning tasks, we further use automated construction and filtering procedures for our newly built datasets. Across 3,443 instances and six recent multimodal models, the best overall score is only 0.2847, indicating that multimodal context learning remains far from saturated. Moreover, InternVL3.5-30B-A3B performs best on context grounding and new knowledge learning, while Qwen3.5-Plus performs best on new information application. We further analyze judge reliability, context length, image count, and representative failure cases. Code is available at https://github.com/IamLihua/CLBench-V.

发表机构

  • School of Computer Science, Shanghai Jiao Tong University(上海交通大学计算机科学与工程学院)
  • Zhongguancun Academy(中关村科学城创新中心)
  • Shanghai Innovation Institute(上海创新研究院)

机构由 AI 辅助整理,请以论文原文为准。

↑