VTCBench: Can Vision-Language Models Understand Long Context with Vision-Text Compression?
VTCBench: 视觉-语言模型能否通过视觉-文本压缩理解长上下文?
机构 * 1 Institute of Automation, Chinese Academy of Sciences ; 2 School of Artificial Intelligence, University of Chinese Academy of Sciences ; 3 Centre for Artificial Intelligence ; Robotics, Hong Kong Institute of Science \& Innovation, CAS ; 4 Independent Researcher
专题命中 视觉问答 :vision-language model(title,abstract);分类 cs.CV、cs.AI
AI总结 VTCBench评估视觉-文本压缩对视觉语言模型长上下文理解能力的影响,发现多数模型在处理压缩信息时表现不佳。