VGI-Bench: Probing Visual Intelligence in Video Generation Models
VGI-BENCH:探究视频生成模型的视觉智能
Xuan He, Cong Wei, Yuhao Cheng, Linrui Ma, Yuxuan Zhang, Zuojun Li, Yuhao Wen, Jize Jiang, Zeyi Liu, Yuren Hao, Songcheng Cai, Keming Wu, Penghui Du, Kai Zou, Rui Yang, Chenkai Sun, Ke Yang, Ping Nie, Kelsey R Allen, Chenglong Wang, Michel Galley, Jianfeng Gao, ChengXiang Zhai
机构
*
University of Illinois Urbana Champaign(伊利诺伊大学厄巴纳-香槟分校)
;
Tsinghua University(清华大学)
;
University of Waterloo(滑铁卢大学)
;
Massachusetts Institute of Technology(麻省理工学院)
;
University of British Columbia(不列颠哥伦比亚大学)
;
Vector Institute(矢量研究所)
;
Microsoft Research(微软研究院)
;
Etude AI
Weihao Tan, Changjiu Jiang, Yu Duan, Mingcong Lei, Jiageng Li, Yitian Hong, Xinrun Wang, Bo An
机构
*
Nanyang Technological University(南洋理工大学)
;
Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳))
;
The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
;
East China University of Science and Technology(华东理工大学)
;
Singapore Management University(新加坡管理大学)
专题命中
视觉推理
:multimodal large language model(abstract);分类 cs.AI
GTPred: Benchmarking MLLMs for Interpretable Geo-localization and Time-of-capture Prediction
GTPred:评估多模态大语言模型在可解释地理定位和拍摄时间预测中的基准测试
Jinnao Li, Tingzhu Chen, Changbo Wang
机构
*
School of Computer Science and Technology, East China Normal University(东华大学计算机科学与技术学院)
;
Institute of Image Communication and Information Processing, Shanghai Jiao Tong University(上海交通大学图像通信与信息处理研究院)
;
School of Humanities, Shanghai Jiao Tong University(上海交通大学人文学院)
;
Shanghai AI Laboratory(上海人工智能实验室)
Comments17 pages; Previously this version appeared as arXiv:2510.15430 (https://arxiv.org/abs/2510.15430) which was submitted as a new work by accident
机构
*
State Key Laboratory of Novel Software Technology, Nanjing University, China(南京大学新型软件技术国家重点实验室)
;
School of Artificial Intelligence, Nanjing University, China(南京大学人工智能学院)
专题命中
VLM训练与架构
:grounding(abstract);multimodal large language model(abstract);分类 cs.AI、cs.LG