VideoReasonBench: Can MLLMs Perform Vision-Centric Complex Video Reasoning?
VideoReasonBench: 大规模语言模型能否进行以视觉为中心的复杂视频推理?
Yuanxin Liu, Kun Ouyang, Haoning Wu, Yi Liu, Lin Sui, Xinhao Li, Yan Zhong, Y. Charles, Xinyu Zhou, Xu Sun
机构
*
State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University(多媒体信息处理国家重点实验室,计算机科学学院,北京大学)
;
Moonshot AI
;
Nanjing University(南京大学)
;
School of Mathematical Sciences, Peking University(数学科学学院,北京大学)
机构
*
School of Computer Science and Technology, Wuhan University of Science and Technology(武汉科技大学计算机科学与技术学院)
;
Hubei Province Key Laboratory of Intelligent Information Processing and Real-time Industrial System(湖北省智能信息处理与实时工业系统重点实验室)
;
Harbin Institute of Technology Zhengzhou Research Institute(哈尔滨工业大学郑州研究所)
;
Guangdong Laboratory of Artificial Intelligence and Digital Economy (SZ)(广东省人工智能与数字经济实验室(深圳))
;
Huawei Technologies Ltd(华为技术有限公司)
机构
*
National Key Laboratory of Human-Machine Hybrid Augmented Intelligence, Institute of Artificial Intelligence and Robotics, Xi’an Jiaotong University(国家人类-机器混合增强智能重点实验室,人工智能与机器人研究院,西安交通大学)
;
Institute of Artificial Intelligence (TeleAI), China Telecom(人工智能研究院(TeleAI),中国电信)
;
University of Science and Technology Beijing(北京科技大学)
机构
*
Shanghai AI Lab(上海人工智能实验室)
;
Northwestern Polytechnical University(西北工业大学)
;
Shanghai Jiao Tong University(上海交通大学)
;
Peking University(北京大学)
;
Nanyang Technological University(南洋理工大学)
;
Beihang University(北京航空航天大学)
;
Sichuan University(四川大学)
;
Tsinghua University(清华大学)
;
The Chinese University of Hong Kong(香港中文大学)
;
Fudan University(复旦大学)
;
Hong Kong University of Science and Technology(香港科技大学)
机构
*
Laboratory of Complex Systems Modeling and Simulation, School of Computer Science and Technology, Hangzhou Dianzi University(电子科技大学复杂系统建模与仿真实验室,计算机科学与技术学院)
;
Zhejiang Key Laboratory of Space Information Sensing and Transmission, Hangzhou Dianzi University(浙江省空间信息感知与传输重点实验室,电子科技大学)
;
Department of Psychological and Cognitive Sciences, Tsinghua University(清华大学心理与认知科学系)
Learning Next Action Predictors from Human-Computer Interaction
从人机交互中学习下一步动作预测器
Omar Shaikh, Valentin Teutschbein, Kanishk Gandhi, Yikun Chi, Nick Haber, Thomas Robinson, Nilam Ram, Byron Reeves, Sherry Yang, Michael S. Bernstein, Diyi Yang
机构
*
Stanford University(斯坦福大学)
;
Hasso Plattner Institute(哈索普拉特纳研究所)
;
New York University(纽约大学)
机构
*
Jockey Club STEM Laboratory of Quantitative Remote Sensing(裘英俊STEM实验室(定量遥感))
;
Department of Geography, the University of Hong Kong(香港大学地理系)
;
School of Information and Electronics, Beijing Institute of Technology(北京理工大学信息电子学院)
;
Beijing Key Laboratory of Fractional Signals and Systems(北京分数信号与系统重点实验室)
;
Department of Electrical and Computer Engineering, University of Delaware(德雷塞尔大学电气与计算机工程系)