ShotBench: Expert-Level Cinematic Understanding in Vision-Language Models
ShotBench:视觉语言模型中的专家级电影叙事理解
Hongbo Liu, Jingwen He, Yi Jin, Dian Zheng, Yuhao Dong, Fan Zhang, Ziqi Huang, Yinan He, Yangguang Li, Weichao Chen, Yu Qiao, Wanli Ouyang, Shengjie Zhao, Ziwei Liu
机构
*
Tongji University(同济大学)
;
The Chinese University of Hong Kong(香港中文大学)
;
Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
;
S-Lab, Nanyang Technological University(南洋理工大学S-Lab)
RVLM: Recursive Vision-Language Models with Adaptive Depth
具有自适应深度的递归视觉-语言模型
Nicanor Mayumu, Zeenath Khan, Melodena Stephens, Patrick Mukala, Farhad Oroumchian
机构
*
Department of Computer Science(计算机科学系)
;
University of Wollongong in Dubai(迪拜沃林戈大学)
;
Dubai Knowledge Park(迪拜知识园区)
;
Mohammed Bin Rashid School of Government(穆罕默德·本·拉希德政府学院)
MARCUS: An agentic, multimodal vision-language model for cardiac diagnosis and management
MARCUS:一种用于心脏诊断和管理的代理式多模态视觉-语言模型
Jack W O'Sullivan, Mohammad Asadi, Lennart Elbe, Akshay Chaudhari, Tahoura Nedaee, Francois Haddad, Michael Salerno, Li Fe-Fei, Ehsan Adeli, Rima Arnaout, Euan A Ashley
机构
*
Division of Cardiology, Department of Medicine, Stanford University(斯坦福大学心脏病学系)
;
Department of Biomedical Data Science, Stanford University(斯坦福大学生物医学数据科学系)
;
Department of Medicine, Radiology, and Pediatrics, UCSF(旧金山大学医学系、放射学与儿科学系)
;
Bakar Institute, UCSF(Bakar研究所,旧金山大学)
;
UCSF–UC Berkeley Joint Program in Computational Precision Health(旧金山大学-伯克利计算精准健康联合计划)
;
Department of Radiology, Stanford University(斯坦福大学放射学系)
;
Department of Psychiatry and Behavioral Sciences, Stanford University(斯坦福大学精神病学与行为科学系)
;
Department of Computer Science, Stanford University(斯坦福大学计算机科学系)
;
Department of Electrical Engineering, Stanford University(斯坦福大学电气工程系)
;
Department of Biology, Stanford University(斯坦福大学生物学系)
机构
*
Institute for Artificial Intelligence, Peking University(北京大学人工智能研究院)
;
School of Psychological and Cognitive Sciences, Peking University(北京大学心理与认知科学学院)
;
Yuanpei College, Peking University(北京大学元培学院)
;
Department of Psychology, Sun Yat-sen University(中山大学心理学系)
;
State Key Lab of General Artificial Intelligence, Peking University(北京大学通用人工智能国家重点实验室)
;
Beijing Key Laboratory of Behavior and Mental Health, Peking University(北京大学行为与心理健康北京市重点实验室)
机构
*
Institute of Digital Twin, Eastern Institute of Technology(数字孪生研究院,东技术学院)
;
Ocean University of China(中国海洋大学)
;
Shanghai Jiao Tong University(上海交通大学)
;
The Hong Kong Polytechnic University(香港理工大学)
;
Munich Center for Machine Learning, LMU Munich(慕尼黑机器学习中心,慕尼黑大学)
;
Ningbo Key Laboratory of Spatial Intelligence and Digital Derivative(宁波空间智能与数字衍生关键实验室)
专题命中
视觉推理
:vision-language model(title);vision language model(abstract);分类 cs.CV
SportR: A Benchmark for Multimodal Large Language Model Reasoning in Sports
SportR:多模态大语言模型在体育中的推理基准
Haotian Xia, Haonan Ge, Junbo Zou, Hyun Woo Choi, Xuebin Zhang, Danny Suradja, Botao Rui, Ethan Tran, Wendy Jin, Zhen Ye, Xiyang Lin, Christopher Lai, Shengjie Zhang, Junwen Miao, Shichao Chen, Rhys Tracy, Vicente Ordonez, Weining Shen, Hanjie Chen
机构
*
Department of Computer Science, Rice University(Rice大学计算机科学系)
;
Ken Kennedy Institute, Rice University(Rice大学肯尼迪研究所)
;
Department of Statistics, University of California, Irvine(伊利诺伊大学欧文分校统计系)
;
College of Sciences, Georgia Institute of Technology(佐治亚理工学院科学学院)
;
Department of Applied Mathematics and Statistics, Johns Hopkins University(约翰霍普金斯大学应用数学与统计学系)
;
Department of Computer Science, University of California, Santa Barbara(加州大学圣芭芭拉分校计算机科学系)
专题命中
视觉推理
:multimodal large language model(title);grounding(abstract);分类 cs.CV
机构
*
Tsinghua University(清华大学)
;
Peking University(北京大学)
;
Fudan University(复旦大学)
;
Microsoft Research Asia(微软亚洲研究院)
;
Hong Kong University of Science and Technology(香港科技大学)
;
Zhejiang University(浙江大学)
CityLens: Evaluating Large Vision-Language Models for Urban Socioeconomic Sensing
CityLens:评估大型视觉-语言模型用于城市社会经济感知
Tianhui Liu, Hetian Pang, Xin Zhang, Tianjian Ouyang, Zhiyuan Zhang, Jie Feng, Yong Li, Pan Hui
机构
*
Information Hub, The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)信息中心)
;
Department of Electronic Engineering, BNRist, Tsinghua University(清华大学电子工程系)
;
School of Electronic and Information Engineering, Beijing Jiaotong University(北京交通大学电子与信息工程学院)
机构
*
Peking University(北京大学)
;
University of Chinese Academy of Sciences(中国科学院大学)
;
Nankai University(南开大学)
;
Beijing Institute of Technology(北京理工大学)
;
Baichuan Inc.(百度文心)
专题命中
视觉推理
:multimodal large language model(title);MLLM(abstract);分类 cs.CV
机构
*
University of Science and Technology of China(中国科学技术大学)
;
Shanghai AI Laboratory(上海人工智能实验室)
;
East China Normal University(华东师范大学)
;
The Chinese University of Hong Kong(香港中文大学)
Red-teaming the Multimodal Reasoning: Jailbreaking Vision-Language Models via Cross-modal Entanglement Attacks
多模态推理的红队测试:通过跨模态纠缠攻击劫持视觉-语言模型
Yu Yan, Sheng Sun, Shengjia Cheng, Teli Liu, Mingfeng Li, Min Liu
机构
*
Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所)
;
University of Chinese Academy of Sciences(中国科学院大学)
;
People’s Public Security University of China(中国人民公安大学)