机构
*
Institute for AI Industry Research (AIR), Tsinghua University(清华大学人工智能产业研究院(AIR))
;
Bosch Corporate Research, China(博世中国企业研究中心)
;
Institute for Interdisciplinary Information Sciences(IIIS), Tsinghua University(清华大学交叉信息研究院)
VURF: A General-purpose Reasoning and Self-refinement Framework for Video Understanding
Ahmad Mahmood, Ashmal Vayani, Muzammal Naseer, Salman Khan, Fahad Shahbaz Khan
机构
*
ETH Zurich(苏黎世联邦理工学院)
;
University of Central Florida(中佛罗里达大学)
;
Khalifa University(哈利法大学)
;
Mohamed Bin Zayed University of AI(穆罕默德·本·扎耶德人工智能大学)
;
Australian National University(澳大利亚国立大学)
;
Linköping University(林雪平大学)
GaussianGraph: 3D Gaussian-based Scene Graph Generation for Open-world Scene Understanding
Xihan Wang, Dianyi Yang, Yu Gao, Yufeng Yue, Yi Yang, Mengyin Fu
机构
*
School of Automation, Beijing Institute of Technology(北京理工大学自动化学院)
;
National Key Lab of Autonomous Intelligent Unmanned Systems, Beijing Institute of Technology(北京理工大学自主智能无人系统全国重点实验室)
MM-OR: A Large Multimodal Operating Room Dataset for Semantic Understanding of High-Intensity Surgical Environments
Ege Özsoy, Chantal Pellegrini, Tobias Czempiel, Felix Tristram, Kun Yuan, David Bani-Harouni, Ulrich Eck, Benjamin Busam, Matthias Keicher, Nassir Navab
机构
*
Technical University of Munich(慕尼黑工业大学)
;
Munich Center for Machine Learning(慕尼黑机器学习中心)
Probing the limitations of multimodal language models for chemistry and materials research
Nawaf Alampara, Mara Schilling-Wilhelmi, Martiño Ríos-García, Indrajeet Mandal, Pranav Khetarpal, Hargun Singh Grover, N. M. Anoop Krishnan, Kevin Maik Jablonka
机构
*
Northwestern Polytechnical University(西北工业大学)
;
Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)
;
Linköping University(林雪平大学)
;
Australian National University(澳大利亚国立大学)
专题命中
视觉推理
:visual language model(abstract);分类 cs.CV
机构
*
Reallm Labs(Reallm实验室)
;
The Hong Kong Polytechnic University(香港理工大学)
;
Zhejiang University(浙江大学)
;
South China University of Technology(华南理工大学)
;
Harbin Institute of Technology(哈尔滨工业大学)
;
Dalian University of Technology(大连理工大学)
;
University of Electronic Science and Technology of China(电子科技大学)
;
Beijing University of Posts and Telecommunications(北京邮电大学)
;
TikTok(字节跳动旗下TikTok)
;
Amazon(亚马逊)
专题命中
视觉推理
:multimodal large language model(abstract);分类 cs.AI
机构
*
Wangxuan Institute of Computer Technology, Peking University(北京大学王选计算机研究所)
;
National Key Laboratory of General Artificial Intelligence, Peking University(北京大学通用人工智能国家重点实验室)
专题命中
视觉推理
:multimodal large language model(abstract);分类 cs.CV
CommentsAccepted to main conference at NAACL 2025; 8 pages;
机构
*
School of Life Science and Technology, University of Electronic Science and Technology of China (UESTC)(电子科技大学生命科学与技术学院)
;
Brain Health Institute, National Center for Mental Disorders, Shanghai Mental Health Center, Shanghai Jiao Tong University School of Medicine(上海交通大学医学院附属精神卫生中心国家精神障碍中心脑健康研究所)
;
State Key Laboratory of Cognitive Neuroscience and Learning, Beijing Normal University(北京师范大学认知神经科学与学习国家重点实验室)
;
School of Systems Science, Beijing Normal University(北京师范大学系统科学学院)
机构
*
University of Electronic Science and Technology of China(电子科技大学)
;
Sun Yat-sen University(中山大学)
;
University of Washington(华盛顿大学)
;
Microsoft(微软公司)
;
The Chinese University of Hong Kong(香港中文大学)
专题命中
视觉推理
:multimodal large language model(abstract);分类 cs.CV