机构
*
School of Vehicle and Mobility, Tsinghua University(清华大学车辆与移动系统学院)
;
Automotive and Robotics, Xiaomi Corporation(小鹏汽车与机器人部)
;
McGill University(麦吉尔大学)
;
University of Wisconsin-Madison(威斯康星大学麦迪逊分校)
机构
*
Università di Roma La Sapienza(罗马La Sapienza大学)
;
Fondazione Bruno Kessler (FBK)(布鲁诺·凯斯勒基金会)
;
Free University of Bozen-Bolzano(博兹纳-博尔扎诺自由大学)
;
Dept. of Computer Science, University of Pisa(比萨大学计算机科学系)
;
CoLing Lab, Dept. of Philology, Literature and Linguistics, University of Pisa(比萨大学语言学、文学与语言学系CoLing实验室)
;
Istituto di Linguistica Computazionale "A. Zampolli" (CNR-ILC), ItaliaNLP Lab, Pisa(A. Zampolli计算语言学研究所(CNR-ILC),意大利NLP实验室,比萨)
专题命中
视觉推理
:vision language model(abstract);visual language model(abstract)
Understanding Space Is Rocket Science -- Only Top Reasoning Models Can Solve Spatial Understanding Tasks
Nils Hoehing, Mayug Maniparambil, Ellen Rushe, Noel E. O'Connor, Anthony Ventresque
机构
*
School of Computer Science(计算机科学学院)
;
University College Dublin(都柏林大学)
;
School of Computing(计算机科学学院)
;
Dublin City University(都柏林城市大学)
;
School of Electronic Engineering(电子工程学院)
;
Trinity College Dublin(都柏林三一学院)
ReasoningTrack: Chain-of-Thought Reasoning for Long-term Vision-Language Tracking
Xiao Wang, Liye Jin, Xufeng Lou, Shiao Wang, Lan Chen, Bo Jiang, Zhipeng Zhang
机构
*
School of Computer Science and Technology, Anhui University(安徽大学计算机科学与技术学院)
;
School of Artificial Intelligence, Shanghai Jiao Tong University(上海交通大学人工智能学院)
;
School of Electronic and Information Engineering, Anhui University(安徽大学电子与信息工程学院)
KeyMPs: One-Shot Vision-Language Guided Motion Generation by Sequencing DMPs for Occlusion-Rich Tasks
Edgar Anarossi, Yuhwan Kwon, Hirotaka Tahara, Shohei Tanaka, Keisuke Shirai, Masashi Hamaya, Cristian C. Beltran-Hernandez, Atsushi Hashimoto, Takamitsu Matsubara
机构
*
Division of Information Science, Graduate School of Science and Technology, Nara Institute of Science and Technology(信息科学系,科学技术研究生学校,科学技术研究所)
;
Department of Electrical and Electronic Engineering, Faculty of Engineering Science, Kansai University(电气电子工程系,工学科学大学)
;
Department of Electronics, Kobe City College of Technology(电子系,神户市立技术学院)
;
OMRON SINIC X Corporation(OMRON SINIC X公司)
ArtGS:3D Gaussian Splatting for Interactive Visual-Physical Modeling and Manipulation of Articulated Objects
Qiaojun Yu, Xibin Yuan, Yu jiang, Junting Chen, Dongzhe Zheng, Ce Hao, Yang You, Yixing Chen, Yao Mu, Liu Liu, Cewu Lu
机构
*
Shanghai Jiao Tong University(上海交通大学)
;
Shanghai AI Laboratory(上海人工智能实验室)
;
National University of Singapore(国立新加坡大学)
;
Princeton University(普林斯顿大学)
;
Stanford University(斯坦福大学)
;
Hefei University of Technology(合肥工业大学)
机构
*
School of Artificial Intelligence, Nanjing University, China(人工智能学院,南京大学)
;
National Key Laboratory for Novel Software Technology, Nanjing University, China(新型软件技术国家重点实验室,南京大学)
;
AI Business, Alibaba Group(阿里集团人工智能业务)
专题命中
视觉推理
:multimodal large language model(abstract);分类 cs.CV、cs.AI、cs.LG
机构
*
The Chinese University of Hong Kong(香港中文大学)
;
State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University(北京大学计算机科学学院多媒体信息处理国家重点实验室)
;
AI 2 Robotics(人工智能与机器人)
;
Beijing Academy of Artificial Intelligence (BAAI)(北京人工智能研究院)