机构
*
School of Software Engineering, Xi’an Jiaotong University(西安交通大学软件工程学院)
;
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
;
Harbin Institute of Technology, Shenzhen(深圳哈尔滨工业大学)
;
University of Science and Technology Beijing(北京科技大学)
专题命中
视觉推理
:multimodal large language model(abstract);分类 cs.CV
World2Mind: Cognition Toolkit for Allocentric Spatial Reasoning in Foundation Models
World2Mind: 用于基础模型的方位认知工具包
Shouwei Ruan, Bin Wang, Zhenyu Wu, Qihui Zhu, Yuxiang Zhang, Hang Su, Yubin Wang
机构
*
Institute of Artificial Intelligence, Beihang University(北京航空航天大学人工智能研究院)
;
Huawei Noah’s Ark Lab(华为诺亚方舟实验室)
;
Dept. of Comp. Sci. and Tech., Institute for AI, Tsinghua-Bosch Joint ML Center, THBI Lab, BNRist Center, Tsinghua University(清华大学人工智能院计算机科学与技术系、清华大学-博世联合机器学习中心、THBI实验室、BNRist中心、清华大学)
机构
*
School of Software Technology, Zhejiang University(浙江大学软件技术学院)
;
College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院)
;
Ant Group(蚂蚁集团)
Context Matters! Relaxing Goals with LLMs for Feasible 3D Scene Planning
情境至关重要!利用LLMs进行可行的3D场景规划
Emanuele Musumeci, Michele Brienza, Francesco Argenziano, Abdel Hakim Drid, Vincenzo Suriani, Daniele Nardi, Domenico D. Bloisi
机构
*
Department of Computer, Automation and Management Engineering, Sapienza University of Rome(Sapienza大学计算机、自动化与管理工程系)
;
Department of Electronics and Automation, Mohamed Khider University of Biskra(Mohamed Khider大学电子与自动化系)
;
International University of Rome UNINT, Rome(罗马国际大学UNINT)
机构
*
Faculty of Computing, Harbin Institute of Technology(哈尔滨工业大学计算机学院)
;
Tsinghua University(清华大学)
;
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
Institute of Microelectronics of the Chinese Academy of Sciences(中国科学院微电子研究所)
专题命中
视觉推理
:multimodal large language model(abstract);分类 cs.CV
OralGPT-Plus: Learning to Use Visual Tools via Reinforcement Learning for Panoramic X-ray Analysis
OralGPT-Plus:通过强化学习学习使用视觉工具进行全景X射线分析
Yuxuan Fan, Jing Hao, Hong Chen, Jiahao Bao, Yihua Shao, Yuci Liang, Kuo Feng Hung, Hao Tang
机构
*
The Hong Kong University of Science and Technology (GZ)(香港科学与技术大学)
;
Faculty of Dentistry, The University of Hong Kong(香港大学牙医学院)
;
School of Computer Science, Peking University(北京大学计算机学院)
;
Shanghai Jiao Tong University(上海交通大学)
;
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
College of Computer Science and Software Engineering, Shenzhen University(深圳大学计算机科学与软件工程学院)
MultiHaystack: Benchmarking Multimodal Retrieval and Reasoning over 40K Images, Videos, and Documents
MultiHaystack:用于40,000张图像、视频和文档的多模态检索和推理的基准测试
Dannong Xu, Zhongyu Yang, Jun Chen, Yingfang Yuan, Ming Hu, Lei Sun, Luc Van Gool, Danda Pani Paudel, Chun-Mei Feng
机构
*
INSAIT
;
Lanzhou University(兰州大学)
;
King Abdullah University of Science and Technology(国王阿卜杜勒-阿齐兹科学与技术大学)
;
Heriot-Watt University(赫瑞-沃德大学)
;
Monash University(莫纳什大学)
;
University College Dublin(都柏林大学学院)
专题命中
视觉推理
:multimodal large language model(abstract);分类 cs.CV
VideoChat-M1: Collaborative Policy Planning for Video Understanding via Multi-Agent Reinforcement Learning
VideoChat-M1: 通过多智能体强化学习实现视频理解的协作策略规划
Boyu Chen, Zikang Wang, Zhengrong Yue, Kainan Yan, Chenyun Yu, Yi Huang, Zijun Liu, Yafei Wen, Xiaoxin Chen, Yang Liu, Peng Li, Yali Wang
机构
*
Shenzhen Key Lab of Computer Vision and Pattern Recognition(深圳计算机视觉与模式识别重点实验室)
;
Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(中国科学院深圳先进技术研究院)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
;
VIVO AI Lab(VIVO人工智能实验室)
;
Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
;
Shenzhen Campus of Sun Yat-sen University(孙逸仙大学深圳校区)
;
Shanghai Jiao Tong University(上海交通大学)
;
Institute for AI Industry Research (AIR), Tsinghua University(清华大学人工智能产业研究院)
;
Dept. of Comp. Sci. & Tech., Institute for AI, Tsinghua University(清华大学计算机科学与技术系,人工智能研究院)
专题命中
视觉推理
:multimodal large language model(abstract);分类 cs.CV
OmniFashion: Towards Generalist Fashion Intelligence via Multi-Task Vision-Language Learning
OmniFashion: 通过多任务视觉-语言学习实现通用时尚智能
Zhengwei Yang, Andi Long, Hao Li, Zechao Hu, Kui Jiang, Zheng Wang
机构
*
National Engineering Research Center for Multimedia Software, Institute of Artificial Intelligence, School of Computer Science, Wuhan University(国家多媒体软件工程技术研究中心、人工智能研究院、计算机科学学院、武汉大学)
;
Harbin Institute of Technology(哈尔滨工业大学)