CommentsThis work was submitted without the consent of my current adviser. Additionally, it overlaps with my unpublished research work. In order to avoid potential academic and authorship conflicts, I am requesting withdrawal of the paper
机构
*
College of Computing and Data Science, Nanyang Technological University(计算与数据科学学院,南洋理工大学)
;
CFAR and I2R, Agency for Science, Technology and Research(CFAR和I2R,科技研究局)
;
Boston Children’s Hospital, Harvard Medical School(哈佛医学院儿童医院)
C^2ROPE: Causal Continuous Rotary Positional Encoding for 3D Large Multimodal-Models Reasoning
C^2ROPE: 3D 大多模态模型推理中的因果连续旋转位置编码
Guanting Ye, Qiyan Zhao, Wenhao Yu, Xiaofeng Zhang, Jianmin Ji, Yanyong Zhang, Ka-Veng Yuen
机构
*
State Key Laboratory of Internet of Things for Smart City, University of Macau(物联网智能城市国家重点实验室,澳门大学)
;
Department of Automation, Shanghai Jiaotong University(上海交通大学自动化系)
;
Institute of Advanced Technology, University of Science and Technology of China(中国科学技术大学先进技术研究院)
;
School of Computer Science and Technology, USTC(中国科学技术大学计算机科学与技术学院)
;
School of Artificial Intelligence and Data Science, USTC(中国科学技术大学人工智能与数据科学学院)
Talk2DM: Enabling Natural Language Querying and Commonsense Reasoning for Vehicle-Road-Cloud Integrated Dynamic Maps with Large Language Models
Talk2DM: 通过大语言模型实现车辆-道路-云集成动态地图的自然语言查询与常识推理
Lu Tao, Jinxuan Luo, Yousuke Watanabe, Zhengshu Zhou, Yuhuan Lu, Shen Ying, Pan Zhang, Fei Zhao, Hiroaki Takada
机构
*
School of Resource and Environmental Sciences, Wuhan University(武汉大学资源与环境科学学院)
;
School of Earth Sciences, Yunnan University(云南大学地球科学学院)
;
College of Artificial Intelligence, Tianjin University of Science & Technology(天津科技大学人工智能学院)
;
Department of Computer and Information Engineering, Khalifa University(卡利法大学计算机与信息工程系)
;
NVIDIA
;
Institutes of Innovation for Future Society, Nagoya University(名古屋大学创新未来社会研究所)
VISOR: VIsual Spatial Object Reasoning for Language-driven Object Navigation
VISOR:基于语言驱动的对象导航的视觉空间对象推理
Francesco Taioli, Shiping Yang, Sonia Raychaudhuri, Marco Cristani, Unnat Jain, Angel X Chang
机构
*
Polytechnic of Turin(都灵理工大学)
;
Simon Fraser University(Simon Fraser大学)
;
University of Verona(威尼斯大学)
;
University of Reykjavik(雷克雅未克大学)
;
University of California, Irvine(加州大学尔湾分校)
Visual serial processing deficits explain divergences in human and VLM reasoning
视觉序列处理缺陷解释了人类与VLM推理之间的差异
Nicholas Budny, Kia Ghods, Declan Campbell, Raja Marjieh, Amogh Joshi, Sreejan Kumar, Jonathan D. Cohen, Taylor W. Webb, Thomas L. Griffiths
机构
*
Princeton Neuroscience Institute(普林斯顿神经科学研究所)
;
Department of Psychology, Princeton University(普林斯顿大学心理学系)
;
Department of Psychology, Université de Montréal(蒙特利尔大学心理学系)
;
Mila - Quebec AI Institute(魁北克AI研究所)
;
Department of Computer Science, Princeton University(普林斯顿大学计算机科学系)
MedVL-SAM2: A unified 3D medical vision-language model for multimodal reasoning and prompt-driven segmentation
MedVL-SAM2:一种统一的3D医学视觉-语言模型,用于多模态推理和基于提示的分割
Yang Xing, Jiong Wu, Savas Ozdemir, Ying Zhang, Yang Yang, Wei Shao, Kuang Gong
机构
*
Department of Biomedical Engineering, University of Florida(佛罗里达大学生物医学工程系)
;
Department of Radiology, University of Florida(佛罗里达大学放射学系)
;
Research Computing, University of Florida(佛罗里达大学研究计算中心)
;
Department of Medicine, University of Florida(佛罗里达大学医学系)
;
Department of Radiology, UC San Francisco(旧金山大学放射学系)