机构
*
State Key Laboratory of General Artificial Intelligence, BIGAI(通用人工智能国家重点实验室(BIGAI))
;
Beijing Institute of Technology(北京理工大学)
;
Institute for Artificial Intelligence, Peking University(北京大学人工智能研究院)
;
State Key Lab of General AI, School of Intelligence Science and Technology, Peking University(通用人工智能国家重点实验室、北京大学智能科学与技术学院)
MPCoT: Reward-Guided Multi-Path Latent Reasoning for Test-Time Scalable Vision-Language-Action
MPCoT: 奖励引导的多路径潜在推理用于测试时可扩展的视觉-语言-动作
Boyang Zhang, Lianlei Shan
机构
*
Department of Electrical and Computer Engineering, Boston University(波士顿大学电气与计算机工程系)
;
Department of Computer Science, Tsinghua University(清华大学计算机系)
GroupForward: Building Referable 3D Scenes via Instance-Grouped Feed-Forward Gaussian Splatting
GroupForward:通过实例分组前馈高斯溅射构建可参考的3D场景
Qijian Tian, Zimeng Wu, Xuhong Wang, Lizhuang Ma, Xin Tan
机构
*
School of Computer Science, Shanghai Jiao Tong University(上海交通大学计算机科学学院)
;
School of Computer Science and Engineering, Beihang University(北京航空航天大学计算机科学与工程学院)
;
Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
;
School of Computer Science and Technology, East China Normal University(华东师范大学计算机科学与技术学院)
机构
*
Department of Computer Science and Engineering, Hong Kong University of Science and Technology(香港科技大学计算机科学与工程系)
;
Department of Pathology, Nanfang Hospital, Southern Medical University(南方医科大学南芳医院病理科)
;
Department of Pathology, School of Basic Medical Sciences, Southern Medical University(南方医科大学基础医学学院病理科)
;
Department of Anatomical and Cellular Pathology, Chinese University of Hong Kong(香港中文大学解剖与细胞病理学系)
;
Guangdong Provincial Key Laboratory of Molecular Tumor Pathology(广东省分子肿瘤病理学重点实验室)
;
Jinfeng Laboratory(锦风实验室)
;
Department of Chemical and Biological Engineering, Hong Kong University of Science and Technology(香港科技大学化学与生物工程系)
;
Division of Life Science, Hong Kong University of Science and Technology(香港科技大学生命科学系)
;
State Key Laboratory of Nervous System Disorders, The Hong Kong University of Science and Technology(香港科技大学神经系统疾病国家重点实验室)
;
HKUST Shenzhen-Hong Kong Collaborative Innovation Research Institute, The Hong Kong University of Science and Technology(香港科技大学深圳-香港协同创新研究院)
GazeXPErT: An Expert Eye-tracking Dataset for Interpretable and Explainable AI in Oncologic FDG-PET/CT Scans
GazeXPErT: 一种用于肿瘤FDG-PET/CT扫描可解释和可解释AI的专家眼动数据集
Joy T Wu, Daniel Beckmann, Sarah Miller, Alexander Lee, Elizabeth Theng, Stephan Altmayer, Ken Chang, David Kersting, Tomoaki Otani, Brittany Z Dashevsky, Hye Lim Park, Matteo Novello, Kip Guja, Curtis Langlotz, Ismini Lourentzou, Daniel Gruhl, Benjamin Risse, Guido A Davidzon
AeroGround: A Comprehensive Benchmark for Aerial-Ground Collaborative Reasoning
AeroGround:用于空-地协同推理的综合基准
Shenghong Yi, Lin Zhang, Muzian Li, Jiakang Yuan, Haoyu Zhang, Peng Ye, Jiayuan Fan, Huafeng Qin, Tao Chen
机构
*
Shanghai Innovation Institute(上海创新研究院)
;
College of Future Information Technology, Fudan University(复旦大学未来信息技术学院)
;
College of Intelligent Robotics and Advanced Manufacturing, Fudan University(复旦大学智能机器人与先进制造学院)
;
Chongqing Technology and Business University(重庆工商大学)
;
The Chinese University of Hong Kong(香港中文大学)
How Good are Foundation Models in Longitudinal MRI Disease Progression Reasoning?
基础模型在纵向MRI疾病进展推理中的表现如何?
Wafa Al Ghallabi, Ritesh Thawkar, Sara Ghaboura, Omkar Thawakar, Numan Saeed, Dana Al Nuaimi, Ajnas Alkatheeri, Salman Khan, Fahad Shahbaz Khan
机构
*
Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)
;
Department of Health Abu Dhabi(阿布扎比卫生部)
;
Fatima College of Health Sciences(法蒂玛健康科学学院)
;
Linköping University(林雪平大学)
GESTO: Human-Centric Spatio-Temporal Memory for Reasoning in Dynamic Scenes
GESTO:面向动态场景推理的以人为中心的时空记忆
Ermanno Bartoli, Buwei He, Dennis Rotondi, Sebastian Koch, Federico Tombari, Kai O. Arras, Patric Jensfelt, Yixi Cai, Iolanda Leite
机构
*
KTH Royal Institute of Technology(瑞典皇家理工学院)
;
University of Stuttgart(斯图加特大学)
;
Max Planck Research School for Intelligent Systems(马克斯·普朗克智能系统研究所)
;
Google Research(谷歌研究院)
;
TU Munich(慕尼黑工业大学)
Advantage-Guided Gate: Reshaping Open-Ended Reasoning for Vision-Based Spatial Intelligence
优势引导门:重塑基于视觉的空间智能的开放式推理
Ling Lin, Yang Bai, Congcong Zhu, Jiangming Shi, Meng Wang, Yang Long, Jingrun Chen, Ling Shao, Huazhu Fu
机构
*
University of Science and Technology of China(中国科学技术大学)
;
Suzhou Institute for Advanced Research, USTC(中国科学技术大学苏州高等研究院)
;
Institute of Advanced Intelligence and Computing (IAIC), Agency for Science, Technology and Research (A*STAR)(新加坡科技研究局高级智能与计算研究所)
;
East China Normal University(华东师范大学)
;
National University of Singapore(新加坡国立大学)
;
Durham University(杜伦大学)
;
University of Chinese Academy of Sciences(中国科学院大学)
专题命中
视觉推理
:multimodal large language model(abstract);分类 cs.CV
SCOUT: Self-Checking and Recovery-Aware Tool-Thought Agents for Ultra-Long Egocentric Video Reasoning
SCOUT:面向超长时长自我中心视频推理的自检查与恢复感知工具思维智能体
Keyang Zhong, Kuo Wang, Peng Liu, Quanlong Zheng, Junlin Xie, Zhijia Liang, Yanhao Zhang, Guanbin Li
机构
*
Sun Yat-sen University(中山大学)
;
Shenzhen Loop Area Institute(深圳河套学院)
;
Guangdong OPPO Mobile Telecommunications Corp., Ltd.(广东欧珀移动通信有限公司)
;
OPPO AI Center(OPPO人工智能中心)
;
The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))