Synthetic Vasculature and Pathology Enhance Vision-Language Model Reasoning
合成血管和病理增强视觉-语言模型推理
Chenjun Li, Cheng Wan, Laurin Lux, Alexander Berger, Richard B. Rosen, Martin J. Menten, Johannes C. Paetzold
机构
*
Cornell University(康奈尔大学)
;
Weill Cornell Medicine(韦尔·康奈尔医学)
;
Technical University of Munich(慕尼黑技术大学)
;
New York Eye and Ear Infirmary of Mount Sinai(圣文森特医院)
;
Cornell Tech(康奈尔科技)
CAPE: A CLIP-Aware Pointing Ensemble of Complementary Heatmap Cues for Embodied Reference Understanding
CAPE:一种基于CLIP的互补热图线索点集用于具身参照理解
Fevziye Irem Eyiokur, Dogucan Yaman, Hazım Kemal Ekenel, Alexander Waibel
机构
*
Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院)
;
Istanbul Technical University(伊斯坦布尔技术大学)
;
Carnegie Mellon University(卡内基梅隆大学)
;
KIT Campus Transfer GmbH (KCT)(KIT校园转移有限责任公司)
机构
*
School of Control Science and Engineering, Shandong University, Jinan, 250061, China(控制科学与工程学院,山东大学,济南,250061,中国)
;
Engineering Research Center of Intelligent Unmanned System, Ministry of Education, China(智能无人系统工程研究中心,教育部,中国)
;
Department of Geriatric Neurology, Qilu Hospital of Shandong University, Jinan, 250061, China(老年神经内科,山东大学齐鲁医院,济南,250061,中国)
;
Shandong Inspur Science Research Institute Co., Ltd.(山东浪潮科学研究院有限公司)
机构
*
Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
;
Shanghai JiaoTong University(上海交通大学)
;
University of Chinese Academy of Sciences(中国科学院大学)
;
Zhejiang University(浙江大学)
Steering Vision-Language-Action Models as Anti-Exploration: A Test-Time Scaling Approach
引导视觉-语言-动作模型作为反探索:一种测试时间缩放方法
Siyuan Yang, Yang Zhang, Haoran He, Ling Pan, Xiu Li, Chenjia Bai, Xuelong Li
机构
*
Institute of Artificial Intelligence, China Telecom(中国电信人工智能研究院)
;
University of Science and Technology of China(中国科学技术大学)
;
Tsinghua University(清华大学)
;
The Hong Kong University of Science and Technology(香港科学与技术大学)
机构
*
MoE Key Laboratory of Brain-inspired Intelligent Perception and Cognition, University of Science and Technology of China(脑启发智能感知与认知联合实验室,中国科学技术大学)
;
ZheJiang University(浙江大学)
;
The Hong Kong University of Science and Technology(香港科技大学)