机构
*
Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳))
;
Microsoft Research(微软研究院)
;
National University of Singapore(新加坡国立大学)
;
Technical University of Dresden(德累斯顿技术大学)
GeoSurDepth: Harnessing Foundation Model for Spatial Geometry Consistency-Oriented Self-Supervised Surround-View Depth Estimation
GeoSurDepth:利用基础模型实现以空间几何一致性为导向的自监督周围视图深度估计
Weimin Liu, Wenjun Wang, Joshua H. Meng
机构
*
State Key Laboratory of Intelligent Green Vehicle and Mobility, School of Vehicle and Mobility, Tsinghua University, Beijing 100084, China(1 智能绿色车辆与移动国家重点实验室,车辆与移动学院,清华大学,北京100084,中国)
;
California PATH, University of California, Berkeley, CA, United States(2 加州PATH,加州大学伯克利分校,加州,美国)
GeoReason: Aligning Thinking And Answering In Remote Sensing Vision-Language Models Via Logical Consistency Reinforcement Learning
GeoReason: 通过逻辑一致性强化学习对遥感视觉-语言模型中的思考与回答进行对齐
Wenshuai Li, Xiantai Xiang, Zixiao Wen, Guangyao Zhou, Ben Niu, Feng Wang, Lijia Huang, Qiantong Wang, Yuxin Hu
机构
*
Aerospace Information Research Institute, Chinese Academy of Sciences(中国科学院航空信息研究所)
;
Key Laboratory of Target Cognition and Application Technology, Chinese Academy of Sciences(中国科学院目标认知与应用技术重点实验室)
;
University of Chinese Academy of Sciences(中国科学院大学)
机构
*
The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州))
;
The Hong Kong University of Science and Technology(香港科学与技术大学)
;
National University of Singapore(新加坡国立大学)
;
Tsinghua University(清华大学)
Chain-of-Anomaly Thoughts with Large Vision-Language Models
基于大视觉-语言模型的异常思维链
Pedro Domingos, João Pereira, Vasco Lopes, João Neves, David Semedo
机构
*
NOVALINCS, NOVA School of Science and Technology(NOVALINCS、NOVA科学与技术学院)
;
NOVALINCS, Universidade da Beira Interior(NOVALINCS、贝拉内尔大学)
;
DeepNeuronic
专题命中
推理与问题求解
:language model(title,abstract)
AI总结
本文提出CoAT框架,通过引入异常偏见提升大视觉-语言模型在异常检测和分类任务中的性能。
Comments2 pages, 3 figures, 1 table. Accepted for RECPAD 2025
R4: Retrieval-Augmented Reasoning for Vision-Language Models in 4D Spatio-Temporal Space
R4:在4D时空空间中为视觉语言模型引入检索增强推理
Tin Stribor Sohn, Maximilian Dillitzer, Jason J. Corso, Eric Sax
机构
*
Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院)
;
Esslingen University of Applied Sciences(埃斯林根应用科学大学)
;
Dr. Ing. h.c. F. Porsche AG(德意志联邦汽车工业协会)
;
University of Michigan(密歇根大学)
;
Voxel51 Inc.(Voxel51公司)
Synthetic Vasculature and Pathology Enhance Vision-Language Model Reasoning
合成血管和病理增强视觉-语言模型推理
Chenjun Li, Cheng Wan, Laurin Lux, Alexander Berger, Richard B. Rosen, Martin J. Menten, Johannes C. Paetzold
机构
*
Cornell University(康奈尔大学)
;
Weill Cornell Medicine(韦尔·康奈尔医学)
;
Technical University of Munich(慕尼黑技术大学)
;
New York Eye and Ear Infirmary of Mount Sinai(圣文森特医院)
;
Cornell Tech(康奈尔科技)
MindDrive: An All-in-One Framework Bridging World Models and Vision-Language Model for End-to-End Autonomous Driving
MindDrive: 一个整合世界模型和视觉-语言模型的全功能框架,用于端到端自动驾驶
Bin Sun, Yaoguang Cao, Yan Wang, Rui Wang, Jiachen Shang, Xiejie Feng, Jiayi Lu, Jia Shi, Shichun Yang, Xiaoyu Yan, Ziying Song
机构
*
School of Transportation Science and Engineering, Beihang University(北京航空航天大学交通科学与工程学院)
;
State Key Laboratory of Intelligent Transportation System, Beihang University(北京航空航天大学智能交通系统国家重点实验室)
;
Hangzhou International Innovation Institute, Beihang University(北京航空航天大学杭州国际创新院)
;
Contemporary Amperex Technology Co., Limited (CATL)(当代电动车技术有限公司(CATL))
;
Research Institute of Aero-Engine, Beihang University(北京航空航天大学航空发动机研究院)
;
School of Computer Science and Technology, Beijing Jiaotong University(北京交通大学计算机科学与技术学院)
;
China Automotive Engineering Research Institute Co., Ltd.(中国汽车工程研究院股份有限公司)