机构
*
Dwarkadas Jivanlal Sanghvi College of Engineering(达沃拉斯·吉万拉尔·桑格维工程学院)
;
King’s College London(伦敦国王学院)
;
Indian Institute of Technology Jodhpur(印度理工学院朱罗普尔)
TraversalBench: Challenging Paths to Follow for Vision Language Models
TraversalBench: 为视觉语言模型设计的复杂路径挑战测试集
Clara Petrova, Zhuo Chen, Marin Soljačić
机构
*
Massachusetts Institute of Technology, Department of Physics(麻省理工学院物理系)
;
Massachusetts Institute of Technology, Institute for Data, Systems, and Society(麻省理工学院数据、系统与社会研究所)
;
NSF AI Institute for Artificial Intelligence and Fundamental Interactions(国家科学基金会人工智能与基本相互作用AI研究所)
Efficient Onboard Vision-Language Inference in UAV-Enabled Low-Altitude Economy Networks via LLM-Enhanced Optimization
基于LLM增强优化的无人机低空经济网络高效机载视觉-语言推理
Yang Li, Ruichen Zhang, Yinqiu Liu, Guangyuan Liu, Abbas Jamalipour, Xianbin Wang, Dong In Kim
机构
*
College of Computing and Data Science, Nanyang Technological University, Singapore(计算与数据科学学院、新加坡国立科技大学)
;
The University of Sydney, Sydney, Australia(悉尼大学、澳大利亚悉尼)
;
Department of Electrical and Computer Engineering, Western University, London, Canada(电气与计算机工程系、西方大学、加拿大伦敦)
;
Department of Electrical and Computer Engineering, Sungkyunkwan University, South Korea(电气与计算机工程系、全州大学、韩国)
CL-CLIP: CLIP-Based Continual Learning Framework with Cost-Volume Category Decoupling for Object Detection
CL-CLIP: 基于CLIP的持续学习框架与代价体积类别解耦用于目标检测
Zihan Liu, Yuguang Yang, Shengjie Su, Jianing Pang, Linlin Yang, Chunyu Xie, Nikolai Yu. Zolotykh, Baochang Zhang
机构
*
National College for Excellent Engineers, Beihang University(卓越工程师学院,北京航空航天大学)
;
AI Research, Qihoo 360(360人工智能研究院,奇虎360)
;
School of Electronic Information Engineering, Beihang University(电子信息学院,北京航空航天大学)
;
School of Cyber Science and Technology, Beihang University(网络安全科学与技术学院,北京航空航天大学)
;
School of Computer Science and Engineering, Beihang University(计算机科学与工程学院,北京航空航天大学)
;
State Key Laboratory of Media Convergence and Communication, Communication University of China(媒体融合与传播国家重点实验室,中国传媒大学)
;
Institute of Information Technology, Mathematics and Mechanics, Lobachebsky University(信息技术、数学与力学学院,洛瓦茨基大学)
;
School of Artificial Intelligence, Beihang University(人工智能学院,北京航空航天大学)
RAPTOR+: A Visually Grounded Vision-Language Framework to Improve Clinical Trust and Auditability in Automated Cancer Referral Processing
RAPTOR+: 一种基于视觉的视觉-语言框架,用于提高自动化癌症转诊处理中的临床信任度和可审计性
Sofiat Abioye, Ufaq Khan, Shazad Ashraf, Anusha Jose, Adam Byfield, Lukman Akanbi, Muhammad Bilal
机构
*
Birmingham City University(伯明翰城市大学)
;
Mohamed bin Zayed University of Artificial Intelligence (MBZUAI)(穆罕默德·本·扎耶德人工智能大学)
;
University Hospitals Birmingham NHS Foundation Trust(伯明翰大学医院 NHS 基础信托)
;
NHS England(英格兰国家卫生服务体系)
CommentsThis manuscript has been withdrawn by the authors. It reproduced the methodology of Gardinazzi et al., arXiv:2410.11042, without citation, and utilized code and data from the associated repository (github.com/RitAreaSciencePark/ZigZagLLMs) without disclosure or violate the MIT License. A revised future version with full attribution may be prepared. For any feedback, please contact Pengcheng Zheng
VOLD: Reasoning Transfer from LLMs to Vision-Language Models via On-Policy Distillation
VOLD:通过在线蒸馏将LLM推理能力转移到视觉语言模型
Walid Bousselham, Hilde Kuehne, Cordelia Schmid
机构
*
Tuebingen AI Center(图宾根人工智能中心)
;
University of Tuebingen(图宾根大学)
;
MIT-IBM Watson AI Lab(MIT-IBM沃森人工智能实验室)
;
Inria, École Normale Supérieure, CNRS, PSL Research University(法国国家科学研究院、巴黎-萨克勒大学、École Normale Supérieure、PSL研究大学)
Hierarchical Semantic-Augmented Navigation: Optimal Transport and Graph-Driven Reasoning for Vision-Language Navigation
层级语义增强导航:面向视觉语言导航的最优传输与图驱动推理
Xiang Fang, Wanlong Fang, Changshuo Wang
机构
*
School of Software Engineering, Huazhong University of Science and Technology(华中科技大学软件学院)
;
Interdisciplinary Graduate Programme, Nanyang Technological University, Singapore(新加坡南洋理工大学交叉学科研究生项目)
;
University College London(伦敦大学学院)
机构
*
Indian Institute of Technology Delhi, India(印度理工学院德里分校)
;
NVIDIA AI Technology Center, India(NVIDIA AI技术中心)
;
Jawaharlal Nehru University, India(贾瓦哈拉尔·尼赫鲁大学)