DepthLM: Metric Depth From Vision Language Models
机构 * Meta ; Princeton University(普林斯顿大学)
专题命中 VLM训练与架构 :vision language model(title,abstract);VLM(abstract);分类 cs.CV
AI 大模型
视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。
机构 * Meta ; Princeton University(普林斯顿大学)
专题命中 VLM训练与架构 :vision language model(title,abstract);VLM(abstract);分类 cs.CV
机构 * CUNY Graduate Center(纽约大学研究生中心) ; Borough of Manhattan Community College(曼哈顿社区学院) ; The City College of New York(纽约城市学院)
专题命中 VLM训练与架构 :vision-language model(title);vision language model(abstract);LLaVA(abstract);分类 cs.CV
机构 * School of Computing, National University of Singapore(计算学院,新加坡国立大学)
专题命中 VLM训练与架构 :VLM(title,abstract);vision-language model(abstract);分类 cs.CV
机构 * Graduate School of Artificial Intelligence, POSTECH(POSTECH人工智能研究生院) ; Department of Computer Science and Engineering, POSTECH(POSTECH计算机科学与工程系) ; Australian National University(澳大利亚国立大学)
专题命中 VLM训练与架构 :VLM(title,abstract);vision-language model(abstract);分类 cs.CV
Comments Accepted to EMNLP-Findings 2025
机构 * Gianmarco Spinaci Department of Classical Philology and Italian Studies, University of Bologna, Italy Villa i Tatti, The Harvard University Center for Italian Renaissance Studies, Florence, Italy(Gianmarco Spinaci 文艺复兴研究系,博洛尼亚大学,意大利 塔蒂别墅,哈佛大学意大利文艺复兴研究中心,佛罗伦萨,意大利) ; Lukas Klic Villa i Tatti, The Harvard University Center for Italian Renaissance Studies, Florence, Italy(Lukas Klic 塔蒂别墅,哈佛大学意大利文艺复兴研究中心,佛罗伦萨,意大利) ; Giovanni Colavizza Department of Classical Philology and Italian Studies, University of Bologna, Italy Department of Communication, University of Copenhagen, Denmark(Giovanni Colavizza 文艺复兴研究系,博洛尼亚大学,意大利 传播系,哥本哈根大学,丹麦)
专题命中 VLM训练与架构 :multimodal large language model(title,abstract);vision language model(abstract);分类 cs.CV
Comments 11 pages, 2 figures
机构 * MBZUAI ; Institute of Automation, Chinese Academy of Sciences(自动化研究所,中国科学院)
专题命中 VLM训练与架构 :vision language model(title,abstract);InternVL(abstract);分类 cs.CV
机构 * University of Waterloo(滑铁卢大学)
专题命中 VLM训练与架构 :multimodal large language model(title);visual reasoning(abstract);MLLM(abstract);分类 cs.CV
Comments To appear at EMNLP 2025
机构 * School of Data Science, University of Virginia(数据科学学院,弗吉尼亚大学)
专题命中 VLM训练与架构 :vision language model(title,abstract);VLM(abstract);分类 cs.CV
机构 * State Key Lab of CAD&CG, Zhejiang University(浙江大学CAD与CG国家重点实验室) ; Laboratory of Art and Archaeology Image (Zhejiang University), Ministry of Education, China(浙江大学艺术与考古图像实验室(教育部,中国)) ; Zhejiang University of Finance&Economics(浙江财经大学) ; Zhejiang University(浙江大学)
专题命中 VLM训练与架构 :MLLM(title,abstract);multimodal large language model(abstract);分类 cs.CV
机构 * ZERON Shanghai, China(上海零点科技有限公司)
专题命中 VLM训练与架构 :vision language model(title,abstract);VLM(abstract);分类 cs.CV
Comments 2nd place in CVPR 2024 End-to-End Driving at Scale Challenge
机构 * Xidian University(西电大学)
专题命中 VLM训练与架构 :vision-language model(title,abstract);LLaVA(abstract);分类 cs.CV
机构 * Kempner Institute for the Study of Natural and Artificial Intelligence at Harvard University(哈佛大学自然与人工智能研究所) ; Department of Computer Science, Harvard University(哈佛大学计算机科学系)
专题命中 VLM训练与架构 :vision-language model(title,abstract);VLM(abstract);分类 cs.CV
Comments COLM 2025
机构 * National Taiwan University(国立台湾大学) ; The University of Tokyo(东京大学) ; Reichman University(里奇曼大学)
专题命中 VLM训练与架构 :VLM(title,abstract);vision-language model(abstract);分类 cs.CV
Comments 11 pages, Hsiao-Yuan Chin and I-Chao Shen contributed equally to the paper
专题命中 VLM训练与架构 :vision-language model(title,abstract);VLM(abstract);分类 cs.CV
Comments accepted by EMNLP 2025
机构 * Eastern Institute of Technology, Ningbo, China(东部技术研究所) ; Westlake University(西湖大学) ; USTC(中国科学技术大学)
专题命中 VLM训练与架构 :MLLM(title,abstract);multimodal large language model(abstract);分类 cs.CV
专题命中 VLM训练与架构 :vision-language model(title,abstract);VLM(abstract);分类 cs.CV
专题命中 VLM训练与架构 :multimodal large language model(title,abstract);LLaVA(abstract);分类 cs.AI
机构 * College of Computer Science & Technology, Zhejiang University(浙江大学计算机科学与技术学院) ; Transvascular Implantation Devices Research Institute and Liangzhu Laboratory(血管植入物研究机构和良渚实验室) ; Ant Group(蚂蚁集团) ; University of Notre Dame(圣母大学) ; HKUST (Guangzhou)(香港科技大学(广州))
专题命中 VLM训练与架构 :MLLM(title);LLaVA(abstract);multimodal large language model(abstract);分类 cs.CV
Comments Accepted by ICCV 2025
机构 * State Key Laboratory of Multimodal Artificial Intelligence Systems(多模态人工智能系统国家重点实验室) ; Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) ; University of Chinese Academy of Sciences(中国科学院大学) ; University of Science and Technology Beijing(北京科技大学) ; Hunan University of Science and Technology(湖南科技大学) ; Shanxi University of Finance and Economics(山西财经大学) ; Beijing Friendship Hospital, Capital Medical University(首都医科大学北京友谊医院)
专题命中 VLM训练与架构 :vision-language model(title,abstract);VLM(abstract);分类 cs.CV
Comments Camera ready version for ICONIP 2025
专题命中 VLM训练与架构 :multimodal large language model(title,abstract);MLLM(abstract);分类 cs.CV
机构 * Friedrich-Alexander-Universität Erlangen-Nürnberg(弗里德里希-亚历山大-厄林根-纽伦堡大学) ; Ingolstadt University of Applied Sciences(因戈尔施塔特应用科学大学) ; Flensburg University of Applied Sciences(弗拉森堡应用科学大学) ; Julius-Maximilians-Universität Würzburg(朱利叶斯-马克斯-魏扎克大学)
专题命中 VLM训练与架构 :vision-language model(title,abstract);VLM(abstract);分类 cs.CV
机构 * School of Architecture, Tsinghua University(清华大学建筑学院) ; Department of Electronic Engineering, Tsinghua University(清华大学电子工程系) ; BNRist, Tsinghua University(清华大学BNRist)
专题命中 VLM训练与架构 :visual language model(title,abstract);VLM(abstract);分类 cs.CV
Journal ref Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics. Volume 1: Long Papers (2025) 11571-11590
机构 * King’s College London(伦敦国王学院) ; Imperial College London(帝国理工学院) ; Columbia University(哥伦比亚大学)
专题命中 VLM训练与架构 :multimodal large language model(title,abstract);MLLM(abstract);分类 cs.AI
Comments Submitted to the NeurIPS 2025 Workshop GenAI4Health. Conference website: https://aihealth.ischool.utexas.edu/GenAI4HealthNeurips2025/
机构 * NEC Laboratories, America(美国 NEC 实验室)
专题命中 VLM训练与架构 :vision-language model(title,abstract);VLM(abstract);分类 cs.CV
Comments Accepted to ICCV 2025 Workshop (4th DataCV Workshop and Challenge)
机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
专题命中 VLM训练与架构 :vision-language model(title,abstract);VLM(abstract);分类 cs.CV
Comments This paper has been accepted by ACM Multimedia 2025 (ACM MM 2025)
机构 * University of Science and Technology of China(中国科学技术大学) ; WeChat Vision, Tencent Inc.(腾讯公司)
专题命中 VLM训练与架构 :vision-language model(title,abstract);VLM(abstract);分类 cs.CV
机构 * Waymo ; Boston University(波士顿大学) ; Korea University(韩国大学)
专题命中 VLM训练与架构 :VLM(title,abstract);vision-language model(abstract);分类 cs.CV
专题命中 VLM训练与架构 :vision-language model(title,abstract);VLM(abstract);分类 cs.CV
Comments ICCV 2025
机构 * Technical University of Munich, Germany(慕尼黑技术大学,德国) ; Helmholtz Munich, Munich Center for Machine Learning, Germany(海德堡慕尼黑,慕尼黑机器学习中心,德国) ; University of Tübingen, Tübingen AI Center, Germany(图宾根大学,图宾根人工智能中心,德国) ; University of Trento, Italy(特伦托大学,意大利) ; Beijing University of Posts and Telecommunications, China(北京邮电大学,中国)
专题命中 VLM训练与架构 :multimodal large language model(title,abstract);LLaVA(abstract);分类 cs.CV
Comments Accepted at GCPR 2025
专题命中 VLM训练与架构 :vision-language model(title,abstract);VLM(abstract);分类 cs.CV
Comments Accepted by ICCV 2025