DRISHTIKON: Visual Grounding at Multiple Granularities in Documents
专题命中 视觉定位与Grounding :grounding(title,abstract);vision-language model(abstract);visual question answering(abstract);分类 cs.CV
Comments Work in Progress
AI 大模型
视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。
专题命中 视觉定位与Grounding :grounding(title,abstract);vision-language model(abstract);visual question answering(abstract);分类 cs.CV
Comments Work in Progress
机构 * School of EIC, Huazhong University of Science & Technology(华中科技大学电子信息学院) ; vivo AI Lab(vivo人工智能实验室)
专题命中 视觉定位与Grounding :grounding(title,abstract);vision-language model(abstract);VLM(abstract);分类 cs.CV
Comments To appear at ICCV 2025. Code: https://github.com/hustvl/GroundingSuite
机构 * School of Software Engineering, South China University of Technology(软件工程学院,华南理工大学) ; School of Future Technology, South China University of Technology(未来技术学院,华南理工大学) ; Shien-Ming Wu School of Intelligent Engineering, South China University of Technology(智能工程学院,华南理工大学)
专题命中 视觉定位与Grounding :vision-language model(title,abstract);VLM(abstract);grounding(abstract);分类 cs.AI
Comments submitted to NeurIPS 2025
机构 * University of Oxford(牛津大学) ; Zhejiang University(浙江大学) ; National University of Singapore(新加坡国立大学)
专题命中 视觉定位与Grounding :grounding(title,abstract);visual question answering(abstract);multimodal large language model(abstract);分类 cs.CV
机构 * University of Electronic Science and Technology of China(电子科技大学) ; Tongji University(同济大学)
专题命中 视觉定位与Grounding :MLLM(title,abstract);grounding(abstract);multimodal large language model(abstract);分类 cs.CV
Comments Accepted by TPAMI 2025
机构 * University of Chinese Academy of Sciences (UCAS)(中国科学院大学) ; New Laboratory of Pattern Recognition (NLPR), CASIA(中国科学院模式识别新技术实验室) ; State Key Laboratory of Multimodal Artificial Intelligence Systems (MAIS), CASIA(中国科学院多模态人工智能系统国家重点实验室) ; Hong Kong Institute of Science & Innovation, CASIA(香港科学与创新研究院) ; The Hong Kong Polytechnic University(香港理工大学) ; Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
专题命中 视觉定位与Grounding :grounding(title,abstract);vision-language model(abstract);VLM(abstract);分类 cs.CV
Comments Accepted to ACL 2025 Main
机构 * UCL(伦敦大学学院) ; Technical University of Munich(慕尼黑技术大学) ; University of Oxford(牛津大学) ; University of Glasgow(格拉斯哥大学)
专题命中 视觉定位与Grounding :multimodal large language model(title,abstract);LLaVA(abstract);grounding(abstract);分类 cs.CV
机构 * Mohamed bin Zayed University of Artificial Intelligence (MBZUAI)(莫扎德·本·扎耶德人工智能大学)
专题命中 视觉定位与Grounding :vision-language model(title,abstract);VLM(abstract);LLaVA(abstract);分类 cs.CV
Comments 14 pages, 5 figures, 3 tables
机构 * Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院) ; AMAP, Alibaba Group(阿里云研究院)
专题命中 视觉定位与Grounding :grounding(title,abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV
机构 * Beijing University of Posts and Telecommunications(北京邮电大学) ; Tianjin University(天津大学) ; South China Hospital, Medical School, Shenzhen University(深圳大学医学院南方医院)
专题命中 视觉定位与Grounding :grounding(title,abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV
机构 * Department of Computer Science, George Mason University(计算机科学系,乔治·马歇尔大学)
专题命中 视觉定位与Grounding :grounding(title,abstract);vision-language model(abstract);visual question answering(abstract);分类 cs.CV
Comments Accepted at CVPR-W
机构 * Microsoft Research(微软研究院)
专题命中 视觉定位与Grounding :grounding(title,abstract);vision language model(abstract);visual question answering(abstract);分类 cs.CV
机构 * East China Normal University(东华大学)
专题命中 视觉定位与Grounding :multimodal large language model(title,abstract);grounding(abstract);MLLM(abstract);分类 cs.CV
专题命中 视觉定位与Grounding :MLLM(title,abstract);VLM(abstract);multimodal large language model(abstract);分类 cs.CV
专题命中 视觉定位与Grounding :grounding(title,abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV
专题命中 视觉定位与Grounding :vision-language model(title,abstract);VLM(abstract);grounding(abstract);分类 cs.AI
专题命中 视觉定位与Grounding :grounding(title,abstract);vision-language model(abstract);VLM(abstract);分类 cs.CV
Comments 9 pages, 5 figures
专题命中 视觉定位与Grounding :vision-language model(title,abstract);visual reasoning(abstract);grounding(abstract);分类 cs.CV
专题命中 视觉定位与Grounding :vision-language model(title,abstract);vision language model(abstract);VLM(abstract);分类 cs.CV
Comments Codes are available at https://github.com/wkfdb/MarvelOVD
专题命中 视觉定位与Grounding :vision-language model(title,abstract);LLaVA(abstract);grounding(abstract);分类 cs.CV
专题命中 视觉定位与Grounding :multimodal large language model(title,abstract);LLaVA(abstract);MLLM(abstract);分类 cs.CV
专题命中 视觉定位与Grounding :multimodal large language model(title,abstract);grounding(abstract);MLLM(abstract);分类 cs.CV
专题命中 视觉定位与Grounding :vision-language model(title,abstract);VLM(abstract);grounding(abstract);分类 cs.AI
Comments 9 pages, 7 figures. Presented at ICML 2023
专题命中 视觉定位与Grounding :vision language model(title,abstract);VLM(abstract);grounding(abstract);分类 cs.CV
Comments Accepted to ICLR2023
专题命中 视觉定位与Grounding :vision-language model(title,abstract);VLM(abstract,comments);分类 cs.CV、cs.AI、cs.LG
Comments ICCV 2025 VLM 3d Workshop
机构 * The University of Tokyo(东京大学) ; S-Lab, Nanyang Technological University(南洋理工大学S实验室) ; Duke University(杜克大学) ; Salesforce AI Research(Salesforce人工智能研究) ; Nanyang Technological University(南洋理工大学) ; University of Wisconsin–Madison(威斯康星大学麦迪逊分校) ; Tokyo University of Science(东京科学大学)
专题命中 视觉定位与Grounding :vision language model(title,abstract);VLM(abstract,comments);分类 cs.CV、cs.AI、cs.LG
Comments Accepted at TMLR2025. Survey paper. We welcome questions, issues, and paper requests via https://github.com/AtsuMiyai/Awesome-OOD-VLM
专题命中 视觉定位与Grounding :vision-language model(title,abstract);VLM(abstract,comments);分类 cs.CV、cs.AI、cs.LG
Comments NIPS 2024 Camera Ready, Codes are available at \url{https://github.com/YBZh/OpenOOD-VLM}
MedPlex:用于临床基础医学分割的深度视觉-语言协同适配
机构 * Wayne State University(韦恩州立大学) ; Henry Ford Health(亨利福特医疗集团) ; Institute for AI and Data Science Wayne State University(韦恩州立大学人工智能与数据科学研究院)
专题命中 视觉定位与Grounding :VLM(summary_cn,abstract);vision-language model(abstract);分类 cs.CV、cs.AI
AI总结 MedPlex是一种端到端VLM框架,通过双向融合和两级概念对齐,实现医学图像分割的视觉-语言协同适配,在CT、MR的多类医学分割任务中达到最优性能。
Comments Accepted By BMVC-2026
开放词汇BEV分割:基于3D感知几何约束的方法
机构 * KAIST AI(韩国科学技术院人工智能) ; NAVER LABS(NAVER实验室)
专题命中 视觉定位与Grounding :VLM(summary_cn,abstract);vision-language model(abstract);分类 cs.CV、cs.LG
AI总结 针对开放词汇BEV分割中2D VLM语义到BEV的3D几何不一致问题,提出OVBEVSeg框架,通过三阶段几何约束(伪标签、场景优化、知识蒸馏)实现高效在线推理,在nuScenes上未见类别mIoU超越闭集方法15.3,推理速度提升2.5倍。
Comments This paper has been accepted to ECCV 2026
基于自主虚拟现实的危险环境态势感知风险检测
机构 * Interactive Robotics and Language Lab, University of Maryland Baltimore County(马里兰大学巴尔的摩县分校交互式机器人与语言实验室) ; DEVCOM Army Research Laboratory(陆军研究实验室)
专题命中 视觉定位与Grounding :VLM(summary_cn,abstract);vision language model(abstract);分类 cs.CV、cs.AI
AI总结 研究在高风险环境中,提出基于VR的人机交互框架,利用VLM辅助机器人识别潜在危险并标注兴趣点,通过VR界面呈现给操作员,经实验验证该方法能有效支持态势感知,提升交互体验。
Comments 7 Pages, Accepted to RO-MAN 2026