GUI-G$^2$: Gaussian Reward Modeling for GUI Grounding
机构 * Zhejiang University(浙江大学) ; Ant Group(蚂蚁集团)
专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV、cs.AI、cs.LG
AI 大模型
视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。
机构 * Zhejiang University(浙江大学) ; Ant Group(蚂蚁集团)
专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV、cs.AI、cs.LG
机构 * Shandong University(山东大学) ; Xi’an Jiaotong University(西安交通大学)
专题命中 视觉定位与Grounding :multimodal large language model(title,abstract);MLLM(abstract)
机构 * Robotics Institute, Carnegie Mellon University(卡内基梅隆大学机器人研究所) ; Department of Electrical and Computer Engineering, Carnegie Mellon University(卡内基梅隆大学电气与计算机工程系) ; College of Aeronautics & Engineering, Kent State University(肯特州立大学航空与工程学院)
专题命中 视觉定位与Grounding :VLM(title,abstract);vision-language model(abstract)
机构 * Computer Science and Artificial Intelligence Laboratory (CSAIL) at the Massachusetts Institute of Technology (MIT)(麻省理工学院计算机科学与人工智能实验室)
专题命中 视觉定位与Grounding :multimodal large language model(title,abstract);MLLM(abstract)
机构 * SKL-IOTSC, CIS, University of Macau(SKL-IOTSC、CIS、澳门大学)
专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV、cs.AI、cs.LG
Comments 16 pages
机构 * Queen Mary University of London(伦敦玛丽女王大学) ; Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) ; Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) ; The Hong Kong University of Science and Technology (GZ)(香港科学与技术大学)
专题命中 视觉定位与Grounding :MLLM(title,abstract);multimodal large language model(abstract)
机构 * School of Communication Engineering, Jilin University(吉林大学通信工程学院) ; College of Software Engineering, Jilin University(吉林大学软件工程学院) ; School of Mathematics, Jilin University(吉林大学数学学院) ; College of Electronic Science and Engineering, Jilin University(吉林大学电子科学工程学院)
专题命中 视觉定位与Grounding :VLM(title,abstract);vision-language model(abstract)
Comments 8 pages, 5 figures
机构 * Nanyang Technological University(南洋理工大学) ; Nanjing University of Science and Technology(南京理工大学)
专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV、cs.AI、cs.LG
机构 * Volkswagen AG(大众汽车集团) ; Technical University Berlin(柏林工业大学) ; University of Siegen(锡根大学)
专题命中 视觉定位与Grounding :vision language model(title,abstract);分类 cs.CV、cs.AI、cs.LG
专题命中 视觉定位与Grounding :vision-language model(title,abstract);VLM(abstract)
机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV、cs.AI、cs.LG
Comments Accepted to CVPR 2025. Project page: https://plan-lab.github.io/calico/
机构 * The University of Iowa(爱荷华大学) ; Max Planck Institute for Intelligent Systems(马克斯·普朗克智能系统研究所) ; Auburn University(奥本大学)
专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV、cs.AI、cs.LG
Comments Accepted in CVPR 2025
机构 * University Sorbonne Paris Nord(巴黎北索邦大学) ; University Paris-Saclay(巴黎-萨克雷大学) ; UVSQ(凡尔赛大学)
专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV、cs.AI、cs.LG
机构 * Goethe University Frankfurt(法兰克福大学) ; Tuebingen AI Center(蒂宾根人工智能中心) ; University of Tuebingen(蒂宾根大学) ; MPI for Informatics(马克斯·普朗克信息学研究所) ; MIT-IBM Watson AI Lab(麻省理工学院-IBM沃森人工智能实验室)
专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV、cs.AI、cs.LG
专题命中 视觉定位与Grounding :visual language model(title,abstract);VLM(abstract)
Journal ref Computer Graphics forum Volume 44 (2025), Number 3
机构 * University of Michigan(密歇根大学) ; New York University(纽约大学)
专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV、cs.AI、cs.LG
Comments CVPR 2025. Project website: https://3d-grand.github.io
专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV、cs.AI、cs.LG
专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV、cs.AI、cs.LG
专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV、cs.AI、cs.LG
Comments To be published in the Proceedings of AAAI 2025. The first three authors contributed equally. Project: https://github.com/DirtyHarryLYL/HAKE-AVA
专题命中 视觉定位与Grounding :grounding(title,abstract);visual question answering(abstract)
Comments Accepted to NAACL 2025 Main
专题命中 视觉定位与Grounding :grounding(title,abstract);multimodal large language model(abstract)
Comments Accepted for ICRA 2025. Project page: https://sites.google.com/umn.edu/etog-etrg/home
专题命中 视觉定位与Grounding :vision language model(title,abstract);VLM(abstract)
Comments Accepted to COLING2025
专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV、cs.AI、cs.LG
专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV、cs.AI、cs.LG
Comments NeurIPS 2024 Datasets and Benchmarks Track
专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV、cs.AI、cs.LG
Comments NeurIPS 2024
专题命中 视觉定位与Grounding :vision-language model(title,abstract);VLM(abstract)
Comments 17 pages, 9 figures
专题命中 视觉定位与Grounding :VLM(title,abstract);vision-language model(abstract)
专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV、cs.AI、cs.LG
Comments ECCV2024; Codes and Supp. are available at: https://github.com/YBZh/LAPT
专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV、cs.AI、cs.LG
Comments Accepted at ECCV 2024
专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV、cs.AI、cs.LG