arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-09-15 至 2025-09-15 共收录 5 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 5 篇

2503.16365 2025-09-15 cs.CV cs.AI 84%

JARVIS-VLA: Post-Training Large-Scale Vision Language Models to Play Visual Games with Keyboards and Mouse

Muyao Li, Zihao Wang, Kaichen He, Xiaojian Ma, Yitao Liang

机构 * Peking University(北京大学) BIGAI(大疆创新)

专题命中 视觉定位与Grounding :vision language model(title);visual language model(abstract);grounding(abstract);分类 cs.CV、cs.AI

Comments Accepted by ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.08336 2025-09-15 cs.CV 83%

Talk2PC: Enhancing 3D Visual Grounding through LiDAR and Radar Point Clouds Fusion for Autonomous Driving

Runwei Guan, Jianan Liu, Ningwei Ouyang, Shaofeng Liang, Daizong Liu, Xiaolou Sun, Lianqing Zheng, Ming Xu, Yutao Yue, Guoqiang Mao, Hui Xiong

机构 * Thrust of Artificial Intelligence, Hong Kong University of Science and Technology (Guangzhou), China(香港理工大学(广州)人工智能研究所) Mononai AI, Sweden(蒙诺人工智能) College of Computer Science and Technology, China University of Petroleum (East China)(中国石油大学(华东)计算机科学与技术学院) School of Advanced Technology, Xi’an Jiaotong-Liverpool University(西安交通大学利物浦大学先进科技学院) Wangxuan Institute of Computer Technology, Peking University(北京大学王轩计算机技术研究所) School of Automation, Southeast University(东南大学自动化学院) College of Science and Engineering, James Cook University(詹姆斯库克大学科学与工程学院) School of Transportation, Southeast University(东南大学交通运输学院)

专题命中 视觉定位与Grounding :grounding(title,abstract);vision-language model(abstract);分类 cs.CV

Comments 13 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10278 2025-09-15 cs.CV 79%

Detecting Text Manipulation in Images using Vision Language Models

Vidit Vidit, Pavel Korshunov, Amir Mohammadi, Christophe Ecabert, Ketan Kotwal, Sébastien Marcel

机构 * IDIAP Research Institute(IDIAP研究 institute)

专题命中 视觉定位与Grounding :vision language model(title,abstract);分类 cs.CV

Comments Accepted in Synthetic Realities and Biometric Security Workshop BMVC-2025. For paper page see https://www.idiap.ch/paper/textvlmdet/

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.04191 2025-09-15 cs.CV cs.RO 70%

GROVE: A Generalized Reward for Learning Open-Vocabulary Physical Skill

Jieming Cui, Tengyu Liu, Ziyu Meng, Jiale Yu, Ran Song, Wei Zhang, Yixin Zhu, Siyuan Huang

专题命中 视觉定位与Grounding :vision language model(abstract);VLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.13776 2025-09-15 cs.CL cs.AI 57%

Aggregation Artifacts in Subjective Tasks Collapse Large Language Models' Posteriors

Georgios Chochlakis, Alexandros Potamianos, Kristina Lerman, Shrikanth Narayanan

机构 * University of Southern California(南加州大学) National Technical University of Athens(雅典国立技术大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments 16 pages, 12 figures, 3 tables

Journal ref Proceedings of the 2025 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 5513-5528, Albuquerque, New Mexico, April 2025

详情

展开后加载摘要…

URL PDF HTML 收藏