arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-09-29 至 2025-09-29 共收录 9 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 9 篇

2509.21788 2025-09-29 cs.CV 83%

MIRG-RL: Multi-Image Reasoning and Grounding with Reinforcement Learning

Lihao Zheng, Jiawei Chen, Xintian Shen, Hao Ma, Tao Wei

机构 * Li Auto Inc.(利亚 Auto 公司)

专题命中 视觉定位与Grounding :grounding(title,abstract);visual language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.05271 2025-09-29 cs.CV 81%

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Zhe Chen, Weiyun Wang, Yue Cao, Yangzhou Liu, Zhangwei Gao, Erfei Cui, Jinguo Zhu, Shenglong Ye, Hao Tian, Zhaoyang Liu, Lixin Gu, Xuehui Wang, Qingyun Li, Yiming Ren, Zixuan Chen, Jiapeng Luo, Jiahao Wang, Tan Jiang, Bo Wang, Conghui He, Botian Shi, Xingcheng Zhang, Han Lv, Yi Wang, Wenqi Shao, Pei Chu, Zhongying Tu, Tong He, Zhiyong Wu, Huipeng Deng, Jiaye Ge, Kai Chen, Kaipeng Zhang, Limin Wang, Min Dou, Lewei Lu, Xizhou Zhu, Tong Lu, Dahua Lin, Yu Qiao, Jifeng Dai, Wenhai Wang

机构 * Shanghai AI Laboratory(上海人工智能实验室) SenseTime Research(商汤科技研究院) Tsinghua University(清华大学) Nanjing University(南京大学) Fudan University(复旦大学) The Chinese University of Hong Kong(香港中文大学) Shanghai Jiao Tong University(上海交通大学)

专题命中 视觉定位与Grounding :InternVL(abstract);grounding(abstract);multimodal large language model(abstract);MLLM(abstract)

Comments Technical Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.23768 2025-09-29 cs.CL cs.CV 79%

Texture or Semantics? Vision-Language Models Get Lost in Font Recognition

Zhecheng Li, Guoxian Song, Yujun Cai, Zhen Xiong, Junsong Yuan, Yiwei Wang

机构 * University of California, San Diego(加州大学圣地亚哥分校) ByteDance(字节跳动) The University of Queensland(昆士兰大学) University of Southern California(南加州大学) University at Buffalo(布法罗大学) University of California, Merced(加州大学默塞德分校)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV

Comments Accepted to COLM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21356 2025-09-29 cs.CV cs.AI 73%

Phrase-grounded Fact-checking for Automatically Generated Chest X-ray Reports

Razi Mahmood, Diego Machado-Reyes, Joy Wu, Parisa Kaviani, Ken C. L. Wong, Niharika D'Souza, Mannudeep Kalra, Ge Wang, Pingkun Yan, Tanveer Syeda-Mahmood

机构 * Rensselaer Polytechnic Institute, NY, USA(罗文学院) IBM Research, Almaden, CA, USA(IBM研究院) Stanford University, CA, USA(斯坦福大学) Massachusetts General Hospital (MGH), Boston, USA(麻省总医院)

专题命中 视觉定位与Grounding :vision language model(abstract);VLM(abstract);分类 cs.CV、cs.AI

Comments In proceedings MICCAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.03334 2025-09-29 cs.CV cs.DB 70%

OS-W2S: An Automatic Labeling Engine for Language-Guided Open-Set Aerial Object Detection

Guoting Wei, Yu Liu, Xia Yuan, Xizhe Xue, Linlin Guo, Yifan Yang, Chunxia Zhao, Zongwen Bai, Haokui Zhang, Rong Xiao

机构 * Nanjing University of Science and Technology(南京理工大学) Intellifusion Inc.(Intellifusion公司) Northwestern Polytechnical University(西北工业大学) Zhejiang Lab(浙江实验室) Yan’an University(延安大学) Beijing University of Posts and Telecommunications(北京邮电大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);grounding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.12881 2025-09-29 cs.CL 67%

TEXT2AFFORD: Probing Object Affordance Prediction abilities of Language Models solely from Text

Sayantan Adak, Daivik Agrawal, Animesh Mukherjee, Somak Aditya

机构 * IIT, Kharagpur(印度Kharagpur理工学院)

专题命中 视觉定位与Grounding :vision-language model(abstract);grounding(abstract)

Comments Accepted at Conference on Computational Natural Language Learning 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22019 2025-09-29 cs.CV 57%

EgoInstruct: An Egocentric Video Dataset of Face-to-face Instructional Interactions with Multi-modal LLM Benchmarking

Yuki Sakai, Ryosuke Furuta, Juichun Yen, Yoichi Sato

机构 * The University of Tokyo(东京大学)

专题命中 视觉定位与Grounding :multimodal large language model(abstract);分类 cs.CV

Comments Accepted to the I-HFM Workshop at ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21367 2025-09-29 cs.CR cs.AI 57%

Design and Implementation of a Secure RAG-Enhanced AI Chatbot for Smart Tourism Customer Service: Defending Against Prompt Injection Attacks -- A Case Study of Hsinchu, Taiwan

Yu-Kai Shih, You-Kai Kang

机构 * Department of Information Management, National Dong Hwa University(信息管理系,国立东华大学) By The Student (BTS) Experimental Education Program (Non-School Type)(学生(BTS)实验教育计划(非学校类型))

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments 12 pages, 7 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21978 2025-09-29 cs.CL 50%

MotivGraph-SoIQ: Integrating Motivational Knowledge Graphs and Socratic Dialogue for Enhanced LLM Ideation

Xinping Lei, Tong Zhou, Yubo Chen, Kang Liu, Jun Zhao

机构 * The Key Laboratory of Cognition and Decision Intelligence for Complex Systems(认知与决策智能复杂系统重点实验室) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Hunan Provincial Key Laboratory of Philosophy and Social Sciences of Artificial Intelligence and Precision International, Hunan Normal University(湖南省人工智能与精准哲学社会科学省级重点实验室,湖南师范大学) Beijing University of Posts and Telecommunications(北京邮电大学)

专题命中 视觉定位与Grounding :grounding(abstract)

Comments EMNLP2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏