arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-10-09 至 2025-10-09 共收录 24 信号源:cs.CV, cs.AI, cs.LG

1. 视觉问答 4 篇

2412.09278 2025-10-09 cs.CV cs.AI 88%

Towards a Multimodal Large Language Model with Pixel-Level Insight for Biomedicine

Xiaoshuang Huang, Lingdong Shen, Jia Liu, Fangxin Shang, Hongxiang Li, Haifeng Huang, Yehui Yang

机构 * Baidu Inc.(百度公司)

专题命中 视觉问答 :multimodal large language model(title,abstract);visual question answering(abstract);grounding(abstract);MLLM(abstract)

Comments Accepted by AAAI2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07098 2025-10-09 cs.CL 75%

TALENT: Table VQA via Augmented Language-Enhanced Natural-text Transcription

Guo Yutong, Wanying Wang, Yue Wu, Zichen Miao, Haoyu Wang

机构 * Johns Hopkins University(约翰霍普金斯大学) Purdue University(普渡大学) University at Albany(阿尔巴尼大学)

专题命中 视觉问答 :vision-language model(abstract);VLM(abstract);visual question answering(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06240 2025-10-09 cs.CL cs.AI cs.DB 57%

Knowledge Graph-Guided Multi-Agent Distillation for Reliable Industrial Question Answering with Datasets

Jiqun Pan, Zhenke Duan, Jiani Tu, Anzhi Cheng, Yanqing Wang

机构 * Zhongnan University of Economics and Law(中南财经政法大学)

专题命中 视觉问答 :grounding(abstract);分类 cs.AI

Comments 41 pages, 12 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04022 2025-10-09 cs.CV 57%

Video-in-the-Loop: Span-Grounded Long Video QA with Interleaved Reasoning

Chendong Wang, Donglin Bai, Yifan Yang, Xiao Jin, Anlan Zhang, Rui Wang, Shiqi Jiang, Yuqing Yang, Hao Wu, Qi Dai, Chong Luo, Ting Cao, Lili Qiu, Suman Banerjee

机构 * University of Wisconsin–Madison(威斯康星大学麦迪逊分校) Microsoft Research(微软研究院) Columbia University(哥伦比亚大学) University of Southern California(南加州大学) Fudan University(复旦大学)

专题命中 视觉问答 :grounding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 视觉推理 2 篇

2408.12009 2025-10-09 cs.CV 77%

CaRDiff: Video Salient Object Ranking Chain of Thought Reasoning for Saliency Prediction with Diffusion

Yolo Yunlong Tang, Gen Zhan, Li Yang, Yiting Liao, Chenliang Xu

机构 * ByteDance(字节跳动)

专题命中 视觉推理 :grounding(abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV

Comments Accepted to AAAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23867 2025-10-09 cs.CL cs.AI 57%

InfiMed: Low-Resource Medical MLLMs with Advancing Understanding and Reasoning

Zeyu Liu, Zhitian Hou, Guanghao Zhu, Zhijie Sang, Congkai Xie, Hongxia Yang

机构 * The Hong Kong Polytechnic University(香港理工大学) Sun Yat-sen University(中山大学)

专题命中 视觉推理 :multimodal large language model(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视觉定位与Grounding 6 篇

2504.07836 2025-10-09 cs.CV cs.AI 81%

AerialVG: A Challenging Benchmark for Aerial Visual Grounding by Exploring Positional Relations

Junli Liu, Qizhi Chen, Zhigang Wang, Yiwen Tang, Yiting Zhang, Chi Yan, Dong Wang, Xuelong Li, Bin Zhao

机构 * Northwestern Polytechnical University(西北工业大学) Shanghai AI Laboratory(上海人工智能实验室) Zhejiang University(浙江大学) TeleAI

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV、cs.AI

Comments 8 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03363 2025-10-09 cs.CV cs.AI eess.IV 62%

Unified Unsupervised Anomaly Detection via Matching Cost Filtering

Zhe Zhang, Mingxiu Cai, Gaochang Wu, Jing Zhang, Lingqiao Liu, Dacheng Tao, Tianyou Chai, Xiatian Zhu

机构 * State Key Laboratory of Synthetical Automation for Process Industries, Northeastern University, Shenyang, China(合成过程工业综合自动化国家重点实验室,东北大学,沈阳,中国) University of Surrey(Surrey大学) School of Computer Science, Wuhan University(武汉大学计算机学院) School of Computer Science, The University of Adelaide(阿德莱德大学计算机学院) College of Computing & Data Science, Nanyang Technological University(南洋理工大学计算机与数据科学学院) Surrey Institute for People-Centred Artificial Intelligence, and Centre for Vision, Speech and Signal Processing, University of Surrey(Surrey人本人工智能研究所,以及视觉、语音和信号处理中心,Surrey大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.AI

Comments 63 pages (main paper and supplementary material), 39 figures, 58 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06512 2025-10-09 cs.CV cs.AI 62%

LogSTOP: Temporal Scores over Prediction Sequences for Matching and Retrieval

Avishree Khare, Hideki Okamoto, Bardh Hoxha, Georgios Fainekos, Rajeev Alur

机构 * seas.upenn.edu(宾夕法尼亚大学塞as学院)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07000 2025-10-09 cs.CL cs.AI 57%

Pragyaan: Designing and Curating High-Quality Cultural Post-Training Datasets for Indian Languages

Neel Prabhanjan Rachamalla, Aravind Konakalla, Gautam Rajeev, Ashish Kulkarni, Chandra Khatri, Shubham Agarwal

机构 * Krutrim AI

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06411 2025-10-09 cs.CL 50%

Instructional Goal-Aligned Question Generation for Student Evaluation in Virtual Lab Settings: How Closely Do LLMs Actually Align?

R. Alexander Knipper, Indrani Dey, Souvika Sarkar, Hari Narayanan, Sadhana Puntambekar, Santu Karmaker

机构 * Department of EdPsych, University of Wisconsin-Madison(威斯康星大学麦迪逊分校教育心理学系) Department of CS, Wichita State University(威斯康星州立大学Wichita分校计算机科学系) Department of CSSE, Auburn University(阿伯茨罕大学计算机科学与工程系)

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06354 2025-10-09 cs.CL 50%

LLM Bias Detection and Mitigation through the Lens of Desired Distributions

Ingroj Shrestha, Padmini Srinivasan

机构 * University of Iowa(爱荷华大学)

专题命中 视觉定位与Grounding :grounding(abstract)

Comments Accepted to EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 文档图表理解 1 篇

2503.16096 2025-10-09 cs.CV 57%

MarkushGrapher: Joint Visual and Textual Recognition of Markush Structures

Lucas Morin, Valéry Weber, Ahmed Nassar, Gerhard Ingmar Meijer, Luc Van Gool, Yawei Li, Peter Staar

专题命中 文档图表理解 :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

5. GUI与屏幕智能体 2 篇

2510.07077 2025-10-09 cs.RO cs.AI cs.CV cs.LG 67%

Vision-Language-Action Models for Robotics: A Review Towards Real-World Applications

Kento Kawaharazuka, Jihoon Oh, Jun Yamada, Ingmar Posner, Yuke Zhu

机构 * Department of Computer Science, The University of Texas at Austin(德克萨斯大学计算机科学系)

专题命中 GUI与屏幕智能体 :vision-language model(abstract);分类 cs.CV、cs.AI、cs.LG

Comments Accepted to IEEE Access, website: https://vla-survey.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13367 2025-10-09 cs.RO 50%

Uncertainty-Informed Active Perception for Open Vocabulary Object Goal Navigation

Utkarsh Bajpai, Julius Rückin, Cyrill Stachniss, Marija Popović

机构 * Center for Robotics, University of Bonn(波恩大学机器人中心) MAVLab, TU Delft(代尔夫特理工大学MAVLab) Lamarr Institute for Machine Learning and Artificial Intelligence(机器学习与人工智能拉马尔研究所)

专题命中 GUI与屏幕智能体 :vision-language model(abstract)

Comments 7 pages, 3 figures

Journal ref Proceedings of the 2025 European Conference on Mobile Robots (ECMR)

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 幻觉与鲁棒性 3 篇

2501.19017 2025-10-09 cs.CL 82%

Benchmarking Gaslighting Negation Attacks Against Multimodal Large Language Models

Bin Zhu, Yinxuan Gui, Huiyan Qi, Jingjing Chen, Chong-Wah Ngo, Ee-Peng Lim

机构 * Singapore Management University(新加坡管理大学) Fudan University(复旦大学)

专题命中 幻觉与鲁棒性 :multimodal large language model(title,abstract);LLaVA(abstract)

Comments Project website: https://yxg1005.github.io/GaslightingNegationAttacks/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06280 2025-10-09 cs.CY cs.AI cs.CV 81%

Surgeons Are Indian Males and Speech Therapists Are White Females: Auditing Biases in Vision-Language Models for Healthcare Professionals

Zohaib Hasan Siddiqui, Dayam Nadeem, Mohammad Masudur Rahman, Mohammad Nadeem, Shahab Saquib Sohail, Beenish Moalla Chaudhry

机构 * University of Louisiana at Lafayette(路易斯安那大学拉斐特分校) Aligarh Muslim University(阿尔瓦格穆斯林大学)

专题命中 幻觉与鲁棒性 :vision-language model(title);vision language model(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07035 2025-10-09 cs.LG cs.AI 62%

Unified Molecule Pre-training with Flexible 2D and 3D Modalities: Single and Paired Modality Integration

Tengwei Song, Min Wu, Yuan Fang

机构 * Computational Bioscience Research Center, King Abdullah University of Science and Technology(国王阿卜杜勒·阿齐兹科技大学计算生物科学研究中心) Institute for Infocomm Research, A*STAR(信息与通信研究机构,A*STAR) School of Computing and Information Systems, Singapore Management University(新加坡管理大学计算机与信息系统学院)

专题命中 幻觉与鲁棒性 :vision-language model(abstract);分类 cs.AI、cs.LG

Comments CIKM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

7. VLM训练与架构 4 篇

2510.06529 2025-10-09 cs.CV 70%

VUGEN: Visual Understanding priors for GENeration

Xiangyi Chen, Théophane Vallaeys, Maha Elbayad, John Nguyen, Jakob Verbeek

机构 * Meta Fundamental AI Research(Meta基础人工智能研究) FAIR at Meta(Meta的FAIR)

专题命中 VLM训练与架构 :vision-language model(abstract);VLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.12353 2025-10-09 cs.CV cs.AI 62%

V2Xum-LLM: Cross-Modal Video Summarization with Temporal Prompt Instruction Tuning

Hang Hua, Yolo Yunlong Tang, Chenliang Xu, Jiebo Luo

专题命中 VLM训练与架构 :vision-language model(abstract);分类 cs.CV、cs.AI

Comments Accepted to AAAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06235 2025-10-09 eess.IV cs.AI cs.CV q-bio.NC 62%

Stacked Regression using Off-the-shelf, Stimulus-tuned and Fine-tuned Neural Networks for Predicting fMRI Brain Responses to Movies (Algonauts 2025 Report)

Robert Scholz, Kunal Bagga, Christine Ahrends, Carlo Alberto Barbano

机构 * Université Paris Cité(巴黎Cité大学) University of Oxford(牛津大学) University of Turin(都灵大学) Universität Leipzig(莱比锡大学) Max Planck School of Cognition(马克斯·普朗克认知科学学院)

专题命中 VLM训练与架构 :vision-language model(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06679 2025-10-09 cs.CV 57%

DreamOmni2: Multimodal Instruction-based Editing and Generation

Bin Xia, Bohao Peng, Yuechen Zhang, Junjia Huang, Jiyang Liu, Jingyao Li, Haoru Tan, Sitong Wu, Chengyao Wang, Yitong Wang, Xinglong Wu, Bei Yu, Jiaya Jia

机构 * CUHK(香港中文大学) HKUST(香港科技大学) HKU(香港大学) ByteDance Inc(字节跳动公司)

专题命中 VLM训练与架构 :VLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

8. 其他VLM 2 篇

2505.20444 2025-10-09 cs.LG cs.CV 81%

HoPE: Hybrid of Position Embedding for Long Context Vision-Language Models

Haoran Li, Yingjie Qin, Baoyuan Ou, Lai Xu, Ruiwen Xu

机构 * Carnegie Mellon University(卡内基梅隆大学) Xiaohongshu Inc.(小红书公司)

专题命中 其他VLM :vision-language model(title,abstract);分类 cs.CV、cs.LG

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06664 2025-10-09 cs.CL 50%

ToolMem: Enhancing Multimodal Agents with Learnable Tool Capability Memory

Yunzhong Xiao, Yangmin Li, Hewei Wang, Yunlong Tang, Zora Zhiruo Wang

机构 * Carnegie Mellon University(卡内基梅隆大学) University of Rochester(罗切斯特大学)

专题命中 其他VLM :vision-language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏