arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-09-05 至 2025-09-05 共收录 27 信号源:cs.CV, cs.AI, cs.LG

1. 视觉问答 2 篇

2509.03805 2025-09-05 cs.CL cs.AI 77%

Measuring How (Not Just Whether) VLMs Build Common Ground

Saki Imai, Mert İnan, Anthony Sicilia, Malihe Alikhani

机构 * Northeastern University(东北大学)

专题命中 视觉问答 :vision language model(abstract);VLM(abstract);grounding(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04162 2025-09-05 cs.AR 67%

Real Time FPGA Based Transformers & VLMs for Vision Tasks: SOTA Designs and Optimizations

Safa Mohammed Sali, Mahmoud Meribout, Ashiyana Abdul Majeed

专题命中 视觉问答 :vision-language model(abstract);visual question answering(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 视觉推理 2 篇

2509.03837 2025-09-05 cs.LG cs.IT math.IT 79%

Vehicle-to-Infrastructure Collaborative Spatial Perception via Multimodal Large Language Models

Kimia Ehsani, Walid Saad

机构 * Bradley Department of Electrical and Computer Engineering(电气与计算机工程系)

专题命中 视觉推理 :multimodal large language model(title,abstract);分类 cs.LG

Comments Accepted at IEEE GLOBECOM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02175 2025-09-05 cs.CV cs.AI cs.CL cs.LG 67%

Understanding Space Is Rocket Science -- Only Top Reasoning Models Can Solve Spatial Understanding Tasks

Nils Hoehing, Mayug Maniparambil, Ellen Rushe, Noel E. O'Connor, Anthony Ventresque

机构 * School of Computer Science(计算机科学学院) University College Dublin(都柏林大学) School of Computing(计算机科学学院) Dublin City University(都柏林城市大学) School of Electronic Engineering(电子工程学院) Trinity College Dublin(都柏林三一学院)

专题命中 视觉推理 :VLM(abstract);分类 cs.CV、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视觉定位与Grounding 10 篇

2509.04243 2025-09-05 cs.CV cs.AI 84%

Learning Active Perception via Self-Evolving Preference Optimization for GUI Grounding

Wanfu Wang, Qipeng Huang, Guangquan Xue, Xiaobo Liang, Juntao Li

机构 * Wanfu Wang, Qipeng Huang, Guangquan Xue, Xiaobo Liang, Juntao Li(作者)

专题命中 视觉定位与Grounding :grounding(title,abstract);vision language model(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03800 2025-09-05 cs.CV 83%

MedVista3D: Vision-Language Modeling for Reducing Diagnostic Errors in 3D CT Disease Detection, Understanding and Reporting

Yuheng Li, Yenho Chen, Yuxiang Lai, Jike Zhong, Vanessa Wildman, Xiaofeng Yang

机构 * Department of Biomedical Engineering(生物医学工程系) Georgia Institute of Technology(佐治亚理工学院) Department of Machine Learning(机器学习系) Department of Radiation Oncology(放射肿瘤科) Emory University School of Medicine(埃默里大学医学院) University of Southern California(南加州大学)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);visual question answering(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14904 2025-09-05 cs.CV cs.AI 81%

TriCLIP-3D: A Unified Parameter-Efficient Framework for Tri-Modal 3D Visual Grounding based on CLIP

Fan Li, Zanyi Wang, Zeyi Huang, Guang Dai, Jingdong Wang, Mengmeng Wang

机构 * Xi’an Jiaotong University(西安交通大学) SGIT AI Lab(SGIT人工智能实验室) Zhejiang University of Technology(浙江工业大学) Huawei(华为)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03961 2025-09-05 cs.CV cs.AI 73%

Multimodal Feature Fusion Network with Text Difference Enhancement for Remote Sensing Change Detection

Yijun Zhou, Yikui Zhai, Zilu Ying, Tingfeng Xian, Wenlve Zhou, Zhiheng Zhou, Xiaolin Tian, Xudong Jia, Hongsheng Zhang, C. L. Philip Chen

机构 * College of Electronics and Information Engineering, Wuyi University(威怡大学电子与信息工程学院) School of Electronic and Information Engineering and the Key Laboratory of Big Data and Intelligent Robot, Ministry of Education, South China University of Technology(电子与信息工程学院和大数据与智能机器人重点实验室,华南理工大学) State Key Laboratory of Lunar and Planetary Sciences, Macau University of Science and Technology(澳门大学地球和行星科学国家重点实验室) College of Engineering and Computer Science, California State University, Northridge(工程与计算机科学学院,加州大学北岭分校) Department of Geography, The University of Hong Kong(地理系,香港大学) Faculty of Computer Science and Engineering, S(计算机科学与工程学院,S)

专题命中 视觉定位与Grounding :vision language model(abstract);VLM(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04326 2025-09-05 cs.CV 70%

Efficient Odd-One-Out Anomaly Detection

Silvio Chito, Paolo Rabino, Tatiana Tommasi

机构 * Politecnico di Torino(托斯尼亚理工学院)

专题命中 视觉定位与Grounding :visual reasoning(abstract);multimodal large language model(abstract);分类 cs.CV

Comments Accepted at ICIAP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04180 2025-09-05 cs.CV cs.AI 62%

VisioFirm: Cross-Platform AI-assisted Annotation Tool for Computer Vision

Safouane El Ghazouali, Umberto Michelucci

机构 * TOELT LLC AI lab(TOELT LLC人工智能实验室)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03793 2025-09-05 cs.MA cs.AI 57%

SAMVAD: A Multi-Agent System for Simulating Judicial Deliberation Dynamics in India

Prathamesh Devadiga, Omkaar Jayadev Shetty, Pooja Agarwal

机构 * PES University(PES大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03626 2025-09-05 cs.AI 57%

Explainable Knowledge Graph Retrieval-Augmented Generation (KG-RAG) with KG-SMILE

Zahra Zehtabi Sabeti Moghaddam, Zeinab Dehghani, Maneeha Rani, Koorosh Aslansefat, Bhupesh Kumar Mishra, Rameez Raja Kureshi, Dhavalkumar Thakker

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04254 2025-09-05 cs.HC 50%

MuMTAffect: A Multimodal Multitask Affective Framework for Personality and Emotion Recognition from Physiological Signals

Meisam Jamshidi Seikavandi, Fabricio Batista Narcizo, Ted Vucurevich, Andrew Burke Dittberner, Paolo Burelli

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.12065 2025-09-05 cs.CL cs.FL 50%

Autoformalization in the Wild: Assessing LLMs on Real-World Mathematical Definitions

Lan Zhang, Marco Valentino, Andre Freitas

机构 * Department of Computer Science, University of Manchester(曼彻斯特大学计算机科学系) School of Computer Science, University of Sheffield(谢菲尔德大学计算机科学学院) Idiap Research Institute(Idiap研究机构) National Biomarker Centre, CRUK Manchester Institute(国家生物标志物中心、CRUK曼彻斯特研究所)

专题命中 视觉定位与Grounding :grounding(abstract)

Comments EMNLP 2025 Camera-Ready Version

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 文档图表理解 1 篇

2509.03615 2025-09-05 cs.CL cs.AI 70%

E-ARMOR: Edge case Assessment and Review of Multilingual Optical Character Recognition

Aryan Gupta, Anupam Purwar

机构 * Sprinklr

专题命中 文档图表理解 :vision-language model(abstract);InternVL(abstract);分类 cs.AI

Comments Sprinklr OCR provides a fast and compute light way of performing OCR

详情

展开后加载摘要…

URL PDF HTML 收藏

5. GUI与屏幕智能体 1 篇

2509.03536 2025-09-05 cs.AI cs.HC 57%

PG-Agent: An Agent Powered by Page Graph

Weizhi Chen, Ziwei Wang, Leyang Yang, Sheng Zhou, Xiaoxuan Tang, Jiajun Bu, Yong Li, Wei Jiang

机构 * Zhejiang Key Lab of Accessible Perception \& Intelligent Systems, Zhejiang University Hangzhou China Zhejiang University

专题命中 GUI与屏幕智能体 :multimodal large language model(abstract);分类 cs.AI

Comments Paper accepted to ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 幻觉与鲁棒性 4 篇

2507.19346 2025-09-05 cs.LG 57%

Short-Form Video Recommendations with Multimodal Embeddings: Addressing Cold-Start and Bias Challenges

Andrii Dzhoha, Katya Mirylenka, Egor Malykh, Marco-Andrea Buchmann, Francesca Catino

机构 * Zalando SE Berlin Germany(泽尔安多德国分公司) Zalando Switzerland AG Zürich Switzerland(泽尔安多瑞士分公司)

专题命中 幻觉与鲁棒性 :vision-language model(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.15537 2025-09-05 cs.CV 57%

MUNBa: Machine Unlearning via Nash Bargaining

Jing Wu, Mehrtash Harandi

机构 * Department of Data Science & AI(数据科学与人工智能系) Department of Electrical and Computer Systems Engineering(电气与计算机系统工程系)

专题命中 幻觉与鲁棒性 :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04214 2025-09-05 cs.CR 50%

An Automated, Scalable Machine Learning Model Inversion Assessment Pipeline

Tyler Shumaker, Jessica Carpenter, David Saranchak, Nathaniel D. Bastian

专题命中 幻觉与鲁棒性 :vision language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03787 2025-09-05 cs.IR cs.CL 50%

Evaluating the Robustness of Retrieval-Augmented Generation to Adversarial Evidence in the Health Domain

Shakiba Amirshahi, Amin Bigdeli, Charles L. A. Clarke, Amira Ghenai

机构 * University of Waterloo(滑铁卢大学) Toronto Metropolitan University(多伦多 Metropolitan 大学)

专题命中 幻觉与鲁棒性 :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

7. VLM训练与架构 3 篇

2509.03895 2025-09-05 cs.CV 79%

Attn-Adapter: Attention Is All You Need for Online Few-shot Learner of Vision-Language Model

Phuoc-Nguyen Bui, Khanh-Binh Nguyen, Hyunseung Choo

机构 * Sungkyunkwan University(顺天大学) Deakin University(德金大学)

专题命中 VLM训练与架构 :vision-language model(title,abstract);分类 cs.CV

Comments ICCV 2025 - LIMIT Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03863 2025-09-05 cs.AI 70%

Expedition & Expansion: Leveraging Semantic Representations for Goal-Directed Exploration in Continuous Cellular Automata

Sina Khajehabdollahi, Gautier Hamon, Marko Cvjetko, Pierre-Yves Oudeyer, Clément Moulin-Frier, Cédric Colas

机构 * Flowers AI & CogSci Lab, Inria, France(Flowers AI与认知科学实验室,Inria,法国) MIT, USA(麻省理工学院,美国)

专题命中 VLM训练与架构 :vision-language model(abstract);VLM(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.10118 2025-09-05 cs.CV cs.AI 62%

Image Embedding Sampling Method for Diverse Captioning

Sania Waheed, Na Min An

机构 * University of Southhampton(南安普顿大学) KAIST(韩国科学技术院)

专题命中 VLM训练与架构 :VLM(abstract);分类 cs.CV、cs.AI

Comments 17 pages, 5 figures, 9 tables

详情

展开后加载摘要…

URL PDF HTML 收藏

8. 其他VLM 4 篇

2509.03893 2025-09-05 cs.CV 70%

Weakly-Supervised Learning of Dense Functional Correspondences

Stefan Stojanov, Linan Zhao, Yunzhi Zhang, Daniel L. K. Yamins, Jiajun Wu

机构 * Stanford University(斯坦福大学)

专题命中 其他VLM :vision-language model(abstract);vision language model(abstract);分类 cs.CV

Comments Accepted at ICCV 2025. Project website: https://dense-functional-correspondence.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02349 2025-09-05 cs.SD cs.AI cs.LG 62%

AudioCodecBench: A Comprehensive Benchmark for Audio Codec Evaluation

Lu Wang, Hao Chen, Siyu Wu, Zhiyue Wu, Hao Zhou, Chengfeng Zhang, Ting Wang, Haodi Zhang

机构 * Lu Wang(卢王) Hao Chen(何晨) Siyu Wu(武士) Zhiyue Wu(吴致岳) Hao Zhou(周浩) Chengfeng Zhang(张成峰) Ting Wang(王婷) Haodi Zhang(张浩迪)

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.12722 2025-09-05 cs.CV cs.AI cs.CR 62%

Defending LVLMs Against Vision Attacks through Partial-Perception Supervision

Qi Zhou, Tianlin Li, Qing Guo, Dongxia Wang, Yun Lin, Yang Liu, Jin Song Dong

机构 * College of Control Science and Engineering, Zhejiang University, China(控制科学与工程学院,浙江大学,中国) Huzhou Institute of Industrial Control Technology, China(湖州工业控制技术研究所,中国) School of Computer Science and Engineering, Nanyang Technological University, Singapore(计算机科学与工程学院,南洋理工大学,新加坡) School of Computer Science, Shanghai Jiao Tong University, China(计算机科学学院,上海交通大学,中国) School of Computing, National University of Singapore, Singapore(计算学院,新加坡国立大学,新加坡) Research, Singapore(研究,新加坡)

专题命中 其他VLM :vision language model(abstract);分类 cs.CV、cs.AI

Comments Accepted to ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18201 2025-09-05 cs.CL cs.CV cs.HC 57%

Deciphering Emotions in Children Storybooks: A Comparative Analysis of Multimodal LLMs in Educational Applications

Bushra Asseri, Estabraq Abdelaziz, Maha Al Mogren, Tayef Alhefdhi, Areej Al-Wabil

机构 * College of Engineering & Advanced Computing(工程与高级计算学院)

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏