arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-09-12 至 2025-09-12 共收录 30 信号源:cs.CV, cs.AI, cs.LG

1. 视觉问答 4 篇

2505.19455 2025-09-12 cs.CV cs.AI cs.LG 82%

MM-Prompt: Cross-Modal Prompt Tuning for Continual Visual Question Answering

Xu Li, Fan Lyu

机构 * Khoury College of Computer Sciences, Northeastern University(东北大学克劳尔计算机科学学院) New Laboratory of Pattern Recognition, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所模式识别新实验室)

专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.07084 2025-09-12 cs.RO 82%

DriveSOTIF: Advancing Perception SOTIF Through Multimodal Large Language Models

Shucheng Huang, Freda Shi, Chen Sun, Jiaming Zhong, Minghao Ning, Yufeng Yang, Yukun Lu, Hong Wang, Amir Khajepour

机构 * MVSLab, Department of Mechanical and Mechatronics Engineering, University of Waterloo(滑铁卢大学机械与机电工程系MVSLab) CompLING Lab, David R. Cheriton School of Computer Science, University of Waterloo(滑铁卢大学大卫·R·切里顿计算机科学学院CompLING Lab) Department of Data and Systems Engineering, University of Hong Kong(香港大学数据与系统工程系) Department of Mechanical Engineering, University of New Brunswick(新不伦瑞克大学机械工程系) School of Vehicle and Mobility, Tsinghua University(清华大学车辆与移动性学院)

专题命中 视觉问答 :multimodal large language model(title,abstract);MLLM(abstract)

Comments This work has been accepted to IEEE Transactions on Vehicular Technology. Please refer to the copyright notice for additional information

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09159 2025-09-12 cs.CV cs.AI 81%

A Knowledge Noise Mitigation Framework for Knowledge-based Visual Question Answering

Zhiyue Liu, Sihang Liu, Jinyuan Liu, Xinru Zhang

机构 * School of Computer, Electronics and Information(计算机、电子与信息学院) Guangxi University(广西大学) Guangxi Key Laboratory of Multimedia Communications and Network Technology(广西多媒体通信与网络技术重点实验室)

专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV、cs.AI

Comments Accepted by the IEEE International Conference on Multimedia and Expo (ICME 2025) for oral presentation. © 2025 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09254 2025-09-12 cs.CV cs.MM 70%

Towards Better Dental AI: A Multimodal Benchmark and Instruction Dataset for Panoramic X-ray Analysis

Jing Hao, Yuxuan Fan, Yanpeng Sun, Kaixin Guo, Lizhuo Lin, Jinrong Yang, Qi Yong H. Ai, Lun M. Wong, Hao Tang, Kuo Feng Hung

机构 * Faculty of Dentistry, The University of Hong Kong(香港大学牙科学院) The Hong Kong University of Science and Technology (GZ)(香港科学与技术大学) National University of Singapore(新加坡国立大学) CVTE Sun Yat-sen University(孙中山大学) Department of Diagnostic Radiology, The University of Hong Kong(香港大学放射科) Imaging and Interventional Radiology, Faculty of Medicine, The Chinese University of Hong Kong(香港中文大学医学院影像与介入放射科) School of Computer Science, Peking University(北京大学计算机科学系)

专题命中 视觉问答 :vision-language model(abstract);visual question answering(abstract);分类 cs.CV

Comments 40 pages, 26 figures, 9 tables

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 视觉推理 5 篇

2509.09013 2025-09-12 cs.CL cs.AI cs.CV 84%

Can Vision-Language Models Solve Visual Math Equations?

Monjoy Narayan Choudhury, Junling Wang, Yifan Hou, Mrinmaya Sachan

机构 * IIIT Bangalore(班加罗尔印度理工学院) ETH Zürich(苏黎世联邦理工学院)

专题命中 视觉推理 :vision-language model(title,abstract);visual reasoning(abstract);分类 cs.CV、cs.AI

Comments Monjoy Narayan Choudhury and Junling Wang contributed equally to this work. Accepted at EMNLP2025 main. Code and datasets are open-sourced with links in the paper

Journal ref Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.07815 2025-09-12 cs.RO cs.CV cs.LG 81%

Imagine, Verify, Execute: Memory-guided Agentic Exploration with Vision-Language Models

Seungjae Lee, Daniel Ekpo, Haowen Liu, Furong Huang, Abhinav Shrivastava, Jia-Bin Huang

专题命中 视觉推理 :vision-language model(title,abstract);分类 cs.CV、cs.LG

Comments Project webpage: https://ive-robot.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01106 2025-09-12 cs.AI cs.CV cs.RO 62%

Robix: A Unified Model for Robot Interaction, Reasoning and Planning

Huang Fang, Mengxi Zhang, Heng Dong, Wei Li, Zixuan Wang, Qifeng Zhang, Xueyun Tian, Yucheng Hu, Hang Li

专题命中 视觉推理 :grounding(abstract);分类 cs.CV、cs.AI

Comments Tech report. Project page: https://robix-seed.github.io/robix/

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09680 2025-09-12 cs.CV cs.CL 57%

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark

Rongyao Fang, Aldrich Yu, Chengqi Duan, Linjiang Huang, Shuai Bai, Yuxuan Cai, Kun Wang, Si Liu, Xihui Liu, Hongsheng Li

机构 * CUHK(香港中文大学) HKU(香港大学) BUAA(北京航空航天大学) Alibaba(阿里巴巴)

专题命中 视觉推理 :vision-language model(abstract);分类 cs.CV

Comments Project page: https://flux-reason-6m.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09263 2025-09-12 cs.CV 57%

DATE: Dynamic Absolute Time Enhancement for Long Video Understanding

Chao Yuan, Yang Yang, Yehui Yang, Zach Cheng

机构 * Beihang University(北京航空航天大学) Dcar, ByteDance(字节跳动Dcar部门) Qfin Holdings,Inc(Qfin控股公司) MAIS, Institute of Automation, Chinese Academy of Sciences(自动化研究所,中国科学院)

专题命中 视觉推理 :multimodal large language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视觉定位与Grounding 4 篇

2509.09584 2025-09-12 cs.CV cs.RO 79%

Visual Grounding from Event Cameras

Lingdong Kong, Dongyue Lu, Ao Liang, Rong Li, Yuhao Dong, Tianshuai Hu, Lai Xing Ng, Wei Tsang Ooi, Benoit R. Cottereau

机构 * NUS(新加坡国立大学) HKUST(GZ)(香港科技大学(广州)) NTU(南洋理工大学) HKUST(香港科技大学) I 2 R, A*STAR(新加坡科技研究局) IPAL, CNRS(法国国家科学研究中心IPAL) CerCo, CNRS(法国国家科学研究中心CerCo)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Abstract Paper (Non-Archival) @ ICCV 2025 NeVi Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09014 2025-09-12 cs.CV cs.CL 70%

COCO-Urdu: A Large-Scale Urdu Image-Caption Dataset with Multimodal Quality Estimation

Umair Hassan

机构 * Independent Researcher(独立研究者)

专题命中 视觉定位与Grounding :vision-language model(abstract);grounding(abstract);分类 cs.CV

Comments 17 pages, 3 figures, 3 tables. Dataset available at https://huggingface.co/datasets/umairhassan02/urdu-translated-coco-captions-subset. Scripts and notebooks to reproduce results available at https://github.com/umair-hassan2/COCO-Urdu

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09172 2025-09-12 cs.CV 57%

Bridging the Gap Between Ideal and Real-world Evaluation: Benchmarking AI-Generated Image Detection in Challenging Scenarios

Chunxiao Li, Xiaoxiao Wang, Meiling Li, Boming Miao, Peng Sun, Yunjian Zhang, Xiangyang Ji, Yao Zhu

机构 * Beijing Normal University(北京师范大学) University of Chinese Academy of Sciences(中国科学院大学) Fudan University(复旦大学) Central University of Finance and Economics(中央财经大学) Tsinghua University(清华大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments ICCV2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09281 2025-09-12 cs.HC 50%

Flip Co-op: Cooperative Takeovers in Shared Autonomy

Sandeep Banik, Naira Hovakimyan

专题命中 视觉定位与Grounding :grounding(abstract)

Comments 11 pages and 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 文档图表理解 1 篇

2509.09286 2025-09-12 cs.CV 77%

Visual Programmability: A Guide for Code-as-Thought in Chart Understanding

Bohao Tang, Yan Ma, Fei Zhang, Jiadi Su, Ethan Chern, Zhulin Hu, Zhixin Wang, Pengfei Liu, Ya Zhang

专题命中 文档图表理解 :vision-language model(abstract);VLM(abstract);visual reasoning(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

5. GUI与屏幕智能体 2 篇

2411.13591 2025-09-12 cs.CV cs.AI cs.CL 87%

Improved GUI Grounding via Iterative Narrowing

Anthony Nguyen

机构 * Algoma University(阿尔戈马大学)

专题命中 GUI与屏幕智能体 :grounding(title,abstract);vision-language model(abstract);VLM(abstract);分类 cs.CV、cs.AI

Comments Code available at https://github.com/ant-8/GUI-Grounding-via-Iterative-Narrowing

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03700 2025-09-12 cs.HC cs.AI 57%

MagicGUI: A Foundational Mobile GUI Agent with Scalable Data Pipeline and Reinforcement Fine-tuning

Liujian Tang, Shaokang Dong, Yijia Huang, Minqi Xiang, Hongtao Ruan, Bin Wang, Shuo Li, Zhiheng Xi, Zhihui Cao, Hailiang Pang, Heng Kong, He Yang, Mingxu Chai, Zhilin Gao, Xingyu Liu, Yingnan Fu, Jiaming Liu, Xuanjing Huang, Yu-Gang Jiang, Tao Gui, Qi Zhang, Kang Wang, Yunke Zhang, Yuran Wang

专题命中 GUI与屏幕智能体 :grounding(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 幻觉与鲁棒性 2 篇

2509.09397 2025-09-12 cs.CV 57%

Decoupling Clinical and Class-Agnostic Features for Reliable Few-Shot Adaptation under Shift

Umaima Rahman, Raza Imam, Mohammad Yaqub, Dwarikanath Mahapatra

机构 * Mohamed Bin Zayed University of Artificial Intelligence(莫扎德人工智能大学) Khalifa University(卡利法大学)

专题命中 幻觉与鲁棒性 :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.03486 2025-09-12 cs.CR cs.CV cs.SI 57%

UnsafeBench: Benchmarking Image Safety Classifiers on Real-World and AI-Generated Images

Yiting Qu, Xinyue Shen, Yixin Wu, Michael Backes, Savvas Zannettou, Yang Zhang

机构 * CISPA Helmholtz Center for Information Security(CISPA赫尔姆霍茨信息安全中心) TU Delft(代尔夫特理工大学)

专题命中 幻觉与鲁棒性 :visual language model(abstract);分类 cs.CV

Comments To Appear in the ACM Conference on Computer and Communications Security (CCS), October 13, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

7. VLM训练与架构 6 篇

2509.08913 2025-09-12 eess.IV 85%

Generalized User-Oriented Image Semantic Coding Empowered by Large Vision-Language Model

Sin-Yu Huang, Vincent W. S. Wong

专题命中 VLM训练与架构 :vision-language model(title,abstract);VLM(abstract);LLaVA(abstract)

Comments Accepted by IEEE Global Communications Conference (GLOBECOM), Taipei, Taiwan, Dec. 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09064 2025-09-12 cs.CV 79%

Enhancing 3D Medical Image Understanding with Pretraining Aided by 2D Multimodal Large Language Models

Qiuhui Chen, Xuancheng Yao, Huping Ye, Yi Hong

机构 * School of Computer Science, Shanghai Jiao Tong University(上海交通大学计算机科学学院)

专题命中 VLM训练与架构 :multimodal large language model(title,abstract);分类 cs.CV

Comments Accepted by IEEE Journal of Biomedical and Health Informatics (JBHI)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09356 2025-09-12 cs.AI cs.RO 70%

Curriculum-Based Multi-Tier Semantic Exploration via Deep Reinforcement Learning

Abdel Hakim Drid, Vincenzo Suriani, Daniele Nardi, Abderrezzak Debilou

机构 * Department of Electrical Engineering - Mohamed Khider, University of Biskra, Biskra (Algeria)(巴尔克拉大学电子工程系) Department of Engineering - University of Basilicata, Potenza (Italy)(巴塞里卡大学工程系) Department of Computer, Control, and Management Engineering ``Antonio Ruberti'', Sapienza University of Rome, Rome (Italy)(罗马萨皮恩扎大学计算机、控制与管理工程系)

专题命中 VLM训练与架构 :vision-language model(abstract);VLM(abstract);分类 cs.AI

Comments The 19th International Conference on Intelligent Autonomous Systems (IAS 19), 2025, Genoa

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09311 2025-09-12 cs.CV 70%

Image Recognition with Vision and Language Embeddings of VLMs

Illia Volkov, Nikita Kisel, Klara Janouskova, Jiri Matas

机构 * Visual Recognition Group, Faculty of Electrical Engineering, Czech Technical University in Prague(视觉识别组,电气工程学院,布拉格捷克技术大学)

专题命中 VLM训练与架构 :vision-language model(abstract);VLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08940 2025-09-12 cs.CV 57%

Discovering Divergent Representations between Text-to-Image Models

Lisa Dunlap, Joseph E. Gonzalez, Trevor Darrell, Fabian Caba Heilbron, Josef Sivic, Bryan Russell

专题命中 VLM训练与架构 :VLM(abstract);分类 cs.CV

Comments Accepted to ICCV 2025. Code available at https://github.com/adobe-research/CompCon

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09629 2025-09-12 cs.CL 50%

Bridging the Capability Gap: Joint Alignment Tuning for Harmonizing LLM-based Multi-Agent Systems

Minghang Zhu, Zhengliang Shi, Zhiwei Xu, Shiguang Wu, Lingjie Wang, Pengjie Ren, Zhaochun Ren, Zhumin Chen

机构 * Shandong University(山东大学) Leiden University(莱顿大学)

专题命中 VLM训练与架构 :grounding(abstract)

Comments EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏

8. 其他VLM 6 篇

2506.19662 2025-09-12 physics.ed-ph 78%

Multimodal large language models and physics visual tasks: comparative analysis of performance and costs

Giulia Polverini, Bor Gregorcic

专题命中 其他VLM :multimodal large language model(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09307 2025-09-12 cs.CV cs.AI cs.CL cs.MM 62%

Can Multimodal LLMs See Materials Clearly? A Multimodal Benchmark on Materials Characterization

Zhengzhao Lai, Youbin Zheng, Zhenyang Cai, Haonan Lyu, Jinpu Yang, Hongqing Liang, Yan Hu, Benyou Wang

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.21831 2025-09-12 cs.CV cs.AI 62%

Early Exit and Multi Stage Knowledge Distillation in VLMs for Video Summarization

Anas Anwarul Haq Khan, Utkarsh Verma, Ganesh Ramakrishnan

机构 * Department of Computer Science and Engineering, IIT Bombay(印度理工学院班加罗尔计算机科学与工程系) Center of Machine Intelligence and Data Science (C-MInDS), IIT Bombay(印度理工学院班加罗尔人工智能与数据科学中心)

专题命中 其他VLM :vision language model(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09541 2025-09-12 cs.AI 57%

Compositional Concept Generalization with Variational Quantum Circuits

Hala Hawashin, Mina Abbaszadeh, Nicholas Joseph, Beth Pearson, Martha Lewis, Mehrnoosh sadrzadeh

机构 * School of Computer Science Engineering University of New South Wales Sydney, Australia Stanford University California, USA Computer Science University College London London, UK School of Eng. Maths. \& Tech University of Bristol Bristol, UK Inst. Logic Language \& Computation University of Amsterdam Amsterdam, NL

专题命中 其他VLM :vision-language model(abstract);分类 cs.AI

Comments Accepted to: 2025 IEEE International Conference on Quantum Artificial Intelligence (QAI), Naples, Italy, Nov 2-5, 2025. This is the authors' accepted manuscript (AAM). An IEEE copyright notice appears on page 1. The final published version will appear in IEEE Xplore; DOI to be added when available

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.11538 2025-09-12 cs.CL cs.AI eess.AS 57%

MERaLiON-SpeechEncoder: Towards a Speech Foundation Model for Singapore and Beyond

Muhammad Huzaifah, Geyu Lin, Tianchi Liu, Hardik B. Sailor, Kye Min Tan, Tarun K. Vangani, Qiongqiong Wang, Jeremy H. M. Wong, Jinyang Wu, Nancy F. Chen, Ai Ti Aw

机构 * MERaLiON Team(MERaLiON团队) Institute for Infocomm Research (I 2 R), A*STAR, Singapore(信息通信研究所(I2R),A*STAR,新加坡)

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.04077 2025-09-12 cs.CL cs.SD eess.AS 50%

A Novel Data Augmentation Approach for Automatic Speaking Assessment on Opinion Expressions

Chung-Chun Wang, Jhen-Ke Lin, Hao-Chien Lu, Hong-Yun Lin, Berlin Chen

机构 * National Taiwan Normal University(台湾国立台湾师范大学)

专题命中 其他VLM :multimodal large language model(abstract)

Comments submitted to the ISCA SLaTE-2025 Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏