arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-09-23 至 2025-09-23 共收录 73 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 17 篇

2509.17522 2025-09-23 cs.CV 57%

Chat-CBM: Towards Interactive Concept Bottleneck Models with Frozen Large Language Models

Hangzhou He, Lei Zhu, Kaiwen Li, Xinliang Zhang, Jiakui Hu, Ourui Fu, Zhengjian Yao, Yanye Lu

机构 * Department of Biomedical Engineering, College of Future Technology, Peking University(生物医学工程系,未来技术学院,北京大学) Institute of Medical Technology, Peking University Health Science Center, Peking University(医学技术研究所,北京大学医学部,北京大学) National Biomedical Imaging Center, College of Future Technology, Peking University(国家生物医学成像中心,未来技术学院,北京大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16810 2025-09-23 cs.AI 57%

Automated Procedural Analysis via Video-Language Models for AI-assisted Nursing Skills Assessment

Shen Chang, Dennis Liu, Renran Tian, Kristen L. Swartzell, Stacie L. Klingler, Amy M. Nagle, Nan Kong

机构 * Weldon School of Biomedical Engineering, Purdue University(普渡大学生物医学工程学院) Department of Industrial and Operations Engineering, University of Michigan(密歇根大学工业与运作工程系) Edward P. Fitts Department of Industrial and Systems Engineering, North Carolina State University(北卡罗来纳州立大学工业与系统工程系) School of Nursing, Purdue University(普渡大学护理学院)

专题命中 视觉定位与Grounding :VLM(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04484 2025-09-23 cs.CL cs.AI cs.CY 57%

The Good, the Bad and the Constructive: Automatically Measuring Peer Review's Utility for Authors

Abdelrahman Sadallah, Tim Baumgärtner, Iryna Gurevych, Ted Briscoe

机构 * NLP Department, Mohamed Bin Zayed University of Artificial Intelligence(马尔代夫比兹艾兹大学人工智能学院自然语言处理系) Ubiquitous Knowledge Processing Lab, Department of Computer Science(计算机科学系通用知识处理实验室) Hessian Center for AI (hessian.AI), TU Darmstadt(图尔努尔德马斯特大学海斯塞人工智能中心)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments EMNLP 2025 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15108 2025-09-23 cs.CL cs.AI cs.HC 57%

A Risk Ontology for Evaluating AI-Powered Psychotherapy Virtual Agents

Ian Steenstra, Timothy W. Bickmore

机构 * Northeastern University(东北大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments This is a preprint version of the paper accepted to IVA'25

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16502 2025-09-23 cs.LG 57%

GRIL: Knowledge Graph Retrieval-Integrated Learning with Large Language Models

Jialin Chen, Houyu Zhang, Seongjun Yun, Alejandro Mottini, Rex Ying, Xiang Song, Vassilis N. Ioannidis, Zheng Li, Qingjun Cui

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16297 2025-09-23 cs.CY cs.AI cs.CL 57%

How Large Language Models are Designed to Hallucinate

Richard Ackermann, Simeon Emanuilov

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments 23 pages, 2 tables, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17727 2025-09-23 cs.CY cs.IT math.IT 50%

Empirical AI Ethics: Reconfiguring Ethics towards a Situated, Plural, and Transformative Approach

Paula Helm, Selin Gerlek

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17523 2025-09-23 cs.CL eess.AS 50%

Leveraging Audio-Visual Data to Reduce the Multilingual Gap in Self-Supervised Speech Models

María Andrea Cruz Blandón, Zakaria Aldeneh, Jie Chi, Maureen de Seyssel

机构 * Tampere University Apple(塔尔库大学苹果)

专题命中 视觉定位与Grounding :grounding(abstract)

Comments 5 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16670 2025-09-23 cs.SD cs.MM eess.AS 50%

Speech-to-See: End-to-End Speech-Driven Open-Set Object Detection

Wenhuan Lu, Xinyue Song, Wenjun Ke, Zhizhi Yu, Wenhao Yang, Jianguo Wei

机构 * College of Intelligence and Computing, Tianjin University, Tianjin, China(智能与计算学院,天津大学,天津,中国) PipeChina Institute of Science and Technology, Tianjin, China(中石油科技研究院,天津,中国)

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.08851 2025-09-23 cs.RO 50%

OTAS: Open-vocabulary Token Alignment for Outdoor Segmentation

Simon Schwaiger, Stefan Thalhammer, Wilfried Wöber, Gerald Steinbauer-Wagner

机构 * Graz University of Technology, Faculty of Computer Science and Biomedical Engineering, Institute of Software Engineering and Artificial Intelligence(格拉茨技术大学,计算机科学与生物医学工程学院,软件工程与人工智能研究所) University of Applied Sciences Technikum Wien, Faculty of Industrial Engineering, Research Group Digital Manufacturing, Automation and Robotics(应用科学大学技术学院,工业工程学院,数字制造、自动化与机器人研究组) University of Natural Resources and Life Sciences, Department of Integrative Biology and Biodiversity Research, Institute for Integrative Nature Conservation Research(自然资源与生命科学大学,整合生物学与生物多样性研究部门,整合自然保护研究 institute)

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 文档图表理解 3 篇

2509.17481 2025-09-23 cs.CV cs.AI cs.CL 81%

ChartHal: A Fine-grained Framework Evaluating Hallucination of Large Vision Language Models in Chart Understanding

Xingqi Wang, Yiming Cui, Xin Yao, Shijin Wang, Guoping Hu, Xiaoyu Qin

机构 * Department of Computer Science and Technology, Tsinghua University(清华大学计算机科学与技术系) State Key Laboratory of Cognitive Intelligence, iFLYTEK(认知智能国家重点实验室)

专题命中 文档图表理解 :vision language model(title);vision-language model(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.20168 2025-09-23 cs.CV 79%

Seeing is Believing? Mitigating OCR Hallucinations in Multimodal Large Language Models

Zhentao He, Can Zhang, Ziheng Wu, Zhenghao Chen, Yufei Zhan, Yifan Li, Zhao Zhang, Xian Wang, Minghui Qiu

机构 * ByteDance(字节跳动) CASIA(中国科学院自动化研究所) RUC(中国人民大学)

专题命中 文档图表理解 :multimodal large language model(title,abstract);分类 cs.CV

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17589 2025-09-23 cs.AI 70%

Table2LaTeX-RL: High-Fidelity LaTeX Code Generation from Table Images via Reinforced Multimodal Language Models

Jun Ling, Yao Qi, Tao Huang, Shibo Zhou, Yanqin Huang, Jiang Yang, Ziqi Song, Ying Zhou, Yang Yang, Heng Tao Shen, Peng Wang

机构 * School of Computer Science and Engineering, University of Electronic Science and Technology of China(电子科技大学计算机科学与工程学院) Research Center for Scientific Data Hub, Zhejiang Lab, Hangzhou, China(浙江实验室科学数据中心研究中心) School of Computer Science and Technology, Tongji University(同济大学计算机科学与技术学院)

专题命中 文档图表理解 :multimodal large language model(abstract);MLLM(abstract);分类 cs.AI

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

3. GUI与屏幕智能体 6 篇

2509.16452 2025-09-23 cs.CV cs.AI 84%

KRAST: Knowledge-Augmented Robotic Action Recognition with Structured Text for Vision-Language Models

Son Hai Nguyen, Diwei Wang, Jinhyeok Jang, Hyewon Seo

机构 * ETRI(韩国科学技术院)

专题命中 GUI与屏幕智能体 :vision-language model(title,abstract);VLM(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17328 2025-09-23 cs.CV cs.HC 70%

UIPro: Unleashing Superior Interaction Capability For GUI Agents

Hongxin Li, Jingran Su, Jingfan Chen, Zheng Ju, Yuntao Chen, Qing Li, Zhaoxiang Zhang

机构 * University of Chinese Academy of Sciences (UCAS)(中国科学院大学) New Laboratory of Pattern Recognition (NLPR), CASIA(中国科学院自动化所模式识别新实验室) State Key Laboratory of Multimodal Artificial Intelligence Systems (MAIS), CASIA(中国科学院多模态人工智能系统国家重点实验室) Hong Kong Institute of Science & Innovation, CASIA(中国科学院香港创新科学研究院) PolyU Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

专题命中 GUI与屏幕智能体 :vision-language model(abstract);grounding(abstract);分类 cs.CV

Comments Accepted to ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17941 2025-09-23 cs.RO cs.AI cs.CV cs.LG 67%

ComposableNav: Instruction-Following Navigation in Dynamic Environments via Composable Diffusion

Zichao Hu, Chen Tang, Michael J. Munje, Yifeng Zhu, Alex Liu, Shuijing Liu, Garrett Warnell, Peter Stone, Joydeep Biswas

机构 * Department of Computer Science, The University of Texas at Austin(德克萨斯大学计算机科学系) Army Research Laboratory(陆军研究实验室) Sony AI(索尼人工智能)

专题命中 GUI与屏幕智能体 :VLM(abstract);分类 cs.CV、cs.AI、cs.LG

Comments Conference on Robot Learning (CoRL) 2025 Project site: https://amrl.cs.utexas.edu/ComposableNav/

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.08283 2025-09-23 cs.IR 67%

Serendipitous Recommendation with Multimodal LLM

Haoting Wang, Jianling Wang, Hao Li, Fangjun Yi, Mengyu Fu, Youwei Zhang, Yifan Liu, Liang Liu, Minmin Chen, Ed H. Chi, Lichan Hong, Haokai Lu

专题命中 GUI与屏幕智能体 :multimodal large language model(abstract);MLLM(abstract)

Comments Accepted by 2025 Recsys EARL Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22424 2025-09-23 cs.LG cs.AI 62%

Spec-VLA: Speculative Decoding for Vision-Language-Action Models with Relaxed Acceptance

Songsheng Wang, Rucheng Yu, Zhihang Yuan, Chao Yu, Feng Gao, Yu Wang, Derek F. Wong

机构 * NICS-EFC Lab, Department of Electronic Engineering, Tsinghua University(清华大学电子工程系NICS-EFC实验室) Department of Computer and Information Science, University of Macau(澳门大学计算机与信息科学系) Infinigence AI Tsinghua University(清华大学) Zhongguancun Academy(中关村学院)

专题命中 GUI与屏幕智能体 :visual language model(abstract);分类 cs.AI、cs.LG

Comments 13 pages, 5 figures, Accepted by EMNLP 2025 (main conference)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17917 2025-09-23 cs.AI 57%

Orcust: Stepwise-Feedback Reinforcement Learning for GUI Agent

Junyu Lu, Songxin Zhang, Zejian Xie, Zhuoyang Song, Jiaxing Zhang

机构 * Lionrock AI Lab(狮岩人工智能实验室) China Merchants Research Institute of Advanced Technology(中国商人先进科技研究院)

专题命中 GUI与屏幕智能体 :grounding(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 幻觉与鲁棒性 8 篇

2509.16805 2025-09-23 cs.CV 85%

Benchmarking and Mitigating MCQA Selection Bias of Large Vision-Language Models

Md. Atabuzzaman, Ali Asgarov, Chris Thomas

专题命中 幻觉与鲁棒性 :vision-language model(title,abstract);visual reasoning(abstract);visual question answering(abstract);分类 cs.CV

Comments Accepted to EMNLP 2025 (Main Conference)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17265 2025-09-23 cs.LG cs.AI 84%

SUA: Stealthy Multimodal Large Language Model Unlearning Attack

Xianren Zhang, Hui Liu, Delvin Ce Zhang, Xianfeng Tang, Qi He, Dongwon Lee, Suhang Wang

机构 * The Pennsylvania State University(宾夕法尼亚州立大学) Amazon(亚马逊) University of Sheffield(谢菲尔德大学)

专题命中 幻觉与鲁棒性 :multimodal large language model(title,abstract);MLLM(abstract);分类 cs.AI、cs.LG

Comments EMNLP25

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16645 2025-09-23 cs.CV 83%

ADVEDM:Fine-grained Adversarial Attack against VLM-based Embodied Agents

Yichen Wang, Hangtao Zhang, Hewen Pan, Ziqi Zhou, Xianlong Wang, Peijin Guo, Lulu Xue, Shengshan Hu, Minghui Li, Leo Yu Zhang

机构 * National Engineering Research Center for Big Data Technology and System(大数据技术与系统国家工程研究中心) Services Computing Technology and System Lab(服务计算技术与系统实验室) Cluster and Grid Computing Lab(集群与网格计算实验室) Hubei Engineering Research Center on Big Data Security(湖北省大数据安全工程研究中心) Hubei Key Laboratory of Distributed System Security(湖北省分布式系统安全重点实验室) School of Cyber Science and Engineering, Huazhong University of Science and Technology(华中科技大学信息科学与工程学院) School of Computer Science and Technology, Huazhong University of Science and Technology(华中科技大学计算机科学与技术学院) Department of Computer Science, City University of HongKong(香港城市大学计算机科学系) School of Software Engineering, Huazhong University of Science and Technology(华中科技大学软件工程学院) School of Information and Communication Technology, Griffith University(格里菲斯大学信息与通信技术学院)

专题命中 幻觉与鲁棒性 :VLM(title,abstract);vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.21059 2025-09-23 cs.CV cs.AI cs.CR cs.LG 82%

FC-Attack: Jailbreaking Multimodal Large Language Models via Auto-Generated Flowcharts

Ziyi Zhang, Zhen Sun, Zongmin Zhang, Jihui Guo, Xinlei He

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))

专题命中 幻觉与鲁棒性 :multimodal large language model(title,abstract);分类 cs.CV、cs.AI、cs.LG

Comments Accepted to Findings of EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.04039 2025-09-23 cs.CV cs.AI cs.CL 81%

Mitigating Hallucinations in Large Vision-Language Models via Entity-Centric Multimodal Preference Optimization

Jiulong Wu, Zhengliang Shi, Shuaiqiang Wang, Jizhou Huang, Dawei Yin, Lingyong Yan, Min Cao, Min Zhang

机构 * Soochow University(苏州大学) Baidu Inc.(百度公司) Shandong University(山东大学)

专题命中 幻觉与鲁棒性 :vision-language model(title);visual language model(abstract);分类 cs.CV、cs.AI

Comments This paper is accepted by EMNLP2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16611 2025-09-23 cs.RO 67%

Video-to-BT: Generating Reactive Behavior Trees from Human Demonstration Videos for Robotic Assembly

Xiwei Zhao, Yiwei Wang, Yansong Wu, Fan Wu, Teng Sun, Zhonghua Miao, Sami Haddadin, Alois Knoll

机构 * Munich Institute of Robotics and Machine Intelligence (MIRMI)(慕尼黑机器人与机器智能研究所) Technical University of Munich(慕尼黑技术大学) Aalto University(艾尔沃斯大学) Shanghai University(上海大学) Mohamed Bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)

专题命中 幻觉与鲁棒性 :vision-language model(abstract);VLM(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11669 2025-09-23 cs.CV 57%

Co-STAR: Collaborative Curriculum Self-Training with Adaptive Regularization for Source-Free Video Domain Adaptation

Amirhossein Dadashzadeh, Parsa Esmati, Majid Mirmehdi

机构 * University of Bristol, UK(布里斯托大学)

专题命中 幻觉与鲁棒性 :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.19269 2025-09-23 cs.CV 57%

Neural Antidote: Class-Wise Prompt Tuning for Purifying Backdoors in CLIP

Jiawei Kong, Hao Fang, Sihang Guo, Chenxi Qing, Kuofeng Gao, Bin Chen, Shu-Tao Xia, Ke Xu

机构 * Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院,清华大学) School of Computer Science and Technology, Harbin Institute of Technology(哈尔滨工业大学计算机科学与技术学院) Department of Computer Science and Technology, Tsinghua University(清华大学计算机科学与技术系)

专题命中 幻觉与鲁棒性 :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

5. VLM训练与架构 10 篇

2509.11815 2025-09-23 cs.CV cs.AI 86%

SpecVLM: Fast Speculative Decoding in Vision-Language Models

Haiduo Huang, Fuwei Yang, Zhenhua Liu, Xuanwu Yin, Dong Li, Pengju Ren, Emad Barsoum

机构 * Advanced Micro Devices, Inc.(先进微器件公司) Institute of Artificial Intelligence and Robotics(人工智能与机器人研究院)

专题命中 VLM训练与架构 :vision-language model(title,abstract);VLM(abstract);LLaVA(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.13146 2025-09-23 cs.CV cs.LG 84%

Re-Align: Aligning Vision Language Models via Retrieval-Augmented Direct Preference Optimization

Shuo Xing, Peiran Li, Yuping Wang, Ruizheng Bai, Yueqi Wang, Chan-Wei Hu, Chengxuan Qian, Huaxiu Yao, Zhengzhong Tu

机构 * Texas A&M University(德克萨斯大学) University of Michigan(密歇根大学) UIUC(伊利诺伊大学香槟分校) UNC Chapel Hill(北卡罗来纳大学教堂山分校)

专题命中 VLM训练与架构 :vision language model(title,abstract);VLM(abstract);分类 cs.CV、cs.LG

Comments Published at EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17418 2025-09-23 cs.CL cs.CV 83%

Vision Language Models Are Not (Yet) Spelling Correctors

Junhong Liang, Bojun Zhang

机构 * MBZUAI Institute of Automation, Chinese Academy of Sciences(自动化研究所,中国科学院)

专题命中 VLM训练与架构 :vision language model(title,abstract);InternVL(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏