arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-11-06 至 2025-11-06 共收录 27 信号源:cs.CV, cs.AI, cs.LG

1. 视觉问答 4 篇

2511.03328 2025-11-06 cs.CL cs.AI cs.CV cs.LG 82%

Benchmarking the Thinking Mode of Multimodal Large Language Models in Clinical Tasks

Jindong Hong, Tianjie Chen, Lingjie Luo, Chuanyang Zheng, Ting Xu, Haibao Yu, Jianing Qiu, Qianzhong Chen, Suning Huang, Yan Xu, Yong Gui, Yijun He, Jiankai Sun

机构 * Bytedance(字节跳动) Peking University(北京大学) The Chinese University of Hong Kong(香港中文大学) The University of Hong Kong(香港大学) Mohamed bin Zayed University of Artificial Intelligence(马尔代夫人工智能大学) Stanford University(斯坦福大学) University of Michigan(密歇根大学)

专题命中 视觉问答 :multimodal large language model(title,abstract);分类 cs.CV、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03617 2025-11-06 cs.GR cs.AI 79%

Visualization Biases MLLM's Decision Making in Network Data Tasks

Timo Brand, Henry Förster, Stephen G. Kobourov, Jacob Miller

机构 * Technical University of Munich, Heilbronn, Germany(慕尼黑技术大学)

专题命中 视觉问答 :MLLM(title,abstract);分类 cs.AI

Comments This manuscript was presented at VIS x GenAI, a workshop co-located with IEEE VIS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08668 2025-11-06 cs.CV 77%

Hulu-Med: A Transparent Generalist Model towards Holistic Medical Vision-Language Understanding

Songtao Jiang, Yuan Wang, Sibo Song, Tianxiang Hu, Chenyi Zhou, Bin Pu, Yan Zhang, Zhibo Yang, Yang Feng, Joey Tianyi Zhou, Jin Hao, Zijian Chen, Ruijia Wu, Tao Tang, Junhui Lv, Hongxia Xu, Hongwei Wang, Jun Xiao, Bin Feng, Fudong Zhu, Kenli Li, Weidi Xie, Jimeng Sun, Jian Wu, Zuozhu Liu

机构 * College of Computer Science and Technology, Zhejiang University-University of Illinois Urbana-Champaign Institute(浙江大学计算机科学与技术学院) Stomatology Hospital, School of Stomatology, Zhejiang University School of Medicine(浙江大学口腔医院) Alibaba Inc(阿里巴巴集团) College of Computer Science and Electronic Engineering, Hunan University(湖南大学计算机科学与电子工程学院) Angelalign Technology Inc.(Angelalign技术有限公司) CFAR & IHPC, Agency for Science, Technology and Research(CFAR与IHPC,新加坡科技研究局) Department of Orthodontics, Shanghai Ninth People’s Hospital, College of Stomatology, Shanghai Jiao Tong University(上海第九人民医院正畸科,上海交通大学口腔医学院)

专题命中 视觉问答 :vision-language model(abstract);VLM(abstract);visual question answering(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03178 2025-11-06 cs.CV 57%

SurgAnt-ViVQA: Learning to Anticipate Surgical Events through GRU-Driven Temporal Cross-Attention

Shreyas C. Dhake, Jiayuan Huang, Runlong He, Danyal Z. Khan, Evangelos B. Mazomenos, Sophia Bano, Hani J. Marcus, Danail Stoyanov, Matthew J. Clarkson, Mobarak I. Hoque

机构 * UCL Hawkes Institute(UCL哈维斯研究所) University College London(伦敦大学学院) Dept of Medical Physics & Biomedical Engineering(医学物理与生物医学工程系) UCL(伦敦大学学院) Dept of Computer Science(计算机科学系) National Hospital for Neurology and Neurosurgery(神经病学与神经外科国家医院) Division of Informatics, Imaging and Data Science(信息学、成像与数据科学 division)

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

Comments 12 pages

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 视觉推理 8 篇

2506.21448 2025-11-06 eess.AS cs.CV cs.SD 79%

ThinkSound: Chain-of-Thought Reasoning in Multimodal Large Language Models for Audio Generation and Editing

Huadai Liu, Kaicheng Luo, Jialei Wang, Wen Wang, Qian Chen, Zhou Zhao, Wei Xue

机构 * Hong Kong University of Science and Technology (HKUST)(香港理工大学) Tongyi Fun Team, Alibaba Group(阿里云团队) Zhejiang University(浙江大学)

专题命中 视觉推理 :multimodal large language model(title,abstract);分类 cs.CV

Comments Accepted by NeurIPS 2025 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.04220 2025-11-06 cs.CV 77%

Struct2D: A Perception-Guided Framework for Spatial Reasoning in MLLMs

Fangrui Zhu, Hanhui Wang, Yiming Xie, Jing Gu, Tianye Ding, Jianwei Yang, Huaizu Jiang

机构 * Northeastern University(东北大学) Microsoft Research(微软研究院) University of Southern California(南加州大学) University of California, Santa Cruz(加州大学圣克鲁兹分校)

专题命中 视觉推理 :grounding(abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV

Comments NeurIPS 2025, code link: https://github.com/neu-vi/struct2d

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04201 2025-11-06 cs.CV cs.AI 73%

ViFP: A Framework for Visual False Positive Detection to Enhance Reasoning Reliability in VLMs

Ben Zhang, LuLu Yu, Lei Gao, QuanJiang Guo, Jing Liu, Hui Gao

机构 * School of Computer Science and Engineering, University of Electronic Science and Technology of China(电子科技大学计算机科学与工程学院)

专题命中 视觉推理 :vision-language model(abstract);VLM(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03206 2025-11-06 cs.CV cs.AI cs.LG 67%

QG-CoC: Question-Guided Chain-of-Captions for Large Multimodal Models

Kuei-Chun Kao, Hsu Tzu-Yin, Yunqi Hong, Ruochen Wang, Cho-Jui Hsieh

机构 * Department of Computer Science, University of California, Los Angeles(计算机科学系,加州大学洛杉矶分校)

专题命中 视觉推理 :multimodal large language model(abstract);分类 cs.CV、cs.AI、cs.LG

Comments 16 pages

Journal ref EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21861 2025-11-06 cs.LG cs.AI cs.CL 62%

The Mirror Loop: Recursive Non-Convergence in Generative Reasoning Systems

Bentley DeVilling

专题命中 视觉推理 :grounding(abstract);分类 cs.AI、cs.LG

Comments 18 pages, 2 figures. Category: cs.LG. Code and data: https://github.com/Course-Correct-Labs/mirror-loop

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02834 2025-11-06 cs.AI cs.CL cs.LG 62%

Agent-Omni: Test-Time Multimodal Reasoning via Model Coordination for Understanding Anything

Huawei Lin, Yunzhi Shi, Tong Geng, Weijie Zhao, Wei Wang, Ravender Pal Singh

机构 * Amazon(亚马逊公司) Rochester Institute of Technology(罗切斯特理工学院) University of Rochester(罗切斯特大学)

专题命中 视觉推理 :multimodal large language model(abstract);分类 cs.AI、cs.LG

Comments 16 pages, 7 figures, 14 tables. Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03471 2025-11-06 cs.AI cs.HC 57%

Towards Scalable Web Accessibility Audit with MLLMs as Copilots

Ming Gu, Ziwei Wang, Sicen Lai, Zirui Gao, Sheng Zhou, Jiajun Bu

专题命中 视觉推理 :multimodal large language model(abstract);分类 cs.AI

Comments 15 pages. Accepted by AAAI 2026 AISI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18140 2025-11-06 cs.CL 50%

MathOPEval: A Fine-grained Evaluation Benchmark for Visual Operations of MLLMs in Mathematical Reasoning

Xiaoyuan Li, Moxin Li, Wenjie Wang, Rui Men, Yichang Zhang, Fuli Feng, Dayiheng Liu

机构 * University of Science and Technology of China(中国科学技术大学) Alibaba Group(阿里巴巴集团) National University of Singapore(新加坡国立大学)

专题命中 视觉推理 :MLLM(abstract)

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视觉定位与Grounding 9 篇

2505.16495 2025-11-06 cs.CV 70%

ALTo: Adaptive-Length Tokenizer for Autoregressive Mask Generation

Lingfeng Wang, Hualing Lin, Senda Chen, Tao Wang, Changxu Cheng, Yangyang Zhong, Dong Zheng, Wuyue Zhao

机构 * Uni-Ubi Zhejiang University(浙江大学) Tongji University(同济大学)

专题命中 视觉定位与Grounding :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03549 2025-11-06 cs.SE cs.AI 57%

Uncovering Code Insights: Leveraging GitHub Artifacts for Deeper Code Understanding

Ziv Nevo, Orna Raz, Karen Yorav

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments 7 pages, 6 figures, to be published in AISM 2025, see https://aism25.github.io/aism25/

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.04789 2025-11-06 cs.CV 57%

Object-X: Learning to Reconstruct Multi-Modal 3D Object Representations

Gaia Di Lorenzo, Federico Tombari, Marc Pollefeys, Daniel Barath

机构 * ETH Zurich(苏黎世联邦理工学院) Google(谷歌) Microsoft(微软)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.11519 2025-11-06 cs.CV cs.CL 57%

Exploring Typographic Visual Prompts Injection Threats in Cross-Modality Generation Models

Hao Cheng, Erjia Xiao, Yichi Wang, Lingfeng Zhang, Qiang Zhang, Jiahang Cao, Kaidi Xu, Mengshu Sun, Xiaoshuai Hao, Jindong Gu, Renjing Xu

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) University of Oxford(牛津大学) Beijing Academy of Artificial Intelligence(北京人工智能研究院) The Hong Kong University of Science and Technology(香港科学与技术大学) Beijing University of Technology(北京工业大学) Tsinghua University(清华大学) City University of Hong Kong(香港城市大学)

专题命中 视觉定位与Grounding :vision language model(abstract);分类 cs.CV

Comments This paper is accepted by IJCAI2025 Workshop on Deepfake Detection, Localization, and Interpretability as Best Student Paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03186 2025-11-06 cs.AI 57%

Adobe Summit Concierge Evaluation with Human in the Loop

Yiru Chen, Sally Fang, Sai Sree Harsha, Dan Luo, Vaishnavi Muppala, Fei Wu, Shun Jiang, Kun Qian, Yunyao Li

机构 * Adobe Inc.(Adobe公司)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments Accepted by 6th Workshop on Data Science with Human in the Loop @ VLDB 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03181 2025-11-06 cs.RO cs.LG 57%

Learning-based Cooperative Robotic Paper Wrapping: A Unified Control Policy with Residual Force Control

Rewida Ali, Cristian C. Beltran-Hernandez, Weiwei Wan, Kensuke Harada

机构 * Department of Systems Innovation, Graduate School of Engineering Science, Osaka University(大阪大学系统创新部门,工学研究科) OMRON SINIC X Corporation(OMRON SINIC X公司) The National Institute of Advanced Industrial Science and Technology (AIST)(国家先进工业科学与技术研究院)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03047 2025-11-06 cs.LG 57%

Unsupervised Evaluation of Multi-Turn Objective-Driven Interactions

Emi Soroka, Tanmay Chopra, Krish Desai, Sanjay Lall

机构 * Department of Electrical Engineering Stanford University(电气工程系 斯坦福大学) Emissary Technologies

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG

Comments Under review at ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.07661 2025-11-06 cs.CV cs.RO 57%

ROADWork: A Dataset and Benchmark for Learning to Recognize, Observe, Analyze and Drive Through Work Zones

Anurag Ghosh, Shen Zheng, Robert Tamburo, Khiem Vuong, Juan Alvarez-Padilla, Hailiang Zhu, Michael Cardei, Nicholas Dunn, Christoph Mertz, Srinivasa G. Narasimhan

机构 * Carnegie Mellon University(卡内基梅隆大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments ICCV 2025 Accepted Paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03165 2025-11-06 cs.RO 50%

SENT Map -- Semantically Enhanced Topological Maps with Foundation Models

Raj Surya Rajendran Kathirvel, Zach A Chavis, Stephen J. Guy, Karthik Desingh

机构 * Minnesota Robotics Institute (MnRI)(明尼苏达州机器人研究所) Department of Computer Science and Engineering (CS&E)(计算机科学与工程系) University of Minnesota(明尼苏达大学)

专题命中 视觉定位与Grounding :grounding(abstract)

Comments Accepted at ICRA 2025 Workshop on Foundation Models and Neuro-Symbolic AI for Robotics

详情

展开后加载摘要…

URL PDF HTML 收藏

4. GUI与屏幕智能体 2 篇

2511.03497 2025-11-06 cs.RO cs.AI cs.SE 57%

ROSBag MCP Server: Analyzing Robot Data with LLMs for Agentic Embodied AI Applications

Lei Fu, Sahar Salimpour, Leonardo Militano, Harry Edelman, Jorge Peña Queralta, Giovanni Toffetti

专题命中 GUI与屏幕智能体 :VLM(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.13174 2025-11-06 cs.CV 57%

Manipulation Facing Threats: Evaluating Physical Vulnerabilities in End-to-End Vision Language Action Models

Hao Cheng, Erjia Xiao, Yichi Wang, Chengyuan Yu, Mengshu Sun, Qiang Zhang, Jiahang Cao, Yijie Guo, Ning Liu, Kaidi Xu, Jize Zhang, Chao Shen, Philip Torr, Jindong Gu, Renjing Xu

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) University of Oxford(牛津大学) Xi’an Jiaotong University(西安交通大学) The Hong Kong University of Science and Technology(香港科学与技术大学) City University of Hong Kong(香港城市大学) Beijing University of Technology(北京理工大学) Duke University(杜克大学) X-Humanoid Project(X-Humanoid 项目)

专题命中 GUI与屏幕智能体 :multimodal large language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 幻觉与鲁棒性 1 篇

2511.03367 2025-11-06 cs.CV cs.AI cs.LG 78%

Decoupling Augmentation Bias in Prompt Learning for Vision-Language Models

Gahyeon Kim, Sohee Kim, Seokju Lee

机构 * Korea Institute of Energy Technology (KENTECH)(韩国能源技术研究所)

专题命中 幻觉与鲁棒性 :vision-language model(title);分类 cs.CV、cs.AI、cs.LG

Comments Accepted in Pattern Recognition

详情

展开后加载摘要…

URL PDF HTML 收藏

6. VLM训练与架构 2 篇

2511.03332 2025-11-06 cs.CV 83%

Multi-Object Tracking Retrieval with LLaVA-Video: A Training-Free Solution to MOT25-StAG Challenge

Yi Yang, Yiming Xu, Timo Kaiser, Hao Cheng, Bodo Rosenhahn, Michael Ying Yang

机构 * Leibniz University Hannover(莱布尼茨汉诺威大学) University of Twente(特文特大学) University of Bath(巴斯大学)

专题命中 VLM训练与架构 :LLaVA(title,abstract);grounding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02996 2025-11-06 cs.CV 57%

SCALE-VLP: Soft-Weighted Contrastive Volumetric Vision-Language Pre-training with Spatial-Knowledge Semantics

Ailar Mahdizadeh, Puria Azadi Moghadam, Xiangteng He, Shahriar Mirabbasi, Panos Nasiopoulos, Leonid Sigal

机构 * University of British Columbia(不列颠哥伦比亚大学) Vector Institute for AI(人工智能向量研究所)

专题命中 VLM训练与架构 :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

7. 其他VLM 1 篇

2511.03137 2025-11-06 cs.AI 57%

Using Multi-modal Large Language Model to Boost Fireworks Algorithm's Ability in Settling Challenging Optimization Tasks

Shipeng Cen, Ying Tan

机构 * School of Intelligence Science Technology, Institute for Artificial Intellignce, Peking University, Beijing, China Technology, Institute for Artificial Intellignce, National Key Laboratory of General Artificial Intelligence, Peking University, Beijing, China

专题命中 其他VLM :MLLM(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏