arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-11-07 至 2025-11-07 共收录 32 信号源:cs.CV, cs.AI, cs.LG

1. 视觉问答 5 篇

2511.04384 2025-11-07 cs.CV cs.LG 74%

Multi-Task Learning for Visually Grounded Reasoning in Gastrointestinal VQA

Itbaan Safwan, Muhammad Annas Shaikh, Muhammad Haaris, Ramail Khan, Muhammad Atif Tahir

机构 * School of Mathematics and Computer Science, Institute of Business Administration (IBA), Karachi, Pakistan(数学与计算机科学学院,商学院(IBA),巴基斯坦卡里奇)

专题命中 视觉问答 :visual question answering(abstract,comments);grounding(abstract);分类 cs.CV、cs.LG

Comments This is a working paper submitted for Medico 2025: Visual Question Answering (with multimodal explanations) for Gastrointestinal Imaging at MediaEval 2025. 5 pages, 3 figures and 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10528 2025-11-07 cs.CV cs.AI 73%

Med-GLIP: Advancing Medical Language-Image Pre-training with Large-scale Grounded Dataset

Ziye Deng, Ruihan He, Jiaxiang Liu, Yuan Wang, Zijie Meng, Songtao Jiang, Yong Xie, Zuozhu Liu

机构 * Zhejiang University(浙江大学) Guangdong Institute of Intelligence Science(广东智能科学研究院)

专题命中 视觉问答 :visual question answering(abstract);grounding(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21814 2025-11-07 cs.CV cs.AI 62%

Gestura: A LVLM-Powered System Bridging Motion and Semantics for Real-Time Free-Form Gesture Understanding

Zhuoming Li, Aitong Liu, Mengxi Jia, Yubi Lu, Tengxiang Zhang, Changzhi Sun, Dell Zhang, Xuelong Li

机构 * Institute of Artificial Intelligence (TeleAI) of China Telecom(中国电信人工智能研究院(TeleAI))

专题命中 视觉问答 :vision-language model(abstract);分类 cs.CV、cs.AI

Comments IMWUT2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04502 2025-11-07 cs.CL cs.AI 57%

RAGalyst: Automated Human-Aligned Agentic Evaluation for Domain-Specific RAG

Joshua Gao, Quoc Huy Pham, Subin Varghese, Silwal Saurav, Vedhus Hoskere

机构 * University of Houston(德克萨斯大学休斯敦分校)

专题命中 视觉问答 :grounding(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03325 2025-11-07 cs.CV 57%

SurgViVQA: Temporally-Grounded Video Question Answering for Surgical Scene Understanding

Mauro Orazio Drago, Luca Carlini, Pelinsu Celebi Balyemez, Dennis Pierantozzi, Chiara Lena, Cesare Hassan, Danail Stoyanov, Elena De Momi, Sophia Bano, Mobarak I. Hoque

机构 * Dipartimento di Elettronica, Informazione e Bioingegneria (DEIB)(电子、信息与生物工程系) Politecnico di Milano(米兰理工大学) IRCCS Humanitas Research Hospital(IRCCS人类itas研究医院) UCL Hawkes Institute and Department of Computer Science(UCL Hawkes研究所和计算机科学系) University College London(伦敦大学学院) University of Manchester(曼彻斯特大学)

专题命中 视觉问答 :visual reasoning(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 视觉推理 7 篇

2510.11190 2025-11-07 cs.CV 79%

FlexAC: Towards Flexible Control of Associative Reasoning in Multimodal Large Language Models

Shengming Yuan, Xinyu Lyu, Shuailong Wang, Beitao Chen, Jingkuan Song, Lianli Gao

机构 * University of Electronic Science and Technology of China(电子科学与技术大学) Southwestern University of Finance and Economics(西南财经大学) Tongji University(同济大学)

专题命中 视觉推理 :multimodal large language model(title,abstract);分类 cs.CV

Comments 19 pages, 11 figures. Accepted by the 39th Conference on Neural Information Processing Systems (NeurIPS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.12795 2025-11-07 cs.CV 79%

EarthGPT-X: A Spatial MLLM for Multi-level Multi-Source Remote Sensing Imagery Understanding with Visual Prompting

Wei Zhang, Miaoxin Cai, Yaqian Ning, Tong Zhang, Yin Zhuang, Shijian Lu, He Chen, Jun Li, Xuerui Mao

机构 * School of Interdisciplinary Science, Beijing Institute of Technology(交叉科学学院,北京理工大学) College of Computing and Data Science, Nanyang Technological University(计算与数据科学学院,南洋理工大学) National Key Laboratory of Science and Technology on Space-Born Intelligent Information Processing, Beijing Institute of Technology(空间智能信息处理国家重点实验室,北京理工大学) School of Optics and Photonics, Beijing Institute of Technology(光学与 photonics 学院,北京理工大学) State Key Laboratory of Explosion Science and Safety Protection, Beijing(爆炸科学与安全防护国家重点实验室,北京)

专题命中 视觉推理 :MLLM(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03908 2025-11-07 cs.CL 78%

Context informs pragmatic interpretation in vision-language models

Alvin Wei Ming Tan, Ben Prystawski, Veronica Boyce, Michael C. Frank

机构 * Department of Psychology Stanford University(心理学系 斯坦福大学)

专题命中 视觉推理 :vision-language model(title,abstract)

Comments Accepted at CogInterp Workshop, NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03845 2025-11-07 cs.AI cs.LG 73%

To See or To Read: User Behavior Reasoning in Multimodal LLMs

Tianning Dong, Luyi Ma, Varun Vasudevan, Jason Cho, Sushant Kumar, Kannan Achan

机构 * Personalization Team, Walmart Global Tech(Walmart全球科技个性化团队)

专题命中 视觉推理 :multimodal large language model(abstract);MLLM(abstract);分类 cs.AI、cs.LG

Comments Accepted by the 39th Conference on Neural Information Processing Systems (NeurIPS 2025) Workshop: Efficient Reasoning

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18751 2025-11-07 cs.AI cs.CV 73%

Seg the HAB: Language-Guided Geospatial Algae Bloom Reasoning and Segmentation

Patterson Hsieh, Jerry Yeh, Mao-Chi He, Wen-Han Hsieh, Elvis Hsieh

机构 * UC San Diego(加州大学圣地亚哥分校) UC Berkeley(加州大学伯克利分校)

专题命中 视觉推理 :vision-language model(abstract);vision language model(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01833 2025-11-07 cs.CV 70%

TIR-Bench: A Comprehensive Benchmark for Agentic Thinking-with-Images Reasoning

Ming Li, Jike Zhong, Shitian Zhao, Haoquan Zhang, Shaoheng Lin, Yuxiang Lai, Chen Wei, Konstantinos Psounis, Kaipeng Zhang

机构 * Shanghai AI Laboratory(上海人工智能实验室) University of Southern California(南加州大学) Emory University(埃默里大学) Chinese University of Hong Kong(香港中文大学) Rice University(德克萨斯大学奥斯汀分校)

专题命中 视觉推理 :visual reasoning(abstract);multimodal large language model(abstract);分类 cs.CV

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04664 2025-11-07 cs.RO 50%

SAFe-Copilot: Unified Shared Autonomy Framework

Phat Nguyen, Erfan Aasi, Shiva Sreeram, Guy Rosman, Andrew Silva, Sertac Karaman, Daniela Rus

机构 * MIT CSAIL(麻省理工学院计算机科学与人工智能实验室) TRI(丰田研究机构) MIT LIDS(麻省理工学院领导力与决策科学实验室)

专题命中 视觉推理 :vision language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视觉定位与Grounding 6 篇

2511.03757 2025-11-07 cs.LG cs.AI 73%

Laugh, Relate, Engage: Stylized Comment Generation for Short Videos

Xuan Ouyang, Senan Wang, Bouzhou Wang, Siyuan Xiahou, Jinrong Zhou, Yuekang Li

机构 * University of New South Wales(新南威尔士大学) University of Sydney(悉尼大学) The University of Hong Kong(香港大学) University of Southern California(南加州大学)

专题命中 视觉定位与Grounding :multimodal large language model(abstract);MLLM(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01019 2025-11-07 cs.CL cs.AI cs.CE cs.LG physics.ao-ph 62%

OceanAI: A Conversational Platform for Accurate, Transparent, Near-Real-Time Oceanographic Insights

Bowen Chen, Jayesh Gajbhar, Gregory Dusek, Rob Redmon, Patrick Hogan, Paul Liu, DelWayne Bohnenstiehl, Dongkuan Xu, Ruoying He

机构 * North Carolina State University(北卡罗来纳州立大学) NOAA(国家海洋和大气管理局)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG

Comments A related presentation will be given at the AGU(American Geophysical Union) and AMS(American Meteorological Society) Annual Meetings

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04590 2025-11-07 cs.LG cs.IT math.IT 57%

Complexity as Advantage: A Regret-Based Perspective on Emergent Structure

Oshri Naparstek

机构 * IBM Research(IBM研究院)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG

Comments 15 pages. Under preparation for submission to ICML 2026. Feedback welcome

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23769 2025-11-07 cs.CV 57%

TextRegion: Text-Aligned Region Tokens from Frozen Image-Text Models

Yao Xiao, Qiqian Fu, Heyi Tao, Yuqun Wu, Zhen Zhu, Derek Hoiem

机构 * Siebel School of Computing and Data Science(塞比尔计算与数据科学学院) University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments Published in TMLR, with a J2C Certification

Journal ref Transactions on Machine Learning Research, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.04847 2025-11-07 cs.CL cs.AI 57%

Benchmarking LLM Faithfulness in RAG with Evolving Leaderboards

Manveer Singh Tamber, Forrest Sheng Bao, Chenyu Xu, Ge Luo, Suleman Kazi, Minseok Bae, Miaoran Li, Ofer Mendelevitch, Renyi Qu, Jimmy Lin

机构 * University of Waterloo(滑铁卢大学) Vectara(Vectara公司) Iowa State University(爱荷华州立大学) Stanford University(斯坦福大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments EMNLP Industry Track 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04112 2025-11-07 cs.CV 57%

SpatialLock: Precise Spatial Control in Text-to-Image Synthesis

Biao Liu, Yuanzhi Liang

机构 * The Sugon Group(神舟集团) TeleAI China Telecom(中国电信) Shanghai China(上海中国)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments Work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏

4. GUI与屏幕智能体 2 篇

2511.04137 2025-11-07 cs.CV cs.AI 62%

Learning from Online Videos at Inference Time for Computer-Use Agents

Yujian Liu, Ze Wang, Hao Chen, Ximeng Sun, Xiaodong Yu, Jialian Wu, Jiang Liu, Emad Barsoum, Zicheng Liu, Shiyu Chang

机构 * UC Santa Barbara(UC圣芭芭拉大学) Advanced Micro Devices, Inc.(先进微器件公司)

专题命中 GUI与屏幕智能体 :VLM(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02728 2025-11-07 cs.RO 50%

Team Xiaomi EV-AD VLA: Caption-Guided Retrieval System for Cross-Modal Drone Navigation -- Technical Report for IROS 2025 RoboSense Challenge Track 4

Lingfeng Zhang, Erjia Xiao, Yuchen Zhang, Haoxiang Fu, Ruibin Hu, Yanbiao Ma, Wenbo Ding, Long Chen, Hangjun Ye, Xiaoshuai Hao

机构 * Tsinghua University(清华大学) Xiaomi EV(小米电动车) Georgia Institute of Technology(佐治亚理工学院) National University of Singapore(新加坡国立大学) The Chinese University of Hong Kong(香港中文大学) Renmin University of China(中国人民大学)

专题命中 GUI与屏幕智能体 :VLM(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 幻觉与鲁棒性 4 篇

2410.04514 2025-11-07 cs.CL cs.CV 70%

DAMRO: Dive into the Attention Mechanism of LVLM to Reduce Object Hallucination

Xuan Gong, Tianshi Ming, Xinpeng Wang, Zhihua Wei

机构 * Department of Computer Science and Technology, Tongji University(计算机科学与技术系,同济大学)

专题命中 幻觉与鲁棒性 :vision-language model(abstract);LLaVA(abstract);分类 cs.CV

Comments Accepted by EMNLP2024 (Main Conference), add GitHub link

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15482 2025-11-07 cs.CV cs.AI 62%

Comparing Computational Pathology Foundation Models using Representational Similarity Analysis

Vaibhav Mishra, William Lotter

机构 * Dana-Farber Cancer Institute(达纳-法伯癌症研究所) Brigham and Women’s Hospital & Harvard Medical School(布里奇沃特医院及哈佛医学院)

专题命中 幻觉与鲁棒性 :vision-language model(abstract);分类 cs.CV、cs.AI

Comments Proceedings of the 5th Machine Learning for Health (ML4H) Symposium

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04655 2025-11-07 cs.CV 61%

Benchmark Designers Should "Train on the Test Set" to Expose Exploitable Non-Visual Shortcuts

Ellis Brown, Jihan Yang, Shusheng Yang, Rob Fergus, Saining Xie

机构 * New York University(纽约大学)

专题命中 幻觉与鲁棒性 :multimodal large language model(abstract);分类 cs.CV;MLLM(comments)

Comments Project page: https://cambrian-mllm.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03900 2025-11-07 cs.CL cs.LG 57%

GRAD: Graph-Retrieved Adaptive Decoding for Hallucination Mitigation

Manh Nguyen, Sunil Gupta, Dai Do, Hung Le

机构 * Applied Artificial Intelligence Initiative(应用人工智能计划) Deakin University(德金大学)

专题命中 幻觉与鲁棒性 :grounding(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏

6. VLM训练与架构 8 篇

2509.10641 2025-11-07 cs.LG cs.AI 84%

Test-Time Warmup for Multimodal Large Language Models

Nikita Rajaneesh, Thomas Zollo, Richard Zemel

机构 * Columbia University(哥伦比亚大学)

专题命中 VLM训练与架构 :multimodal large language model(title,abstract);MLLM(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21849 2025-11-07 cs.LG cs.AI 81%

TowerVision: Understanding and Improving Multilinguality in Vision-Language Models

André G. Viveiros, Patrick Fernandes, Saul Santos, Sonal Sannigrahi, Emmanouil Zaranis, Nuno M. Guerreiro, Amin Farajian, Pierre Colombo, Graham Neubig, André F. T. Martins

机构 * Instituto Superior Técnico, Universidade de Lisboa(里斯本大学技术高级学院) Instituto de Telecomunicações(电信研究所) Carnegie Mellon University(卡内基梅隆大学) Sword Health(Sword健康) TransPerfect MICS, CentraleSupélec, Université Paris-Saclay(MICS,中央圣艾尔布里大学,巴黎萨克雷大学) ELLIS Unit Lisbon(里斯本ELLIS单位)

专题命中 VLM训练与架构 :vision-language model(title,abstract);分类 cs.AI、cs.LG

Comments 15 pages, 7 figures, submitted to arXiv October 2025. All models, datasets, and training code will be released at https://huggingface.co/collections/utter-project/towervision

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04012 2025-11-07 cs.SE 78%

PSD2Code: Automated Front-End Code Generation from Design Files via Multimodal Large Language Models

Yongxi Chen, Lei Chen

专题命中 VLM训练与架构 :multimodal large language model(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11234 2025-11-07 cs.RO cs.CV 70%

Poutine: Vision-Language-Trajectory Pre-Training and Reinforcement Learning Post-Training Enable Robust End-to-End Autonomous Driving

Luke Rowe, Rodrigue de Schaetzen, Roger Girgis, Christopher Pal, Liam Paull

机构 * Mila - Quebec AI Institute(魁北克人工智能研究所) Université de Montréal(蒙特利尔大学) Polytechnique Montréal(蒙特利尔理工学院) CIFAR AI Chair(CIFAR人工智能主席)

专题命中 VLM训练与架构 :vision-language model(abstract);VLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.07416 2025-11-07 cs.CV cs.CL cs.LG 62%

RadZero: Similarity-Based Cross-Attention for Explainable Vision-Language Alignment in Chest X-ray with Zero-Shot Multi-Task Capability

Jonggwon Park, Byungmu Yoon, Soobum Kim, Kyoyun Choi

专题命中 VLM训练与架构 :grounding(abstract);分类 cs.CV、cs.LG

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04601 2025-11-07 cs.CV cs.MM 57%

PixCLIP: Achieving Fine-grained Visual Language Understanding via Any-granularity Pixel-Text Alignment Learning

Yicheng Xiao, Yu Chen, Haoxuan Ma, Jiale Hong, Caorui Li, Lingxiang Wu, Haiyun Guo, Jinqiao Wang

专题命中 VLM训练与架构 :multimodal large language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏