arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-10-13 至 2025-10-13 共收录 11 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 11 篇

2502.03333 2025-10-13 cs.CV cs.AI 84%

RadVLM: A Multitask Conversational Vision-Language Model for Radiology

Nicolas Deperrois, Hidetoshi Matsuo, Samuel Ruipérez-Campillo, Moritz Vandenhirtz, Sonia Laguna, Alain Ryser, Koji Fujimoto, Mizuho Nishio, Thomas M. Sutter, Julia E. Vogt, Jonas Kluckert, Thomas Frauenfelder, Christian Blüthgen, Farhad Nooralahzadeh, Michael Krauthammer

机构 * Department of Radiology, Kobe University(金泽大学放射科) Department of Computer Science, ETH Zurich(苏黎世联邦理工学院计算机科学系) Department of Advanced Imaging in Medical Magnetic Resonance, Kyoto University(京都大学医学磁共振高级成像部门) Department of Quantitative Biomedicine, University of Zurich(苏黎世大学定量生物医学系) Diagnostic and Interventional Radiology, University Hospital Zurich(苏黎世大学医院诊断与介入放射科)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);grounding(abstract);分类 cs.CV、cs.AI

Comments 21 pages, 15 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05404 2025-10-13 cs.CV cs.AI 81%

AD-EE: Early Exiting for Fast and Reliable Vision-Language Models in Autonomous Driving

Lianming Huang, Haibo Hu, Yufei Cui, Jiacheng Zuo, Shangyu Wu, Nan Guan, Chun Jason Xue

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV、cs.AI

Comments We believe that the contribution of this paper is not enough, so we integrated it into another new paper. The arXiv ID of the new paper is arXiv:2510.01795

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.21447 2025-10-13 cs.CV cs.AI 73%

Multimodal Language Models See Better When They Look Shallower

Haoran Chen, Junyan Lin, Xinghao Chen, Yue Fan, Jianfeng Dong, Xin Jin, Hui Su, Jinlan Fu, Xiaoyu Shen

机构 * Zhejiang Gongshang University(浙江工商大学) Ningbo Key Laboratory of Spatial Intelligence and Digital Derivative(宁波空间智能与数字衍生关键实验室) Institute of Digital Twin, Eastern Institute of Technology, Ningbo(数字孪生研究院,东部技术研究所,宁波) Meituan Inc.(美团公司) National University of Singapore(新加坡国立大学)

专题命中 视觉定位与Grounding :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI

Comments 9 pages, 6 figures, accepted by EMNLP2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08978 2025-10-13 cs.CV 70%

HandEval: Taking the First Step Towards Hand Quality Evaluation in Generated Images

Zichuan Wang, Bo Peng, Songlin Yang, Zhenchen Tang, Jing Dong

专题命中 视觉定位与Grounding :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08775 2025-10-13 cs.CV cs.AI 62%

Re-Identifying Kākā with AI-Automated Video Key Frame Extraction

Paula Maddigan, Andrew Lensen, Rachael C. Shaw

机构 * Centre for Data Science and Artificial Intelligence, and School of Engineering and Computer Science(数据科学与人工智能中心,工程与计算机科学学院) Victoria University of Wellington(惠灵顿维多利亚大学) School of Biological Sciences(生物科学学院)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09421 2025-10-13 cs.CL cs.AI 57%

On the Representations of Entities in Auto-regressive Large Language Models

Victor Morand, Josiane Mothe, Benjamin Piwowarski

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments Accepted at BlackBoxNLP@EMNLP2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09274 2025-10-13 cs.CV 57%

MomentSeg: Moment-Centric Sampling for Enhanced Video Pixel Understanding

Ming Dai, Sen Yang, Boqiang Duan, Wankou Yang, Jingdong Wang

机构 * Southeast University(东南大学) Baidu VIS(百度视觉)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08849 2025-10-13 cs.CV 57%

FOLK: Fast Open-Vocabulary 3D Instance Segmentation via Label-guided Knowledge Distillation

Hongrui Wu, Zhicheng Gao, Jin Cao, Kelu Yao, Wen Shen, Zhihua Wei

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21787 2025-10-13 cs.CV cs.CL 57%

DeHate: A Stable Diffusion-based Multimodal Approach to Mitigate Hate Speech in Images

Dwip Dalal, Gautam Vashishtha, Anku Rani, Aishwarya Reganti, Parth Patwa, Mohd Sarique, Chandan Gupta, Keshav Nath, Viswanatha Reddy, Vinija Jain, Aman Chadha, Amitava Das, Amit Sheth, Asif Ekbal

机构 * MIT Media Lab, USA(麻省理工学院媒体实验室) Stanford University, USA(斯坦福大学) Amazon GenAI, USA(亚马逊生成人工智能) University of South Carolina, USA(南卡罗来纳大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments Defactify 3 workshop at AAAI 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19114 2025-10-13 cs.CL cs.IR cs.LG 57%

Understanding and Improving Information Preservation in Prompt Compression for LLMs

Weronika Łajewska, Momchil Hardalov, Laura Aina, Neha Anna John, Hang Su, Lluís Màrquez

机构 * University of Stavanger(斯塔万格大学) AWS AI Labs(AWS AI实验室) Technical University of Catalonia (UPC)(加泰罗尼亚理工大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG

Comments Accepted to EMNLP 2025 (Findings), 22 pages, 6 figures, 24 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10168 2025-10-13 stat.ME math.ST stat.TH 50%

Statistical methods: Basic concepts, interpretations, and cautions

Sander Greenland

专题命中 视觉定位与Grounding :grounding(abstract)

Comments 64 pages. For Pigeot I, Ahrens W, eds., Handbook of Epidemiology, 3rd edn. Springer, 2025, Ch. 54-1

详情

展开后加载摘要…

URL PDF HTML 收藏