arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-10-31 至 2025-10-31 共收录 8 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 8 篇

2510.26151 2025-10-31 cs.CV cs.AI 73%

MV-MLM: Bridging Multi-View Mammography and Language for Breast Cancer Diagnosis and Risk Prediction

Shunjie-Fabian Zheng, Hyeonjun Lee, Thijs Kooi, Ali Diba

机构 * Department of Medicine I, LMU University Hospital, LMU Munich, Germany(慕尼黑大学医学部第一部门,LMU大学医院,慕尼黑,德国) Lunit Inc.(Lunit公司)

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);分类 cs.CV、cs.AI

Comments Accepted to Computer Vision for Automated Medical Diagnosis (CVAMD) Workshop at ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26006 2025-10-31 cs.CV cs.CL 70%

CAVE: Detecting and Explaining Commonsense Anomalies in Visual Environments

Rishika Bhagwatkar, Syrielle Montariol, Angelika Romanou, Beatriz Borges, Irina Rish, Antoine Bosselut

机构 * EPFL(瑞士联邦理工学院) MILA(蒙特利尔人工智能研究院)

专题命中 视觉定位与Grounding :vision-language model(abstract);grounding(abstract);分类 cs.CV

Journal ref 2025 Conference on Empirical Methods in Natural Language Processing

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.02820 2025-10-31 cs.AI cs.CL cs.LG 62%

AutoLibra: Agent Metric Induction from Open-Ended Human Feedback

Hao Zhu, Phil Cuvin, Xinkai Yu, Charlotte Ka Yee Yan, Jason Zhang, Diyi Yang

机构 * Stanford University(斯坦福大学) University of Toronto(多伦多大学) University of Pennsylvania(宾夕法尼亚大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG

Comments https://github.com/Open-Social-World/autolibra

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16556 2025-10-31 cs.CV 57%

Fit for Purpose? Deepfake Detection in the Real World

Guangyu Lin, Li Lin, Christina P. Walker, Daniel S. Schiff, Shu Hu

专题命中 视觉定位与Grounding :multimodal large language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26527 2025-10-31 cs.LG 57%

Polybasic Speculative Decoding Through a Theoretical Perspective

Ruilin Wang, Huixia Li, Yuexiao Ma, Xiawu Zheng, Fei Chao, Xuefeng Xiao, Rongrong Ji

机构 * Key Laboratory of Multimedia Trusted Perception(多媒体可信感知关键实验室) Efficient Computing, Ministry of Education of China, Xiamen University(高效计算、教育部中国 ministry of education、厦门大学) Institute of Artificial Intelligence, Xiamen University(人工智能研究院、厦门大学) Peng Cheng Laboratory, Shenzhen, China(鹏城实验室、深圳中国)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26464 2025-10-31 cs.CV 57%

Towards Fine-Grained Vision-Language Alignment for Few-Shot Anomaly Detection

Yuanting Fan, Jun Liu, Xiaochen Chen, Bin-Bin Gao, Jian Li, Yong Liu, Jinlong Peng, Chengjie Wang

机构 * Tencent Youtu Lab(腾讯优图实验室)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments 12 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04448 2025-10-31 cs.CV cs.MM 57%

TRUST-VL: An Explainable News Assistant for General Multimodal Misinformation Detection

Zehong Yan, Peng Qi, Wynne Hsu, Mong Li Lee

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments EMNLP 2025 Oral; Project Homepage: https://yanzehong.github.io/trust-vl/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26015 2025-10-31 cs.HC 50%

Designing for Dignity while Driving: Interaction Needs of Blind and Low-Vision Passengers in Fully Automated Vehicles

Zhengtao Ma, Rafael Gomez, Togtokhtur Batbold, Zishuo Zhu, Yueteng Yu, Ronald Schroeter

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏