arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-11-19 至 2025-11-19 共收录 10 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 10 篇

2412.15925 2025-11-19 cs.CV cs.AI 84%

MiniGPT-Pancreas: Multimodal Large Language Model for Pancreas Cancer Classification and Detection

Andrea Moglia, Elia Clement Nastasio, Luca Mainardi, Pietro Cerveri

机构 * Department of Electronics, Information, and Bioengineering(电子、信息与生物工程系) Polytechnic University of Milan(米兰理工学院) Department of Industrial, and Information Engineering(工业与信息工程系) University of Pavia(帕维亚大学)

专题命中 视觉定位与Grounding :multimodal large language model(title,abstract);MLLM(abstract);分类 cs.CV、cs.AI

Journal ref Moglia, A., Nastasio, E.C., Mainardi, L. et al. MiniGPT-Pancreas: Multimodal Large Language Model for Pancreas Cancer Observation and Localization in CT Images. J Healthc Inform Res (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14086 2025-11-19 cs.CV cs.AI cs.CL 81%

Error-Driven Scene Editing for 3D Grounding in Large Language Models

Yue Zhang, Zun Wang, Han Lin, Jialu Li, Jianing Yang, Yonatan Bitton, Idan Szpektor, Mohit Bansal

机构 * UNC Chapel Hill(北卡罗来纳大学教堂山分校) University of Michigan(密歇根大学) Google Research(谷歌研究)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV、cs.AI

Comments Code: https://github.com/zhangyuejoslin/Deer-3D

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13924 2025-11-19 cs.CV 79%

Start Small, Think Big: Curriculum-based Relative Policy Optimization for Visual Grounding

Qingyang Yan, Guangyao Chen, Yixiong Zou

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments AAAI 2026 (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13442 2025-11-19 cs.CV cs.AI 73%

Unlocking the Forgery Detection Potential of Vanilla MLLMs: A Novel Training-Free Pipeline

Rui Zuo, Qinyue Tong, Zhe-Ming Lu, Ziqian Lu

机构 * Zhejiang University(浙江大学) Zhejiang Sci-Tech University(浙江科技学院)

专题命中 视觉定位与Grounding :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05430 2025-11-19 cs.CV cs.AI cs.LG 67%

Explaining Similarity in Vision-Language Encoders with Weighted Banzhaf Interactions

Hubert Baniecki, Maximilian Muschalik, Fabian Fumagalli, Barbara Hammer, Eyke Hüllermeier, Przemyslaw Biecek

机构 * University of Warsaw(华沙大学) Warsaw University of Technology(华沙理工大学) LMU Munich(慕尼黑大学) MCML DFKI(德累斯顿大学) Bielefeld University(比勒菲尔德大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.AI、cs.LG

Comments NeurIPS 2025. Code: https://github.com/hbaniecki/fixlip

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14027 2025-11-19 cs.CL 67%

HiEAG: Evidence-Augmented Generation for Out-of-Context Misinformation Detection

Junjie Wu, Yumeng Fu, Nan Yu, Guohong Fu

机构 * School of Computer Science and Technology, Soochow University(苏州大学计算机科学与技术学院) Institute of Artificial Intelligence, Soochow University(苏州大学人工智能研究院) School of Computer Science and Technology, Harbin Institute of Technology(哈尔滨工业大学计算机科学与技术学院)

专题命中 视觉定位与Grounding :multimodal large language model(abstract);MLLM(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13759 2025-11-19 cs.LG cs.AI 62%

Multi-Agent VLMs Guided Self-Training with PNU Loss for Low-Resource Offensive Content Detection

Han Wang, Deyi Ji, Junyu Lu, Lanyun Zhu, Hailong Zhang, Haiyang Wu, Liqun Liu, Peng Shu, Roy Ka-Wei Lee

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.AI、cs.LG

Comments 8 pages, 4 figures, Fortieth AAAI Conference on Artificial Intelligence (AAAI-26)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19082 2025-11-19 cs.CV 57%

Sa2VA-i: Improving Sa2VA Results with Consistent Training and Inference

Alexey Nekrasov, Ali Athar, Daan de Geus, Alexander Hermans, Bastian Leibe

机构 * RWTH Aachen University(亚琛工业大学) Eindhoven University of Technology(埃因霍温理工大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11777 2025-11-19 cs.RO cs.CV 57%

Large Language Models and 3D Vision for Intelligent Robotic Perception and Autonomy

Vinit Mehta, Charu Sharma, Karthick Thiyagarajan

机构 * Machine Learning Lab IIIT Hyderabad(IIIT Hyderabad 机器学习实验室) SensR Lab Western Sydney University(Western Sydney University SensR实验室)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments 45 pages, 15 figures, MDPI Sensors Journal

Journal ref Sensors 2025, 25(20), 6394

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09030 2025-11-19 cs.LG 57%

Contextual Learning for Anomaly Detection in Tabular Data

Spencer King, Zhilu Zhang, Ruofan Yu, Baris Coskun, Wei Ding, Qian Cui

机构 * Amazon Web Services, Seattle, WA, USA(亚马逊网络服务,西雅图,WA,USA)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG

Comments Submitted to TMLR. 26 pages, 4 figures, 8 tables, 1 algorithm, 8 datasets, contextual anomaly detection framework for tabular data

详情

展开后加载摘要…

URL PDF HTML 收藏