arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-10-20 至 2025-10-20 共收录 3 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 3 篇

2510.15866 2025-10-20 cs.CV cs.NE 83%

BiomedXPro: Prompt Optimization for Explainable Diagnosis with Biomedical Vision Language Models

Kaushitha Silva, Mansitha Eashwara, Sanduni Ubayasiri, Ruwan Tennakoon, Damayanthi Herath

机构 * University of Peradeniya(珀德尼亚大学) RMIT University(皇家墨尔本理工大学)

专题命中 视觉定位与Grounding :vision language model(title);vision-language model(abstract);grounding(abstract);分类 cs.CV

Comments 10 Pages + 15 Supplementary Material Pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.13564 2025-10-20 cs.CV cs.AI 62%

HumorDB: Can AI understand graphical humor?

Vedaant Jain, Felipe dos Santos Alves Feitosa, Gabriel Kreiman

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) University of São Paulo(圣保罗大学) Harvard Medical School(哈佛医学院)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.AI

Comments 10 main figures, 4 additional appendix figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04232 2025-10-20 cs.LG 57%

Rethinking Layer-wise Gaussian Noise Injection: Bridging Implicit Objectives and Privacy Budget Allocation

Qifeng Tan, Shusen Yang, Xuebin Ren, Yikai Zhang

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG

Comments Errors were found in the experimental data preprocessing, which affected the reported results and conclusions. The paper is being revised and a corrected version will be resubmitted

详情

展开后加载摘要…

URL PDF HTML 收藏