arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-08-29 至 2025-08-29 共收录 8 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 8 篇

2508.20188 2025-08-29 cs.CV cs.LG 90%

Grounding Multimodal Large Language Models with Quantitative Skin Attributes: A Retrieval Study

Max Torop, Masih Eskandar, Nicholas Kurtansky, Jinyang Liu, Jochen Weber, Octavia Camps, Veronica Rotemberg, Jennifer Dy, Kivanc Kose

机构 * Northeastern University(东北大学) Memorial Sloan Kettering Cancer Center(纪念斯隆凯特琳癌症中心)

专题命中 视觉定位与Grounding :grounding(title,abstract);multimodal large language model(title,abstract);MLLM(abstract);分类 cs.CV、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20758 2025-08-29 cs.CV cs.AI 88%

SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding

Jiawen Lin, Shiran Bian, Yihang Zhu, Wenbin Tan, Yachao Zhang, Yuan Xie, Yanyun Qu

机构 * School of Informatics, Xiamen University(厦门大学信息学院) School of Computer Science, Nanjing University(南京大学计算机科学学院) School of Computer Science and Technology, East China Normal University(华东师范大学计算机科学与技术学院) Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University(教育部多媒体可信感知与高效计算重点实验室,厦门大学)

专题命中 视觉定位与Grounding :VLM(title,abstract);grounding(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20279 2025-08-29 cs.CV cs.AI cs.CL 86%

How Multimodal LLMs Solve Image Tasks: A Lens on Visual Grounding, Task Reasoning, and Answer Decoding

Zhuoran Yu, Yong Jae Lee

机构 * University of Wisconsin–Madison(威斯康星大学麦迪逊分校)

专题命中 视觉定位与Grounding :grounding(title,abstract);LLaVA(abstract);multimodal large language model(abstract);分类 cs.CV、cs.AI

Comments Accepted by COLM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20830 2025-08-29 cs.CV 83%

Estimating 2D Keypoints of Surgical Tools Using Vision-Language Models with Low-Rank Adaptation

Krit Duangprom, Tryphon Lambrou, Binod Bhattarai

机构 * University of Aberdeen(阿伯丁大学)

专题命中 视觉定位与Grounding :vision-language model(title);vision language model(abstract);VLM(abstract);分类 cs.CV

Comments Accepted to MICCAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.19573 2025-08-29 cs.CV cs.AI 81%

See then Tell: Enhancing Key Information Extraction with Vision Grounding

Shuhang Liu, Zhenrong Zhang, Pengfei Hu, Jiefeng Ma, Jun Du, Qing Wang, Jianshu Zhang, Chenyu Liu

机构 * University of Science and Technology of China(科学技术大学) iFLYTEK Research(iFLYTEK研究院)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.01401 2025-08-29 cs.CV 79%

Language-to-Space Programming for Training-Free 3D Visual Grounding

Boyu Mi, Hanqing Wang, Tai Wang, Yilun Chen, Jiangmiao Pang

机构 * Shanghai Jiao Tong University(上海交通大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20976 2025-08-29 cs.SD cs.AI eess.AS 57%

WoW-Bench: Evaluating Fine-Grained Acoustic Perception in Audio-Language Models via Marine Mammal Vocalizations

Jaeyeon Kim, Heeseung Yun, Sang Hoon Woo, Chao-Han Huck Yang, Gunhee Kim

机构 * Carnegie Mellon University(卡内基梅隆大学) Seoul National University(首尔国立大学) NVIDIA(NVIDIA公司)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments Preprint. Project page: https://jaeyeonkim99.github.io/wow_bench/

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01970 2025-08-29 cs.LG 57%

Improving Hospital Risk Prediction with Knowledge-Augmented Multimodal EHR Modeling

Rituparna Datta, Jiaming Cui, Zihan Guan, Vishal G. Reddy, Joshua C. Eby, Gregory Madden, Rupesh Silwal, Anil Vullikanti

机构 * Department of Computer Science, University of Virginia(大学计算机科学系) University of Virginia School of Medicine(弗吉尼亚大学医学院) Virginia Polytechnic Institute and State University(弗吉尼亚理工学院和州立大学) Biocomplexity Institute and Initiative, University of Virginia(大学生物复杂性研究所) Division of Infectious Diseases & International Health, University of Virginia School of Medicine(大学感染性疾病与国际卫生分会)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏