arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-09-03 至 2025-09-03 共收录 16 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 16 篇

2509.01554 2025-09-03 cs.CV cs.AI cs.LG 85%

Unified Supervision For Vision-Language Modeling in 3D Computed Tomography

Hao-Chih Lee, Zelong Liu, Hamza Ahmed, Spencer Kim, Sean Huver, Vishwesh Nath, Zahi A. Fayad, Timothy Deyer, Xueyan Mei

专题命中 视觉定位与Grounding :vision-language model(title,abstract);VLM(abstract,comments);分类 cs.CV、cs.AI、cs.LG

Comments ICCV 2025 VLM 3d Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16647 2025-09-03 cs.CV cs.AI 81%

Point, Detect, Count: Multi-Task Medical Image Understanding with Instruction-Tuned Vision-Language Models

Sushant Gautam, Michael A. Riegler, Pål Halvorsen

机构 * Simula Metropolitan Center for Digital Engineering (SimulaMet), Norway(Simula数字工程中心(SimulaMet)) Oslo Metropolitan University (OsloMet), Norway(奥斯陆 Metropolitan 大学(OsloMet)) Simula Research Laboratory, Norway(Simula研究实验室)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV、cs.AI

Comments Accepted as a full paper at the 38th IEEE International Symposium on Computer-Based Medical Systems (CBMS) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02324 2025-09-03 cs.RO 75%

Language-Guided Long Horizon Manipulation with LLM-based Planning and Visual Perception

Changshi Zhou, Haichuan Xu, Ningquan Gu, Zhipeng Wang, Bin Cheng, Pengpeng Zhang, Yanchao Dong, Mitsuhiro Hayashibe, Yanmin Zhou, Bin He

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.13919 2025-09-03 cs.CV cs.AI cs.CL cs.LG cs.RO 75%

Temporal Preference Optimization for Long-Form Video Understanding

Rui Li, Xiaohan Wang, Yuhui Zhang, Orr Zohar, Zeyu Wang, Serena Yeung-Levy

机构 * Stanford University(斯坦福大学)

专题命中 视觉定位与Grounding :LLaVA(abstract);grounding(abstract);分类 cs.CV、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00284 2025-09-03 cs.CV cs.AI 73%

Generative AI for Industrial Contour Detection: A Language-Guided Vision System

Liang Gong, Tommy, Wang, Sara Chaker, Yanchen Dong, Fouad Bousetouane, Brenden Morton, Mark Mendez

机构 * The University of Chicago(芝加哥大学) FabTrack

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);分类 cs.CV、cs.AI

Comments 20 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.16680 2025-09-03 cs.CV 70%

AeroReformer: Aerial Referring Transformer for UAV-based Referring Image Segmentation

Rui Li, Xiaowei Zhao

机构 * Intelligent Control \& Smart Energy (ICSE) Research Group, School of Engineering, University of Warwick, Coventry, CV4 7AL, UK

专题命中 视觉定位与Grounding :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04549 2025-09-03 cs.CV cs.AI cs.MM 62%

MSC: A Marine Wildlife Video Dataset with Grounded Segmentation and Clip-Level Captioning

Quang-Trung Truong, Yuk-Kwan Wong, Vo Hoang Kim Tuyen Dang, Rinaldi Gotama, Duc Thanh Nguyen, Sai-Kit Yeung

机构 * The Hong Kong University of Science and Technology(香港科学与技术大学) Ho Chi Minh University of Science(胡志明市科学大学) Deakin University(德肯大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI

Comments Published at ACMMM2025 (Dataset track)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01716 2025-09-03 cs.AI cs.CL 57%

An LLM-enabled semantic-centric framework to consume privacy policies

Rui Zhao, Vladyslav Melnychuk, Jun Zhao, Jesse Wright, Nigel Shadbolt

机构 * University of Oxford(牛津大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.21484 2025-09-03 q-bio.QM cs.LG stat.ML 57%

Data-driven Discovery of Digital Twins in Biomedical Research

Clémence Métayer, Annabelle Ballesta, Julien Martinelli

机构 * Inserm U1331, Institut Curie Saint-Cloud, France(法国国家医学研究院U1331,圣克鲁医院) Aalto University Espoo, Finland(芬兰艾尔托大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10292 2025-09-03 cs.CV cs.CL 57%

StoryReasoning Dataset: Using Chain-of-Thought for Scene Understanding and Grounded Story Generation

Daniel A. P. Oliveira, David Martins de Matos

机构 * Instituto Superior Técnico, Universidade de Lisboa(里斯本大学技术学院) INESC-ID Lisboa(里斯本INESC-ID)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments 31 pages, 14 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.00600 2025-09-03 cs.DB cs.AI cs.CL 57%

Semantic Integrity Constraints: Declarative Guardrails for AI-Augmented Data Processing Systems

Alexander W. Lee, Justin Chan, Michael Fu, Nicolas Kim, Akshay Mehta, Deepti Raghavan, Ugur Cetintemel

机构 * Brown University(布朗大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Journal ref PVLDB, 18(11): 4073 - 4080, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.07268 2025-09-03 cs.MM cs.CL cs.CV 57%

Advancing Grounded Multimodal Named Entity Recognition via LLM-Based Reformulation and Box-Based Segmentation

Jinyuan Li, Ziyan Li, Han Li, Jianfei Yu, Rui Xia, Di Sun, Gang Pan

机构 * College of Intelligence and Computing, Tianjin University(智能与计算学院,天津大学) NJUST(南京理工大学) College of Mathematics, Taiyuan University of Technology(数学学院,太原科技大学) Tianjin University of Science and Technology(天津科技大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments Extension of our Findings of EMNLP 2023 & ACL 2024 paper, IEEE Transactions on Multimedia accepted on July 19, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02425 2025-09-03 cs.RO cs.HC 50%

OpenGuide: Assistive Object Retrieval in Indoor Spaces for Individuals with Visual Impairments

Yifan Xu, Qianwei Wang, Vineet Kamat, Carol Menassa

机构 * Department of Civil and Environmental Engineering, University of Michigan(密歇根大学土木与环境工程系) College of Literature, Science, and the Arts, University of Michigan(密歇根大学文学、科学与艺术学院)

专题命中 视觉定位与Grounding :VLM(abstract)

Comments 32 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02367 2025-09-03 cs.HC 50%

Talking Spell: A Wearable System Enabling Real-Time Anthropomorphic Voice Interaction with Everyday Objects

Xuetong Wang, Ching Christie Pang, Pan Hui

专题命中 视觉定位与Grounding :vision-language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18906 2025-09-03 cs.CL cs.IR 50%

Federated Retrieval-Augmented Generation: A Systematic Mapping Study

Abhijit Chakraborty, Chahana Dahal, Vivek Gupta

机构 * Arizona State University(亚利桑那州立大学) University of Nevada, Las Vegas(内华达大学拉斯维加斯分校)

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00325 2025-09-03 cs.CL cs.IR 50%

GIER: Gap-Driven Self-Refinement for Large Language Models

Rinku Dewri

机构 * University of Denver(丹佛大学)

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏