arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 7441 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7441 篇

2506.21001 2025-06-27 cs.CV 57%

Style-Aligned Image Composition for Robust Detection of Abnormal Cells in Cytopathology

Qiuyi Qi, Xin Li, Ming Kong, Zikang Xu, Bingdi Chen, Qiang Zhu, S Kevin Zhou

机构 * Zhejiang University(浙江大学) University of Science and Technology of China(中国科学技术大学) Suzhou Institute for Advance Research, USTC(中国科学技术大学苏州市先进研究院) The Institute for Biomedical Engineering & Nano Science, Tongji University School of Medicine(同济大学医学院生物医学工程与纳米科学研究院) Zhihui Medical Technology (Shanghai) Co., Ltd.(智辉医疗技术(上海)有限公司)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments MIDL 2025 Oral

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.13898 2025-06-27 cs.CV cs.CL 57%

GroundCap: A Visually Grounded Image Captioning Dataset

Daniel A. P. Oliveira, Lourenço Teodoro, David Martins de Matos

机构 * INESC-ID(INESC-ID研究所)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments 37 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.17726 2025-06-26 cs.CV cs.CL 57%

VICCA: Visual Interpretation and Comprehension of Chest X-ray Anomalies in Generated Report Without Human Feedback

Sayeh Gholipour Picha, Dawood Al Chanti, Alice Caplier

机构 * Univ. Grenoble Alpes(格勒诺布尔阿尔卑斯大学) CNRS(法国国家科学研究中心) Grenoble INP(格勒诺布尔研究所) GIPSA-lab(GIPSA实验室)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Journal ref Machine Learning with Applications, Volume 21, 2025, 100684, ISSN 2666-8270

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.19204 2025-06-25 cs.CV 57%

OpenWildlife: Open-Vocabulary Multi-Species Wildlife Detector for Geographically-Diverse Aerial Imagery

Muhammed Patel, Javier Noa Turnes, Jayden Hsiao, Linlin Xu, David Clausi

机构 * Systems Design Engineering, University of Waterloo(滑铁卢大学系统设计工程学院) Department of Geomatics Engineering,University of Calgary(卡尔加里大学测绘工程系)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18557 2025-06-25 cs.CV 57%

Object-aware Sound Source Localization via Audio-Visual Scene Understanding

Sung Jin Um, Dongjin Kim, Sangmin Lee, Jung Uk Kim

机构 * Kyung Hee University(庆熙大学) KAIST AI(韩国科学技术院人工智能研究所) Sungkyunkwan University(成均馆大学)

专题命中 视觉定位与Grounding :multimodal large language model(abstract);分类 cs.CV

Comments Accepted at CVPR 2025

Journal ref Proceedings of the Computer Vision and Pattern Recognition Conference (CVPR), 2025, pp. 8342-8351

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16058 2025-06-25 cs.CV 57%

Stepping Out of Similar Semantic Space for Open-Vocabulary Segmentation

Yong Liu, SongLi Wu, Sule Bai, Jiahao Wang, Yitong Wang, Yansong Tang

机构 * Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院,清华大学) The University of Hong Kong(香港大学) ByteDance Inc.(字节跳动公司)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.02157 2025-06-25 cs.CV cs.HC 57%

FineCLIPER: Multi-modal Fine-grained CLIP for Dynamic Facial Expression Recognition with AdaptERs

Haodong Chen, Haojian Huang, Junhao Dong, Mingzhe Zheng, Dian Shao

机构 * School of Automation, Northwestern Polytechnical University(西北工业大学自动化学院) The University of Hong Kong(香港大学) Nanyang Technological University(南洋理工大学) School of Computer Science, Northwestern Polytechnical University(西北工业大学计算机学院) Unmanned System Research Institute, Northwestern Polytechnical University(西北工业大学无人系统研究所)

专题命中 视觉定位与Grounding :MLLM(abstract);分类 cs.CV

Comments Accepted to ACM MM 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17285 2025-06-24 cs.IR cs.LG 57%

A Framework for Generating Conversational Recommendation Datasets from Behavioral Interactions

Vinaik Chhetri, Yousaf Reza, Moghis Fereidouni, Srijata Maji, Umar Farooq, AB Siddique

机构 * Louisiana State University(路易斯安那州立大学) Independent Researcher(独立研究者) University of Kentucky(肯塔基大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG

Comments 12 pages, 6 tables,4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16730 2025-06-23 cs.CV 57%

TeSG: Textual Semantic Guidance for Infrared and Visible Image Fusion

Mingrui Zhu, Xiru Chen, Xin Wei, Nannan Wang, Xinbo Gao

机构 * Xidian University(西电大学) Chongqing University of Post and Telecommunications(重庆邮电大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments 11 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10730 2025-06-23 cs.CV 57%

IQE-CLIP: Instance-aware Query Embedding for Zero-/Few-shot Anomaly Detection in Medical Domain

Hong Huang, Weixiang Sun, Zhijian Wu, Jingwen Niu, Donghuan Lu, Xian Wu, Yefeng Zheng

机构 * Westlake University(西湖大学) Simon Fraser University(西蒙·弗雷泽大学) University of Notre Dame(诺特难大学) Shandong University(山东大学) Tencent Jarvis Lab(腾讯Jarvis实验室)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07497 2025-06-23 cs.CV 57%

Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal Consistency

Xiangyu Guo, Zhanqian Wu, Kaixin Xiong, Ziyang Xu, Lijun Zhou, Gangwei Xu, Shaoqing Xu, Haiyang Sun, Bing Wang, Guang Chen, Hangjun Ye, Wenyu Liu, Xinggang Wang

机构 * Huazhong University of Science and Technology(华中科技大学) Xiaomi EV(小米电动车)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.15084 2025-06-19 cs.SE cs.CV cs.HC 57%

An Empirical Study of Bugs in Data Visualization Libraries

Weiqi Lu, Yongqiang Tian, Xiaohan Zhong, Haoyang Ma, Zhenyang Xu, Shing-Chi Cheung, Chengnian Sun

机构 * The Hong Kong University of Science and Technology(香港科学与技术大学) University of Waterloo(滑铁卢大学) University of Waterloo Canada(滑铁卢大学加拿大)

专题命中 视觉定位与Grounding :vision language model(abstract);分类 cs.CV

Comments Proc. ACM Softw. Eng. 2, FSE

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.14936 2025-06-19 cs.AI 57%

CALM: Contextual Analog Logic with Multimodality

Maxwell J. Jacobson, Corey J. Maley, Yexiang Xue

机构 * Purdue University(普渡大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12331 2025-06-17 cs.MA cs.AI 57%

IndoorWorld: Integrating Physical Task Solving and Social Simulation in A Heterogeneous Multi-Agent Environment

Dekun Wu, Frederik Brudy, Bang Liu, Yi Wang

机构 * Université de Montréal & Mila - Quebec AI Institute(蒙特利尔大学及魁北克人工智能研究所) Autodesk Research(Autodesk研究)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10633 2025-06-13 cs.CV 57%

Anatomy-Grounded Weakly Supervised Prompt Tuning for Chest X-ray Latent Diffusion Models

Konstantinos Vilouras, Ilias Stogiannidis, Junyu Yan, Alison Q. O'Neil, Sotirios A. Tsaftaris

机构 * University of Edinburgh(爱丁堡大学) Canon Medical Research Europe Ltd.(佳能医疗欧洲有限公司)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments 14 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2007.02527 2025-06-12 cs.AI 57%

Goal Kernel Planning: Linearly-Solvable Non-Markovian Policies for Logical Tasks with Goal-Conditioned Options

Thomas J. Ringstrom, Mohammadhosein Hasanbeig, Alessandro Abate

机构 * Department of Computer Science, University of Minnesota(明尼苏达大学计算机科学系) Department of Computer Science, University of Oxford(牛津大学计算机科学系)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments 52 Pages total. This is an update to a paper we submitted to a Journal and received reviewer feedback for improvement

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.16124 2025-06-11 cs.AI 57%

On the generalization of learned constraints for ASP solving in temporal domains

Javier Romero, Torsten Schaub, Klaus Strauch

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments 41 pages, 2 figures, Under consideration in Theory and Practice of Logic Programming (TPLP)

Journal ref Theory and Practice of Logic Programming 25 (2025) 197-224

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.08150 2025-06-11 cs.AI cs.LO 57%

Compiling Metric Temporal Answer Set Programming

Arvid Becker, Pedro Cabalar, Martin Diéguez, Javier Romero, Susana Hahn, Torsten Schaub

机构 * A. Becker University of Potsdam, Germany(波恩大学) P. Cabalar University of Corunna, Spain(科鲁纳大学) M. Diéguez LERIA, University of Angers, France(安格里大学) J. Romero University of Potsdam, Germany(波恩大学) S. Hahn University of Potsdam, Germany(波恩大学) T. Schaub University of Potsdam, Germany(波恩大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.14898 2025-06-11 cs.CL cs.AI cs.IR 57%

Retrieval-augmented systems can be dangerous medical communicators

Lionel Wong, Ayman Ali, Raymond Xiong, Shannon Zeijang Shen, Yoon Kim, Monica Agrawal

机构 * MIT CSAIL(麻省理工学院计算机科学与人工智能实验室) Duke University(杜克大学) Stanford University(斯坦福大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments Position paper in Proceedings of the 42 nd International Conference on Machine Learning

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07697 2025-06-10 cs.CV 57%

OpenSplat3D: Open-Vocabulary 3D Instance Segmentation using Gaussian Splatting

Jens Piekenbrinck, Christian Schmidt, Alexander Hermans, Narunas Vaskevicius, Timm Linder, Bastian Leibe

机构 * RWTH Aachen University(亚琛工业大学) Robert Bosch GmbH(博世公司)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.12262 2025-06-10 cs.CL cs.AI 57%

Low-resource Machine Translation: what for? who for? An observational study on a dedicated Tetun language translation service

Raphael Merx, Adérito José Guterres Correia, Hanna Suominen, Ekaterina Vylomova

机构 * School of Computing and Information Systems, The University of Melbourne(墨尔本大学计算与信息学院) Instituto Nacional de Linguística, Dili(迪利国家语言研究所) The Australian National University(澳大利亚国立大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments to be published in LoResMT 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06480 2025-06-10 cs.CV 57%

(LiFT) Lightweight Fitness Transformer: A language-vision model for Remote Monitoring of Physical Training

A. Postlmayr, P. Cosman, S. Dey

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.12606 2025-06-10 cs.CV 57%

ELEV-VISION-SAM: Integrated Vision Language and Foundation Model for Automated Estimation of Building Lowest Floor Elevation

Yu-Hsuan Ho, Longxiang Li, Ali Mostafavi

专题命中 视觉定位与Grounding :vision language model(abstract);分类 cs.CV

Journal ref Comput. Aided Civ. Infrastruct. Eng. 40.1 (2025) 75-90

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.01850 2025-06-09 cs.CV 57%

Foundation Model-Based Apple Ripeness and Size Estimation for Selective Harvesting

Keyi Zhu, Jiajia Li, Kaixiang Zhang, Chaaran Arunachalam, Siddhartha Bhattacharya, Renfu Lu, Zhaojian Li

机构 * Department of Mechanical Engineering, Michigan State University(机械工程系,密歇根州立大学) Department of Electrical and Computer Engineering, Michigan State University(电气与计算机工程系,密歇根州立大学) Department of Computer Science and Engineering, Michigan State University(计算机科学与工程系,密歇根州立大学) United States Department of Agriculture Agricultural Research Service(美国农业部农业研究服务)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.12813 2025-06-09 cs.CV 57%

Described Object Detection: Liberating Object Detection with Flexible Expressions

Chi Xie, Zhao Zhang, Yixuan Wu, Feng Zhu, Rui Zhao, Shuang Liang

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments Accepted by NeurIPS 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05651 2025-06-09 cs.CV 57%

Hallucinate, Ground, Repeat: A Framework for Generalized Visual Relationship Detection

Shanmukha Vellamcheti, Sanjoy Kundu, Sathyanarayanan N. Aakur

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments 22 pages, 9 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.21682 2025-06-09 cs.CV 57%

Visual Text Processing: A Comprehensive Review and Unified Evaluation

Yan Shu, Weichao Zeng, Fangmin Zhao, Zeyu Chen, Zhenhang Li, Xiaomeng Yang, Yu Zhou, Paolo Rota, Xiang Bai, Lianwen Jin, Xu-Cheng Yin, Nicu Sebe

机构 * VCIP & TMCC & DISSec, College of Computer Science, Nankai University(VCIP与TMCC与DISSec,计算机科学学院,南开大学) Department of Information Engineering and Computer Science, University of Trento(信息工程与计算机科学系,特伦托大学) Institute of Information Engineering, Chinese Academy of Sciences(信息工程研究所,中国科学院) School of Cyber Security, University of Chinese Academy of Sciences(网络安全学院,中国科学院大学) Department of Electrical and Computer Engineering, Northeastern University(电气与计算机工程系,东北大学) School of Software Engineering, Huazhong University of Science and Technology(软件工程学院,华中科技大学) School of Electronic and Information Engineering, South China University of Technology(电子与信息工程学院,华南理工大学)

专题命中 视觉定位与Grounding :multimodal large language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.04026 2025-06-05 cs.LG 57%

On the Usage of Gaussian Process for Efficient Data Valuation

Clément Bénesse, Patrick Mesana, Athénaïs Gautier, Sébastien Gambs

机构 * Opsci.ai Paris, France(Opsci.ai巴黎,法国) Université du Québec à Montréal(魁北克大学蒙特利尔分校) HEC Montréal(蒙特利尔HEC商学院) McGill University(麦吉尔大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03546 2025-06-05 cs.RO cs.AI cs.MA 57%

From Virtual Agents to Robot Teams: A Multi-Robot Framework Evaluation in High-Stakes Healthcare Context

Yuanchen Bai, Zijian Ding, Angelique Taylor

机构 * Cornell University(康奈尔大学) University of Maryland, College Park(马里兰大学学院公园分校)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20573 2025-06-05 cs.RO cs.AI 57%

Collision- and Reachability-Aware Multi-Robot Control with Grounded LLM Planners

Jiabao Ji, Yongchao Chen, Yang Zhang, Ramana Rao Kompella, Chuchu Fan, Gaowen Liu, Shiyu Chang

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏