arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 7409 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7409 篇

2508.18651 2025-08-27 cs.CL cs.AI 57%

Breaking the Trade-Off Between Faithfulness and Expressiveness for Large Language Models

Chenxu Yang, Qingyi Si, Zheng Lin

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18067 2025-08-26 cs.CV 57%

Annotation-Free Open-Vocabulary Segmentation for Remote-Sensing Images

Kaiyu Li, Xiangyong Cao, Ruixun Liu, Shihong Wang, Zixuan Jiang, Zhi Wang, Deyu Meng

机构 * School of Software Engineering, Xi’an Jiaotong University(软件工程学院,西安交通大学) School of Computer Science and Technology and Ministry of Education Key Laboratory of Intelligent Networks and Network Security, Xi’an Jiaotong University(计算机科学与技术学院和教育部智能网络与网络安全重点实验室,西安交通大学) School of Automation, Xi’an Jiaotong University(自动化学院,西安交通大学) College of Artificial Intelligence, Xi’an Jiaotong University(人工智能学院,西安交通大学) School of Mathematics and Statistics and Ministry of Education Key Laboratory of Intelligent Networks and Network Security, Xi’an Jiaotong University(数学与统计学院和教育部智能网络与网络安全重点实验室,西安交通大学)

专题命中 视觉定位与Grounding :VLM(abstract);分类 cs.CV

Comments All codes and models will be released at https://github.com/earth-insights/SegEarth-OV-2

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18050 2025-08-26 cs.CV 57%

ArgusCogito: Chain-of-Thought for Cross-Modal Synergy and Omnidirectional Reasoning in Camouflaged Object Segmentation

Jianwen Tan, Huiyao Zhang, Rui Xiong, Han Zhou, Hongfei Wang, Ye Li

机构 * University of Chinese Academy of Sciences(中国科学院大学) Technology and Engineering Center for Space Utilization, Chinese Academy of Sciences(中国科学院空间利用技术与工程中心)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17255 2025-08-26 cs.CV cs.RO 57%

SEER-VAR: Semantic Egocentric Environment Reasoner for Vehicle Augmented Reality

Yuzhi Lai, Shenghai Yuan, Peizheng Li, Jun Lou, Andreas Zell

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17044 2025-08-26 cs.CV cs.RO 57%

M3DMap: Object-aware Multimodal 3D Mapping for Dynamic Environments

Dmitry Yudin

机构 * Moscow Institute of Physics and Technology(莫斯科物理技术学院) AIRI

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments 29 pages, 3 figures, 13 tables. Preprint of the accepted article in Optical Memory and Neural Network Journal

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.08525 2025-08-26 cs.AI cs.CL 57%

Task Memory Engine (TME): Enhancing State Awareness for Multi-Step LLM Agent Tasks

Ye Ye

机构 * Ye Ye(独立研究者)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments 14 pages, 5 figures. Preprint prepared for future submission. Includes implementation and token-efficiency analysis. Code at https://github.com/biubiutomato/TME-Agent

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15930 2025-08-25 cs.CV 57%

Semantic-Aware Ship Detection with Vision-Language Integration

Jiahao Li, Jiancheng Pan, Yuze Sun, Xiaomeng Huang

机构 * Department of Earth System Science, Tsinghua University 100084 Beijing, China(地球系统科学系,清华大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments 5 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15904 2025-08-25 cs.CV 57%

Boosting Pathology Foundation Models via Few-shot Prompt-tuning for Rare Cancer Subtyping

Dexuan He, Xiao Zhou, Wenbin Guan, Liyuan Zhang, Xiaoman Zhang, Sinuo Xu, Ge Wang, Lifeng Wang, Xiaojun Yuan, Xin Sun, Yanfeng Wang, Kun Sun, Ya Zhang, Weidi Xie

机构 * School of Artificial Intelligence, Shanghai Jiao Tong University(上海交通大学人工智能学院) Department of Biomedical Informatics, Harvard Medical School(哈佛医学院生物医学信息学系) Department of Oral Pathology, Shanghai Ninth People’s Hospital, Shanghai Jiao Tong University School of Medicine(上海交通大学医学院第九人民医院口腔病学部) Department of Pathology, Xinhua Hospital Affiliated to Shanghai Jiao Tong University School of Medicine(上海交通大学医学院新华医院病理科) Department of Pediatric Hematology/Oncology, Xinhua Hospital Affiliated to Shanghai Jiao Tong University School of Medicine(上海交通大学医学院新华医院儿童血液肿瘤科) Clinical Research and Innovation Unit, Xinhua Hospital Affiliated to Shanghai Jiao Tong University School of Medicine(上海交通大学医学院新华医院临床研究与创新单元) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15767 2025-08-22 cs.CV 57%

ATLAS: Decoupling Skeletal and Shape Parameters for Expressive Parametric Human Modeling

Jinhyung Park, Javier Romero, Shunsuke Saito, Fabian Prada, Takaaki Shiratori, Yichen Xu, Federica Bogo, Shoou-I Yu, Kris Kitani, Rawal Khirodkar

机构 * Meta Carnegie Mellon University(卡内基梅隆大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments ICCV 2025; Website: https://jindapark.github.io/projects/atlas/

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15641 2025-08-22 cs.CV 57%

When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding

Pengcheng Fang, Yuxia Chen, Rui Guo

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15236 2025-08-22 eess.IV cs.CV 57%

Pathology-Informed Latent Diffusion Model for Anomaly Detection in Lymph Node Metastasis

Jiamu Wang, Keunho Byeon, Jinsol Song, Anh Nguyen, Sangjeong Ahn, Sung Hak Lee, Jin Tae Kwak

机构 * School of Electrical Engineering, Korea University, Seoul 02841, Korea(韩国大学电子工程学院) Department of Pathology, Korea University Anam Hospital and Department of Biomedical Informatics, Korea University College of Medicine, Seoul 02841, Korea(韩国大学医学院病理学系) Department of Hospital Pathology, Seoul St. Mary’s Hospital, College of Medicine, The Catholic University of Korea, Seoul 06591, Korea(韩国天主大学医学院圣玛丽医院医院病理学系)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13560 2025-08-22 cs.CV 57%

DictAS: A Framework for Class-Generalizable Few-Shot Anomaly Segmentation via Dictionary Lookup

Zhen Qu, Xian Tao, Xinyi Gong, ShiChen Qu, Xiaopei Zhang, Xingang Wang, Fei Shen, Zhengtao Zhang, Mukesh Prasad, Guiguang Ding

机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Casivision Longmen Laboratory(龙门实验室) HDU UTS UCLA(加州大学洛杉矶分校) Tsinghua University(清华大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments Accepted by ICCV 2025, Project: https://github.com/xiaozhen228/DictAS

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.22668 2025-08-22 cs.CV 57%

Understanding Co-speech Gestures in-the-wild

Sindhu B Hegde, K R Prajwal, Taein Kwon, Andrew Zisserman

机构 * Visual Geometry Group, Dept. of Engineering Science, University of Oxford(牛津大学视觉几何组)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments Main paper - 11 pages, 4 figures, Supplementary - 6 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14729 2025-08-21 cs.CV 57%

Multiscale Video Transformers for Class Agnostic Segmentation in Autonomous Driving

Leila Cheshmi, Mennatullah Siam

机构 * Ontariotechu(安大略理工学院)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments 6 pages, 2 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14197 2025-08-21 cs.CV 57%

CLIPSym: Delving into Symmetry Detection with CLIP

Tinghan Yang, Md Ashiqur Rahman, Raymond A. Yeh

机构 * Department of Computer Science, Purdue University(计算机科学系,普渡大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14080 2025-08-21 cs.LG 57%

KnowDR-REC: A Benchmark for Referring Expression Comprehension with Real-World Knowledge

Guanghao Jin, Jingpei Wu, Tianpei Guo, Yiyi Niu, Weidong Zhou, Guoyang Liu

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13901 2025-08-20 cs.RO cs.CV 57%

Multimodal Data Storage and Retrieval for Embodied AI: A Survey

Yihao Lu, Hao Tang

机构 * School of Economics and Management, South China Normal University(经济管理学院,华南师范大学) State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University(多媒体信息处理国家重点实验室,计算机学院,北京大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13530 2025-08-20 cs.AI 57%

CrafterDojo: A Suite of Foundation Models for Building Open-Ended Embodied Agents in Crafter

Junyeong Park, Hyeonseo Cho, Sungjin Ahn

机构 * CrafterDojo: A Suite of Foundation Models for Building Open-Ended Embodied Agents in Crafter(CrafterDojo:构建开放性具身智能体的基础模型集合)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.04678 2025-08-20 eess.IV cs.CV 57%

RadGPT: Constructing 3D Image-Text Tumor Datasets

Pedro R. A. S. Bassi, Mehmet Can Yavuz, Kang Wang, Xiaoxi Chen, Wenxuan Li, Sergio Decherchi, Andrea Cavalli, Yang Yang, Alan Yuille, Zongwei Zhou

机构 * Johns Hopkins University(约翰霍普金斯大学) University of Bologna(博洛尼亚大学) Italian Institute of Technology(意大利理工学院) University of California, San Francisco(加州大学旧金山分校) Istanbul Medipol University(伊斯坦布尔Medipol大学) University of Zurich(苏黎世大学) ETH AI Center(ETH人工智能中心) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) École Polytechnique Fédérale de Lausanne(洛桑联邦理工学院)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13043 2025-08-19 cs.CV 57%

IntelliCap: Intelligent Guidance for Consistent View Sampling

Ayaka Yasunaga, Hideo Saito, Dieter Schmalstieg, Shohei Mori

机构 * Keio University(庆应大学) University of Stuttgart(斯图加特大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments This work is a pre-print version of a paper that has been accepted to the IEEE International Symposium on Mixed and Augmented Reality for future publication. Project Page: https://mediated-reality.github.io/projects/yasunaga_ismar25/

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07402 2025-08-19 cs.CV 57%

ForensicsSAM: Toward Robust and Unified Image Forgery Detection and Localization Resisting to Adversarial Attack

Rongxuan Peng, Shunquan Tan, Chenqi Kong, Anwei Luo, Alex C. Kot, Jiwu Huang

机构 * Shenzhen Key Laboratory of Media Security, Faculty of Electronic and Information Engineering, Shenzhen University, China(深圳媒体安全重点实验室,电子与信息工程学院,深圳大学,中国) Rapid-Rich Object Search (ROSE) Lab, School of Electrical and Electronic Engineering, Nanyang Technological University, Singapore(快速富对象搜索(ROSE)实验室,电气电子工程学院,南洋理工大学,新加坡) Guangdong Laboratory of Machine Perception and Intelligent Computing, Faculty of Engineering, Shenzhen MSU-BIT University, China(广东机器感知与智能计算实验室,工程学院,深圳MSU-BIT大学,中国)

专题命中 视觉定位与Grounding :LLaVA(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.24039 2025-08-19 cs.CV cs.HC 57%

Foundation Models for Zero-Shot Segmentation of Scientific Images without AI-Ready Data

Shubhabrata Mukherjee, Jack Lang, Obeen Kwon, Iryna Zenyuk, Valerie Brogden, Adam Weber, Daniela Ushizima

机构 * Lawrence Berkeley National Laboratory(伯克利国家实验室) University of California, Irvine(加州大学尔湾分校) University of California, Berkeley(加州大学伯克利分校) Covalent Metrology(协力计量)

专题命中 视觉定位与Grounding :visual reasoning(abstract);分类 cs.CV

Comments This paper has been accepted for presentation at the 59th International Conference on Parallel Processing (ICPP 2025), DRAI workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02356 2025-08-19 cs.CV 57%

InterRVOS: Interaction-aware Referring Video Object Segmentation

Woojeong Jin, Seongchan Kim, Jaeho Lee, Seungryong Kim

专题命中 视觉定位与Grounding :MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11873 2025-08-19 cs.CY cs.AI cs.HC cs.MM 57%

SimInterview: Transforming Business Education through Large Language Model-Based Simulated Multilingual Interview Training System

Truong Thanh Hung Nguyen, Tran Diem Quynh Nguyen, Hoang Loc Cao, Thi Cam Thanh Tran, Thi Cam Mai Truong, Hung Cao

机构 * Analytics Everywhere Lab, University of New Brunswick, Canada(新不伦瑞克大学分析 everywhere 实验室) University of Foreign Language Studies, University of Danang, Vietnam(越南丹绒大学外语学院) Faculty of Information Technology, University of Science, VNU-HCM, Vietnam(越南胡志明市大学信息科技学院) Faculty of Economics and Accounting, Quy Nhon University, Vietnam(越南奎隆大学经济与会计学院) Faculty of Natural Sciences, Quy Nhon University, Vietnam(越南奎隆大学自然科学学院)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments Published as a conference paper at ICEFM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11690 2025-08-19 cs.CY cs.AI 57%

Real Time Child Abduction And Detection System

Tadisetty Sai Yashwanth, Yangalasetty Sruthi Royal, Vankayala Rajeshwari Shreya, Mayank Kashyap, Divyaprabha K N

机构 * Dept. of CSE PES University Bangalore, India(计算机科学与工程系,PES大学,印度班加罗尔)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09023 2025-08-18 cs.DB cs.AI cs.CL 57%

E3-Rewrite: Learning to Rewrite SQL for Executability, Equivalence,and Efficiency

Dongjie Xu, Yue Cui, Weijie Shi, Qingzhi Ma, Hanghui Guo, Jiaming Li, Yao Zhao, Ruiyuan Zhang, Shimin Di, Jia Zhu, Kai Zheng, Jiajie Xu

机构 * stu.suda.edu.cn(苏州大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17593 2025-08-18 cs.HC cs.AI 57%

JELAI: Integrating AI and Learning Analytics in Jupyter Notebooks

Manuel Valle Torre, Thom van der Velden, Marcus Specht, Catharine Oertel

机构 * Delft University of Technology(代尔夫特理工大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments Accepted for AIED 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10681 2025-08-15 cs.CV 57%

IADGPT: Unified LVLM for Few-Shot Industrial Anomaly Detection, Localization, and Reasoning via In-Context Learning

Mengyang Zhao, Teng Fu, Haiyang Yu, Ke Niu, Bin Li

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.18923 2025-08-14 cs.CV 57%

Video SimpleQA: Towards Factuality Evaluation in Large Video Language Models

Meng Cao, Pengfei Hu, Yingyao Wang, Jihao Gu, Haoran Tang, Haoze Zhao, Chen Wang, Jiahua Dong, Wangbo Yu, Ge Zhang, Jun Song, Xiang Li, Bo Zheng, Ian Reid, Xiaodan Liang

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09566 2025-08-14 cs.CV 57%

A Chain of Diagnosis Framework for Accurate and Explainable Radiology Report Generation

Haibo Jin, Haoxuan Che, Sunan He, Hao Chen

机构 * Department of Computer Science and Engineering, Hong Kong University of Science and Technology(计算机科学与工程系,香港科学理工大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments Accepted to IEEE TMI

详情

展开后加载摘要…

URL PDF HTML 收藏