arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 7409 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7409 篇

2509.24171 2025-09-30 cs.LG 57%

Model Correlation Detection via Random Selection Probing

Ruibo Chen, Sheng Zhang, Yihan Wu, Tong Zheng, Peihua Mai, Heng Huang

机构 * University of Maryland, College Park(马里兰大学学院公园分校)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23955 2025-09-30 cs.CV 57%

ColLab: A Collaborative Spatial Progressive Data Engine for Referring Expression Comprehension and Generation

Shilan Zhang, Jirui Huang, Ruilin Yao, Cong Wang, Yaxiong Chen, Peng Xu, Shengwu Xiong

机构 * Wuhan University of Technology(武汉理工大学) Northwestern Polytechnical University(西北工业大学) Tsinghua University(清华大学)

专题命中 视觉定位与Grounding :multimodal large language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23838 2025-09-30 cs.CV 57%

2nd Place Report of MOSEv2 Challenge 2025: Concept Guided Video Object Segmentation via SeC

Zhixiong Zhang, Shuangrui Ding, Xiaoyi Dong, Yuhang Zang, Yuhang Cao, Jiaqi Wang

机构 * Shanghai Jiao Tong University(上海交通大学) Shanghai Innovation Institute(上海创新研究院) Shanghai AI Laboratory(上海人工智能实验室) The Chinese University of Hong Kong(香港中文大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22019 2025-09-29 cs.CV 57%

EgoInstruct: An Egocentric Video Dataset of Face-to-face Instructional Interactions with Multi-modal LLM Benchmarking

Yuki Sakai, Ryosuke Furuta, Juichun Yen, Yoichi Sato

机构 * The University of Tokyo(东京大学)

专题命中 视觉定位与Grounding :multimodal large language model(abstract);分类 cs.CV

Comments Accepted to the I-HFM Workshop at ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21367 2025-09-29 cs.CR cs.AI 57%

Design and Implementation of a Secure RAG-Enhanced AI Chatbot for Smart Tourism Customer Service: Defending Against Prompt Injection Attacks -- A Case Study of Hsinchu, Taiwan

Yu-Kai Shih, You-Kai Kang

机构 * Department of Information Management, National Dong Hwa University(信息管理系,国立东华大学) By The Student (BTS) Experimental Education Program (Non-School Type)(学生(BTS)实验教育计划(非学校类型))

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments 12 pages, 7 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20369 2025-09-26 cs.CY cs.AI cs.HC 57%

AI-driven formative assessment and adaptive learning in data-science education: Evaluating an LLM-powered virtual teaching assistant

Fadjimata I Anaroua, Qing Li, Yan Tang, Hong P. Liu

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.15356 2025-09-26 cs.AI cs.HC 57%

A Decision Theoretic Framework for Measuring AI Reliance

Ziyang Guo, Yifan Wu, Jason Hartline, Jessica Hullman

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19843 2025-09-25 cs.CV cs.RO 57%

PersONAL: Towards a Comprehensive Benchmark for Personalized Embodied Agents

Filippo Ziliotto, Jelin Raphael Akkara, Alessandro Daniele, Lamberto Ballan, Luciano Serafini, Tommaso Campari

机构 * University of Padova(帕多瓦大学) Fondazione Bruno Kessler(布鲁诺·科塞拉基金会)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19326 2025-09-25 cs.CL cs.AI 57%

Unveiling the Merits and Defects of LLMs in Automatic Review Generation for Scientific Papers

Ruochi Li, Haoxuan Zhang, Edward Gehringer, Ting Xiao, Junhua Ding, Haihua Chen

机构 * North Carolina State University(北卡罗来纳州立大学) University of North Texas(德克萨斯大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments Accepted as short paper at 25th IEEE International Conference on Data Mining

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19322 2025-09-25 cs.CL cs.AI 57%

Readme_AI: Dynamic Context Construction for Large Language Models

Millie Vyas, Timothy Blattner, Alden Dima

机构 * Purdue University(普渡大学) National Institute of Standards and Technology(国家标准技术研究院)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16180 2025-09-25 cs.CV cs.CL 57%

Redemption Score: A Multi-Modal Evaluation Framework for Image Captioning via Distributional, Perceptual, and Linguistic Signal Triangulation

Ashim Dahal, Ankit Ghimire, Saydul Akbar Murad, Nick Rahimi

机构 * University of Southern Mississippi(密苏里州南方大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18733 2025-09-24 cs.CV 57%

Knowledge Transfer from Interaction Learning

Yilin Gao, Kangyi Chen, Zhongxing Peng, Hengjie Lu, Shugong Xu

机构 * Shanghai University(上海大学) Xi’an Jiaotong-Liverpool University(西安交通大学利物浦大学)

专题命中 视觉定位与Grounding :vision language model(abstract);分类 cs.CV

Comments Accepted by ICCV2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.05023 2025-09-24 cs.CV 57%

Split Matching for Inductive Zero-shot Semantic Segmentation

Jialei Chen, Xu Zheng, Dongyue Li, Chong Yi, Seigo Ito, Danda Pani Paudel, Luc Van Gool, Hiroshi Murase, Daisuke Deguchi

机构 * Graduate School of Informatics, Nagoya University (NU)(名古屋大学信息研究生院) Artificial Intelligence Thrust, The Hong Kong University of Science and Technology (Guangzhou) (HKUST (GZ))(香港科学与技术大学(广州)人工智能推进部) Institute for Computer Science, Artificial Intelligence and Technology (INSAIT), Sofia University, St. Kliment Ohridski(索菲亚大学计算机科学、人工智能与技术研究所)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments Accepted by BMVC 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.15794 2025-09-24 cs.AI cs.LO 57%

Finite Groundings for ASP with Functions: A Journey through Consistency

Lukas Gerlach, David Carral, Markus Hecher

机构 * Knowledge-Based Systems Group, TU Dresden(图腾大学达姆施塔特分校知识系统小组) LIRMM, Inria, University of Montpellier, CNRS(里尔毫米研究所、法国国家信息与自动化技术研究院、蒙彼利埃大学、国家科学研究中心) Massachusetts Institute of Technology(麻省理工学院)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments to be published at IJCAI 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18096 2025-09-23 cs.CV 57%

Seg4Diff: Unveiling Open-Vocabulary Segmentation in Text-to-Image Diffusion Transformers

Chaehyun Kim, Heeseong Shin, Eunbeen Hong, Heeji Yoon, Anurag Arnab, Paul Hongsuck Seo, Sunghwan Hong, Seungryong Kim

机构 * KAIST AI(韩国科学技术院人工智能研究所) Korea University(韩国大学) ETH Zürich(苏黎世联邦理工学院)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments NeurIPS 2025. Project page: https://cvlab-kaist.github.io/Seg4Diff/

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17671 2025-09-23 cs.CL cs.AI 57%

Turk-LettuceDetect: A Hallucination Detection Models for Turkish RAG Applications

Selva Taş, Mahmut El Huseyni, Özay Ezerceli, Reyhan Bayraktar, Fatma Betül Terzioğlu

机构 * Hidden for Review(保密)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17615 2025-09-23 cs.CV 57%

From Benchmarks to Reality: Advancing Visual Anomaly Detection by the VAND 3.0 Challenge

Lars Heckler-Kram, Ashwin Vaidya, Jan-Hendrik Neudeck, Ulla Scheler, Dick Ameln, Samet Akcay, Paula Ramos

机构 * MVTec Software GmbH(MVTec软件公司) Technical University of Munich(慕尼黑技术大学) Intel(英特尔) Voxel51

专题命中 视觉定位与Grounding :vision language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17522 2025-09-23 cs.CV 57%

Chat-CBM: Towards Interactive Concept Bottleneck Models with Frozen Large Language Models

Hangzhou He, Lei Zhu, Kaiwen Li, Xinliang Zhang, Jiakui Hu, Ourui Fu, Zhengjian Yao, Yanye Lu

机构 * Department of Biomedical Engineering, College of Future Technology, Peking University(生物医学工程系,未来技术学院,北京大学) Institute of Medical Technology, Peking University Health Science Center, Peking University(医学技术研究所,北京大学医学部,北京大学) National Biomedical Imaging Center, College of Future Technology, Peking University(国家生物医学成像中心,未来技术学院,北京大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16810 2025-09-23 cs.AI 57%

Automated Procedural Analysis via Video-Language Models for AI-assisted Nursing Skills Assessment

Shen Chang, Dennis Liu, Renran Tian, Kristen L. Swartzell, Stacie L. Klingler, Amy M. Nagle, Nan Kong

机构 * Weldon School of Biomedical Engineering, Purdue University(普渡大学生物医学工程学院) Department of Industrial and Operations Engineering, University of Michigan(密歇根大学工业与运作工程系) Edward P. Fitts Department of Industrial and Systems Engineering, North Carolina State University(北卡罗来纳州立大学工业与系统工程系) School of Nursing, Purdue University(普渡大学护理学院)

专题命中 视觉定位与Grounding :VLM(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04484 2025-09-23 cs.CL cs.AI cs.CY 57%

The Good, the Bad and the Constructive: Automatically Measuring Peer Review's Utility for Authors

Abdelrahman Sadallah, Tim Baumgärtner, Iryna Gurevych, Ted Briscoe

机构 * NLP Department, Mohamed Bin Zayed University of Artificial Intelligence(马尔代夫比兹艾兹大学人工智能学院自然语言处理系) Ubiquitous Knowledge Processing Lab, Department of Computer Science(计算机科学系通用知识处理实验室) Hessian Center for AI (hessian.AI), TU Darmstadt(图尔努尔德马斯特大学海斯塞人工智能中心)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments EMNLP 2025 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15108 2025-09-23 cs.CL cs.AI cs.HC 57%

A Risk Ontology for Evaluating AI-Powered Psychotherapy Virtual Agents

Ian Steenstra, Timothy W. Bickmore

机构 * Northeastern University(东北大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments This is a preprint version of the paper accepted to IVA'25

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16502 2025-09-23 cs.LG 57%

GRIL: Knowledge Graph Retrieval-Integrated Learning with Large Language Models

Jialin Chen, Houyu Zhang, Seongjun Yun, Alejandro Mottini, Rex Ying, Xiang Song, Vassilis N. Ioannidis, Zheng Li, Qingjun Cui

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16297 2025-09-23 cs.CY cs.AI cs.CL 57%

How Large Language Models are Designed to Hallucinate

Richard Ackermann, Simeon Emanuilov

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments 23 pages, 2 tables, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14312 2025-09-22 cs.CV 57%

CLIPTTA: Robust Contrastive Vision-Language Test-Time Adaptation

Marc Lafon, Gustavo Adolfo Vargas Hakim, Clément Rambour, Christian Desrosier, Nicolas Thome

机构 * Conservatoire National des Arts et Métiers(法国国家艺术与工艺学院) Sorbonne Université(索邦大学) ETS Montreal(蒙特利尔ETS) Institut universitaire de France(法国国家科学研究中心)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Journal ref 39th Conference on Neural Information Processing Systems, NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.01029 2025-09-22 cs.CY cs.AI cs.DB cs.HC 57%

Who is Responsible When AI Fails? Mapping Causes, Entities, and Consequences of AI Privacy and Ethical Incidents

Hilda Hadan, Reza Hadi Mogavi, Leah Zhang-Kennedy, Lennart E. Nacke

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments 63 pages, 7 tables, 7 figures

Journal ref International Journal of Human-Computer Interaction (2025): 1-45

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14985 2025-09-19 cs.CV 57%

PRISM: Product Retrieval In Shopping Carts using Hybrid Matching

Arda Kabadayi, Senem Velipasalar, Jiajing Chen

机构 * Syracuse University(苏塞克斯大学) Amazon(亚马逊)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14623 2025-09-19 cs.SE cs.AI cs.PL cs.SY eess.SY 57%

Automating Modelica Module Generation Using Large Language Models: A Case Study on Building Control Description Language

Hanlong Wan, Xing Lu, Yan Chen, Karthik Devaprasad, Laura Hinkle

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments This is the pre-peer-review version of a journal paper; the repo is available at: https://github.com/pnnl/prompt2control

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13888 2025-09-18 cs.CL cs.AI cs.IR 57%

Combating Biomedical Misinformation through Multi-modal Claim Detection and Evidence-based Verification

Mariano Barone, Antonio Romano, Giuseppe Riccio, Marco Postiglione, Vincenzo Moscato

机构 * University of Naples Federico II(那不勒斯费迪里奇二世大学) Northwestern University(西北大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Journal ref SIGIR '25: Proceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retrieval, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13289 2025-09-17 cs.CV eess.IV 57%

Image Realness Assessment and Localization with Multimodal Features

Lovish Kaushik, Agnij Biswas, Somdyuti Paul

机构 * Indian Institute of Technology, Kharagpur(印度理工学院,克哈拉格普)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.22413 2025-09-17 cs.CR cs.LG 57%

Instance-Level Data-Use Auditing of Visual ML Models

Zonghao Huang, Neil Zhenqiang Gong, Michael K. Reiter

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏