arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 7409 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7409 篇

2509.11840 2025-09-16 cs.CV 57%

Synthetic Captions for Open-Vocabulary Zero-Shot Segmentation

Tim Lebailly, Vijay Veerabadran, Satwik Kottur, Karl Ridgeway, Michael Louis Iuzzolino

机构 * Meta KU Leuven(鲁汶大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments ICCV 2025 CDEL Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11714 2025-09-16 eess.IV cs.LG 57%

EMeRALDS: Electronic Medical Record Driven Automated Lung Nodule Detection and Classification in Thoracic CT Images

Hafza Eman, Furqan Shaukat, Muhammad Hamza Zafar, Syed Muhammad Anwar

机构 * Faculty of Electrical and Electronics Engineering, University of Engineering(电气电子工程学院,工程大学) Department of Engineering Sciences, University of Agder(工程科学系,阿格德大学) Sheikh Zayed Institute for Pediatric Surgical Innovation, Children’s National Hospital(谢赫扎耶德小儿外科创新研究所,儿童医院) School of Medicine and Health Sciences, George Washington University(医学与健康科学学院,乔治华盛顿大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11336 2025-09-16 cs.AI 57%

The power of dynamic causality in observer-based design for soft sensor applications

William Farlessyost, Sebastian Oberst, Shweta Singh

机构 * organization= Agricultural \& Biological Engineering, Purdue University , country= USA organization= Environmental \& Ecological Engineering, Purdue University , country= USA organization= Davidson School of Chemical Engineering, Purdue University , country= USA organization= Mechanical \& Mechatronic Engineering, University of Technology Sydney (UTS) , country= Australia

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10478 2025-09-16 cs.NI cs.LG cs.SY eess.SY 57%

The LLM as a Network Operator: A Vision for Generative AI in the 6G Radio Access Network

Oluwaseyi Giwa, Michael Adewole, Tobi Awodumila, Pelumi Aderinto

机构 * African Institute for Mathematical Sciences(非洲数学科学研究所)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG

Comments Submitted to Workshop on AI and ML for Next-Generation Wireless Communications and Networking, NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.13776 2025-09-15 cs.CL cs.AI 57%

Aggregation Artifacts in Subjective Tasks Collapse Large Language Models' Posteriors

Georgios Chochlakis, Alexandros Potamianos, Kristina Lerman, Shrikanth Narayanan

机构 * University of Southern California(南加州大学) National Technical University of Athens(雅典国立技术大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments 16 pages, 12 figures, 3 tables

Journal ref Proceedings of the 2025 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 5513-5528, Albuquerque, New Mexico, April 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09172 2025-09-12 cs.CV 57%

Bridging the Gap Between Ideal and Real-world Evaluation: Benchmarking AI-Generated Image Detection in Challenging Scenarios

Chunxiao Li, Xiaoxiao Wang, Meiling Li, Boming Miao, Peng Sun, Yunjian Zhang, Xiangyang Ji, Yao Zhu

机构 * Beijing Normal University(北京师范大学) University of Chinese Academy of Sciences(中国科学院大学) Fudan University(复旦大学) Central University of Finance and Economics(中央财经大学) Tsinghua University(清华大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments ICCV2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08222 2025-09-11 cs.AI 57%

Exploratory Retrieval-Augmented Planning For Continual Embodied Instruction Following

Minjong Yoo, Jinwoo Jang, Wei-jin Park, Honguk Woo

机构 * Department of Computer Science and Engineering, Sungkyunkwan University(全南大学计算机科学与工程系) Acryl Inc.(阿克罗尔公司)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments 21 pages. NeurIPS 2024

Journal ref Advances in Neural Information Processing Systems 37, 67034-67060, 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07577 2025-09-11 cs.AI 57%

Towards explainable decision support using hybrid neural models for logistic terminal automation

Riccardo D'Elia, Alberto Termine, Francesco Flammini

机构 * University of Applied Sciences and Arts of Southern Switzerland(应用科学与艺术大学(南瑞士)) Dalle Molle Institute for Artificial Intelligence(达勒莫勒人工智能研究所) University of Florence(佛罗伦萨大学) Department of Mathematics and Computer Science Ulisse Dini(数学与计算机科学系(乌利塞·迪尼))

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20670 2025-09-10 cs.CV cs.MM 57%

"Humor, Art, or Misinformation?": A Multimodal Dataset for Intent-Aware Synthetic Image Detection

Anastasios Skoularikis, Stefanos-Iordanis Papadopoulos, Symeon Papadopoulos, Panagiotis C. Petrantonakis

机构 * Department of Electrical & Computer Engineering, Aristotle University of Thessaloniki(电气与计算机工程系,阿基米德大学塞萨洛尼基分校) Information Technology Institute, Centre for Research & Technology Hellas(信息科技研究所,希腊研究中心)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06401 2025-09-10 cs.DL cs.AI cs.CL cs.IR 57%

A Systematic Literature Review of Retrieval-Augmented Generation: Techniques, Metrics, and Challenges

Andrew Brown, Muhammad Roman, Barry Devereux

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments 58 page

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10781 2025-09-10 cs.CV 57%

Large-scale Pre-training for Grounded Video Caption Generation

Evangelos Kazakos, Cordelia Schmid, Josef Sivic

机构 * Czech Institute of Informatics, Robotics and Cybernetics at the Czech Technical University in Prague(捷克信息技术、机器人与自动化研究所(捷克技术大学)) Inria, École normale supérieure, CNRS, PSL Research University(法国国家信息与自动化研究所、法国高等师范学校、法国国家科学研究中心、巴黎-萨克勒大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments Accepted at ICCV 2025. Erratum: An earlier version reported ablations (Table 6 & Fig. 6) with pre-training on a 50k subset of HowToGround1M + fine-tuning on iGround. In the ICCV camera-ready, Table 6 already used the full dataset, but Fig. 6 and a sentence in the text were mistakenly left on 50k. All now use the full HowToGround1M

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06422 2025-09-09 cs.CV 57%

Phantom-Insight: Adaptive Multi-cue Fusion for Video Camouflaged Object Detection with Multimodal LLM

Hua Zhang, Changjiang Luo, Ruoyu Chen

机构 * Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络安全学院)

专题命中 视觉定位与Grounding :MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.05112 2025-09-08 cs.SE cs.AI 57%

GenAI-based test case generation and execution in SDV platform

Denesa Zyberaj, Lukasz Mazur, Nenad Petrovic, Pankhuri Verma, Pascal Hirmer, Dirk Slama, Xiangwei Cheng, Alois Knoll

机构 * Mercedes-Benz AG, Germany(梅赛德斯-奔驰集团,德国) Technical University of Munich, Germany(慕尼黑技术大学,德国) Ferdinand-Steinbeis-Institut der Steinbeis-Stiftung, Germany(施坦贝格基金会费尔迪南-斯坦贝格研究所,德国)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.17422 2025-09-08 cs.RO cs.CV 57%

Multimodal LLM Guided Exploration and Active Mapping using Fisher Information

Wen Jiang, Boshu Lei, Katrina Ashton, Kostas Daniilidis

机构 * University of Pennsylvania(宾夕法尼亚大学) Archimedes, Athena RC(阿基米德、阿提卡RC)

专题命中 视觉定位与Grounding :multimodal large language model(abstract);分类 cs.CV

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03793 2025-09-05 cs.MA cs.AI 57%

SAMVAD: A Multi-Agent System for Simulating Judicial Deliberation Dynamics in India

Prathamesh Devadiga, Omkaar Jayadev Shetty, Pooja Agarwal

机构 * PES University(PES大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03626 2025-09-05 cs.AI 57%

Explainable Knowledge Graph Retrieval-Augmented Generation (KG-RAG) with KG-SMILE

Zahra Zehtabi Sabeti Moghaddam, Zeinab Dehghani, Maneeha Rani, Koorosh Aslansefat, Bhupesh Kumar Mishra, Rameez Raja Kureshi, Dhavalkumar Thakker

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02924 2025-09-04 cs.MM cs.AI cs.HC 57%

Simulacra Naturae: Generative Ecosystem driven by Agent-Based Simulations and Brain Organoid Collective Intelligence

Nefeli Manoudaki, Mert Toka, Iason Paterakis, Diarmid Flatley

机构 * Media Arts & Technology UC Santa Barbara(媒体艺术与技术大学圣芭芭拉分校)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments to be published in IEEE VISAP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02837 2025-09-04 cs.IR cs.AI 57%

HF-RAG: Hierarchical Fusion-based RAG with Multiple Sources and Rankers

Payel Santra, Madhusudan Ghosh, Debasis Ganguly, Partha Basuchowdhuri, Sudip Kumar Naskar

机构 * University of Glasgow(格拉斯哥大学) Jadavpur University(贾瓦德普尔大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.17642 2025-09-04 cs.CL cs.AI 57%

Banishing LLM Hallucinations Requires Rethinking Generalization

Johnny Li, Saksham Consul, Eda Zhou, James Wong, Naila Farooqui, Yuxin Ye, Nithyashree Manohar, Zhuxiaona Wei, Tian Wu, Ben Echols, Sharon Zhou, Gregory Diamos

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments I want to revisit some of the experiments in this paper, specifically figure 5

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01716 2025-09-03 cs.AI cs.CL 57%

An LLM-enabled semantic-centric framework to consume privacy policies

Rui Zhao, Vladyslav Melnychuk, Jun Zhao, Jesse Wright, Nigel Shadbolt

机构 * University of Oxford(牛津大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.21484 2025-09-03 q-bio.QM cs.LG stat.ML 57%

Data-driven Discovery of Digital Twins in Biomedical Research

Clémence Métayer, Annabelle Ballesta, Julien Martinelli

机构 * Inserm U1331, Institut Curie Saint-Cloud, France(法国国家医学研究院U1331,圣克鲁医院) Aalto University Espoo, Finland(芬兰艾尔托大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10292 2025-09-03 cs.CV cs.CL 57%

StoryReasoning Dataset: Using Chain-of-Thought for Scene Understanding and Grounded Story Generation

Daniel A. P. Oliveira, David Martins de Matos

机构 * Instituto Superior Técnico, Universidade de Lisboa(里斯本大学技术学院) INESC-ID Lisboa(里斯本INESC-ID)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments 31 pages, 14 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.00600 2025-09-03 cs.DB cs.AI cs.CL 57%

Semantic Integrity Constraints: Declarative Guardrails for AI-Augmented Data Processing Systems

Alexander W. Lee, Justin Chan, Michael Fu, Nicolas Kim, Akshay Mehta, Deepti Raghavan, Ugur Cetintemel

机构 * Brown University(布朗大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Journal ref PVLDB, 18(11): 4073 - 4080, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.07268 2025-09-03 cs.MM cs.CL cs.CV 57%

Advancing Grounded Multimodal Named Entity Recognition via LLM-Based Reformulation and Box-Based Segmentation

Jinyuan Li, Ziyan Li, Han Li, Jianfei Yu, Rui Xia, Di Sun, Gang Pan

机构 * College of Intelligence and Computing, Tianjin University(智能与计算学院,天津大学) NJUST(南京理工大学) College of Mathematics, Taiyuan University of Technology(数学学院,太原科技大学) Tianjin University of Science and Technology(天津科技大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments Extension of our Findings of EMNLP 2023 & ACL 2024 paper, IEEE Transactions on Multimedia accepted on July 19, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.21080 2025-09-01 cs.CV cs.RO 57%

2COOOL: 2nd Workshop on the Challenge Of Out-Of-Label Hazards in Autonomous Driving

Ali K. AlShami, Ryan Rabinowitz, Maged Shoman, Jianwu Fang, Lukas Picek, Shao-Yuan Lo, Steve Cruz, Khang Nhut Lam, Nachiket Kamod, Lei-Lei Li, Jugal Kalita, Terrance E. Boult

机构 * University of Colorado Colorado Springs(科罗拉多州立大学) University of Tennessee–Oak Ridge Innovation Institute(田纳西大学-橡树岭创新研究所) Xi’an Jiaotong University(西安交通大学) University of West Bohemia(西波维亚大学) Honda Research Institute USA(本田美国研究机构) University of Notre Dame(诺特丹大学) Can Tho University(庆和大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments 11 pages, 2 figures, Accepted to ICCV 2025 Workshop on Out-of-Label Hazards in Autonomous Driving (2COOOL)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20976 2025-08-29 cs.SD cs.AI eess.AS 57%

WoW-Bench: Evaluating Fine-Grained Acoustic Perception in Audio-Language Models via Marine Mammal Vocalizations

Jaeyeon Kim, Heeseung Yun, Sang Hoon Woo, Chao-Han Huck Yang, Gunhee Kim

机构 * Carnegie Mellon University(卡内基梅隆大学) Seoul National University(首尔国立大学) NVIDIA(NVIDIA公司)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments Preprint. Project page: https://jaeyeonkim99.github.io/wow_bench/

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01970 2025-08-29 cs.LG 57%

Improving Hospital Risk Prediction with Knowledge-Augmented Multimodal EHR Modeling

Rituparna Datta, Jiaming Cui, Zihan Guan, Vishal G. Reddy, Joshua C. Eby, Gregory Madden, Rupesh Silwal, Anil Vullikanti

机构 * Department of Computer Science, University of Virginia(大学计算机科学系) University of Virginia School of Medicine(弗吉尼亚大学医学院) Virginia Polytechnic Institute and State University(弗吉尼亚理工学院和州立大学) Biocomplexity Institute and Initiative, University of Virginia(大学生物复杂性研究所) Division of Infectious Diseases & International Health, University of Virginia School of Medicine(大学感染性疾病与国际卫生分会)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20029 2025-08-28 cs.CV 57%

Segmentation Assisted Incremental Test Time Adaptation in an Open World

Manogna Sreenivas, Soma Biswas

专题命中 视觉定位与Grounding :vision language model(abstract);分类 cs.CV

Comments Accepted at BMVC 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20020 2025-08-28 cs.CV 57%

GS: Generative Segmentation via Label Diffusion

Yuhao Chen, Shubin Chen, Liang Lin, Guangrun Wang

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments 12 pages, 7 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19517 2025-08-28 cs.HC cs.AI 57%

Orchid: Orchestrating Context Across Creative Workflows with Generative AI

Srishti Palani, Gonzalo Ramos

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏