arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 7473 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7473 篇

2507.08335 2025-07-14 cs.CL 50%

MK2 at PBIG Competition: A Prompt Generation Solution

Yuzheng Xu, Tosho Hirasawa, Seiya Kawano, Shota Kato, Tadashi Kozuno

机构 * OMRON SINIC X NexaScience Kyoto Institute of Technology(京都技术大学) Kyoto University(京都大学)

专题命中 视觉定位与Grounding :grounding(abstract)

Comments 9 pages, to appear in the 2nd Workshop on Agent AI for Scenario Planning (AGENTSCEN 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04791 2025-07-08 cs.RO 50%

Safe Bimanual Teleoperation with Language-Guided Collision Avoidance

Dionis Totsila, Clemente Donoso, Enrico Mingo Hoffman, Jean-Baptiste Mouret, Serena Ivaldi

机构 * INRIA, Université de Lorraine, CNRS(INRIA、洛林大学、法国国家科学研究中心)

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.03241 2025-07-08 cs.CL 50%

KinyaColBERT: A Lexically Grounded Retrieval Model for Low-Resource Retrieval-Augmented Generation

Antoine Nzeyimana, Andre Niyongabo Rubungo

机构 * University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校) Princeton University(普林斯顿大学)

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.02694 2025-07-04 cs.CL 50%

Can LLMs Identify Critical Limitations within Scientific Research? A Systematic Evaluation on AI Research Papers

Zhijian Xu, Yilun Zhao, Manasi Patwardhan, Lovekesh Vig, Arman Cohan

机构 * Yale University(耶鲁大学) TCS Research(TCS研究)

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.01967 2025-07-04 q-bio.NC 50%

Ghost in the Machine: Examining the Philosophical Implications of Recursive Algorithms in Artificial Intelligence Systems

Llewellin RG Jegels

专题命中 视觉定位与Grounding :grounding(abstract)

Comments 27 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.00718 2025-07-02 cs.CL 50%

AI Analyst: Framework and Comprehensive Evaluation of Large Language Models for Financial Time Series Report Generation

Elizabeth Fons, Elena Kochkina, Rachneet Kaur, Zhen Zeng, Berowne Hlavaty, Charese Smiley, Svitlana Vyetrenko, Manuela Veloso

机构 * J.P. Morgan AI Research(摩根大通人工智能研究)

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18548 2025-07-02 cs.IR 50%

Rethinking Click Models in Light of Carousel Interfaces: Theory-Based Categorization and Design of Click Models

Jingwei Kang, Maarten de Rijke, Santiago de Leon-Martinez, Harrie Oosterhuis

专题命中 视觉定位与Grounding :grounding(abstract)

Comments Accepted by ICTIR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22210 2025-06-30 cs.IR 50%

UiS-IAI@LiveRAG: Retrieval-Augmented Information Nugget-Based Generation of Responses

Weronika Łajewska, Ivica Kostric, Gabriel Iturra-Bocaz, Mariam Arustashvili, Krisztian Balog

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17808 2025-06-24 cs.CY 50%

The value of human and machine in machine-generated creative contents

Weina Jin

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.05846 2025-06-24 eess.SY cs.SY 50%

Enhanced Rapid Detection of High-impedance Arc Faults in Medium Voltage Electrical Distribution Networks

Kriti Thakur, Divyanshi Dwivedi, K. Victor Sam Moses Babu, Alivelu Manga Parimi, Prasanta K. Panigrahi, Pradeep Kumar Yemula, Pratyush Chakraborty, Mayukha Pal

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.15242 2025-06-23 cs.HC 50%

Agonistic Image Generation: Unsettling the Hegemony of Intention

Andrew Shaw, Andre Ye, Ranjay Krishna, Amy X. Zhang

专题命中 视觉定位与Grounding :grounding(abstract)

Comments Accepted to ACM Fairness, Accountability, Transparency 2025 -- Athens, Greece

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.15522 2025-06-19 cs.CL 50%

Lessons from Training Grounded LLMs with Verifiable Rewards

Shang Hong Sim, Tej Deep Pala, Vernon Toh, Hai Leong Chieu, Amir Zadeh, Chuan Li, Navonil Majumder, Soujanya Poria

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.14561 2025-06-18 stat.AP stat.ML 50%

Bayesian Hybrid Machine Learning of Gallstone Risk

Chitradipa Chakraborty, Nayana Mukherjee

专题命中 视觉定位与Grounding :grounding(abstract)

Comments 25 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.14378 2025-06-18 physics.soc-ph cond-mat.dis-nn cond-mat.stat-mech 50%

A novel approach for converting spatio-temporal series into complex networks

G. Cigdem Yalcin, M. Berk Onder

专题命中 视觉定位与Grounding :grounding(abstract)

Comments 6 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.14040 2025-06-18 cs.CL cs.HC 50%

An Interdisciplinary Review of Commonsense Reasoning and Intent Detection

Md Nazmus Sakib

机构 * University of Maryland, Baltimore County(马里兰大学巴尔的摩分校)

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13178 2025-06-17 cs.CL 50%

Enhancing Large Language Models with Reliable Knowledge Graphs

Qinggang Zhang

机构 * Department of Computing(计算系)

专题命中 视觉定位与Grounding :grounding(abstract)

Comments Thesis

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.07778 2025-06-17 cs.CL 50%

WorldAPIs: The World Is Worth How Many APIs? A Thought Experiment

Jiefu Ou, Arda Uzunoglu, Benjamin Van Durme, Daniel Khashabi

专题命中 视觉定位与Grounding :grounding(abstract)

Comments AAAI 2025 & ACL 2024 NLRSE, 7 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11478 2025-06-16 cs.CL 50%

ImmunoFOMO: Are Language Models missing what oncologists see?

Aman Sinha, Bogdan-Valentin Popescu, Xavier Coubez, Marianne Clausel, Mathieu Constant

机构 * Université de Lorraine(洛林大学) ICANS

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07633 2025-06-10 cs.RO 50%

Blending Participatory Design and Artificial Awareness for Trustworthy Autonomous Vehicles

Ana Tanevska, Ananthapathmanabhan Ratheesh Kumar, Arabinda Ghosh, Ernesto Casablanca, Ginevra Castellano, Sadegh Soudjani

机构 * Uppsala University(乌普萨拉大学) Max Planck Institute for Software Systems(马克斯·普朗克软件系统研究所) Newcastle University(新castle大学)

专题命中 视觉定位与Grounding :grounding(abstract)

Comments Submitted to IEEE RO-MAN 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06091 2025-06-10 cs.CL 50%

MIRIAD: Augmenting LLMs with millions of medical query-response pairs

Qinyue Zheng, Salman Abdullah, Sam Rawal, Cyril Zakka, Sophie Ostmeier, Maximilian Purk, Eduardo Reis, Eric J. Topol, Jure Leskovec, Michael Moor

机构 * Department of Biosystems Science and Engineering, ETH Zurich(生物系统科学与工程系,苏黎世联邦理工学院) Department of Computer Science, Stanford University(计算机科学系,斯坦福大学) Department of Internal Medicine, Mayo Clinic(内科医学系,梅奥诊所) Hugging Face Department of Radiology, Stanford University(放射学系,斯坦福大学) Hasso-Plattner-Institute for Digital Engineering, University of Potsdam(数字工程研究所,波茨坦大学) Center for Artificial Intelligence in Medicine and Imaging, Stanford, CA, USA(医学与成像人工智能中心,斯坦福,CA,美国) Scripps Translational Science Institute, San Diego, CA, USA(斯克里普斯转化科学研究所,圣地亚哥,CA,美国)

专题命中 视觉定位与Grounding :grounding(abstract)

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05142 2025-06-10 cs.CL 50%

Do Large Language Models Judge Error Severity Like Humans?

Diege Sun, Guanyi Chen, Zhao Fan, Xiaorong Cheng, Tingting He

机构 * School of Psychology(心理学系) Hubei Provincial Key Laboratory of Artificial Intelligence and Smart Learning(湖北省人工智能与智能学习重点实验室) National Language Resources Monitor and Research Center for Network Media(网络媒体语言资源监测与研究中心) School of Computer Science(计算机科学学院)

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01287 2025-06-03 cs.HC 50%

How Problematic are Suspenseful Interactions?

Alarith Uhde

专题命中 视觉定位与Grounding :grounding(abstract)

Comments 14 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11726 2025-06-03 cs.CL 50%

Disambiguating Reference in Visually Grounded Dialogues through Joint Modeling of Textual and Multimodal Semantic Structures

Shun Inadumi, Nobuhiro Ueda, Koichiro Yoshino

机构 * Nara Institute of Science and Technology(奈良科学技術研究所) Guardian Robot Project(守護機器人計劃) RIKEN(理化学研究所) Kyoto University(京都大學) Institute of Science Tokyo(東京科學研究院)

专题命中 视觉定位与Grounding :grounding(abstract)

Comments ACL2025 main. Code available at https://github.com/SInadumi/mmrr

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10084 2025-06-03 cs.CL 50%

Why Prompt Design Matters and Works: A Complexity Analysis of Prompt Search Space in LLMs

Xiang Zhang, Juntai Cao, Jiaqi Wei, Chenyu You, Dujian Ding

专题命中 视觉定位与Grounding :grounding(abstract)

Comments ACL 2025 main conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24423 2025-06-02 cs.CL 50%

MMAFFBen: A Multilingual and Multimodal Affective Analysis Benchmark for Evaluating LLMs and VLMs

Zhiwei Liu, Lingfei Qian, Qianqian Xie, Jimin Huang, Kailai Yang, Sophia Ananiadou

机构 * The University of Manchester(曼彻斯特大学) The Fin AI Singapore(Fin AI新加坡) School of Artificial Intelligence, Wuhan University(武汉大学人工智能学院)

专题命中 视觉定位与Grounding :vision-language model(abstract)

Comments Work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22879 2025-05-30 cs.SE cs.DC 50%

Visualizing Cloud-native Applications with KubeDiagrams

Philippe Merle, Fabio Petrillo

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21542 2025-05-29 cs.CY 50%

Toward a Cultural Co-Genesis of AI Ethics

Ammar Younas

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16849 2025-05-29 cs.IR cs.CL 50%

Walk&Retrieve: Simple Yet Effective Zero-shot Retrieval-Augmented Generation via Knowledge Graph Walks

Martin Böckling, Heiko Paulheim, Andreea Iana

机构 * University of Mannheim(曼海姆大学)

专题命中 视觉定位与Grounding :grounding(abstract)

Comments Accepted at the Information Retrieval's Role in RAG Systems (IR-RAG 2025) in conjunction with SIGIR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.16514 2025-05-29 cs.CL 50%

GraphCheck: Breaking Long-Term Text Barriers with Extracted Knowledge Graph-Powered Fact-Checking

Yingjian Chen, Haoran Liu, Yinhong Liu, Jinxiang Xie, Rui Yang, Han Yuan, Yanran Fu, Peng Yuan Zhou, Qingyu Chen, James Caverlee, Irene Li

机构 * University of Tokyo(东京大学) Texas A&M University(德克萨斯A&M大学) University of Cambridge(剑桥大学) Duke-NUS Medical School(杜克-新加坡国立大学医学院) Aarhus University(阿贾克斯大学) Yale University(耶鲁大学)

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.10066 2025-05-29 cs.LO 50%

Non-Ground Congruence Closure

Hendrik Leidinger, Christoph Weidenbach

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏