arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 7473 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7473 篇

2505.19360 2025-05-27 cs.CL 50%

ChartLens: Fine-grained Visual Attribution in Charts

Manan Suri, Puneet Mathur, Nedim Lipka, Franck Dernoncourt, Ryan A. Rossi, Dinesh Manocha

机构 * University of Maryland(马里兰大学) Adobe Research(Adobe研究)

专题命中 视觉定位与Grounding :multimodal large language model(abstract)

Comments ACL 2025 (Main)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09866 2025-05-27 cs.CL 50%

PASS-FC: Progressive and Adaptive Search Scheme for Fact Checking of Comprehensive Claims

Ziyu Zhuang

机构 * Trip.com Group(Trip.com集团)

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17099 2025-05-26 cs.CL 50%

Learning Interpretable Representations Leads to Semantically Faithful EEG-to-Text Generation

Xiaozhao Liu, Dinggang Shen, Xihui Liu

机构 * University of Hong Kong(香港大学) ShanghaiTech University(上海科技大学)

专题命中 视觉定位与Grounding :grounding(abstract)

Comments Code, checkpoint and text samples available at https://github.com/justin-xzliu/GLIM

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15872 2025-05-26 cs.IR cs.CL 50%

InfoDeepSeek: Benchmarking Agentic Information Seeking for Retrieval-Augmented Generation

Yunjia Xi, Jianghao Lin, Menghui Zhu, Yongzhao Xiao, Zhuoying Ou, Jiaqi Liu, Tong Wan, Bo Chen, Weiwen Liu, Yasheng Wang, Ruiming Tang, Weinan Zhang, Yong Yu

机构 * Shanghai Jiao Tong University(上海交通大学) Huawei Noah’s Ark Lab(华为诺亚实验室)

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.12632 2025-05-26 quant-ph 50%

Transferring linearly fixed QAOA angles: performance and real device results

Ryo Sakai, Hiromichi Matsuyama, Wai-Hong Tam, Yu Yamashiro

专题命中 视觉定位与Grounding :grounding(abstract)

Comments 11 pages, 10 figures, submitted to 2025 IEEE International Conference on Quantum Computing and Engineering (QCE25)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16125 2025-05-23 cs.CL 50%

KoBALT: Korean Benchmark For Advanced Linguistic Tasks

Hyopil Shin, Sangah Lee, Dongjun Jang, Wooseok Song, Jaeyoon Kim, Chaeyoung Oh, Hyemi Jo, Youngchae Ahn, Sihyun Oh, Hyohyeong Chang, Sunkyoung Kim, Jinsik Lee

机构 * Seoul National University(首尔国立大学) LG AI Research(LG人工智能研究)

专题命中 视觉定位与Grounding :grounding(abstract)

Comments Under Reveiw

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10454 2025-05-22 cs.HC 50%

Emotion-sensitive Explanation Model

Christian Schütze, Birte Richter, Britta Wrede

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.18132 2025-05-21 cs.CL 50%

MathAgent: Leveraging a Mixture-of-Math-Agent Framework for Real-World Multimodal Mathematical Error Detection

Yibo Yan, Shen Wang, Jiahao Huo, Philip S. Yu, Xuming Hu, Qingsong Wen

机构 * Squirrel Ai Learning The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) The Hong Kong University of Science and Technology(香港科学与技术大学) University of Illinois at Chicago(伊利诺伊大学香槟分校)

专题命中 视觉定位与Grounding :multimodal large language model(abstract)

Comments Accepted by The 63rd Annual Meeting of the Association for Computational Linguistics (ACL Industry 2025, Oral Presentation)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13450 2025-05-21 cs.LO math.FA 50%

Fractal Analysis on the Real Interval: A Constructive Approach via Fractal Countability

Stanislav Semenov

专题命中 视觉定位与Grounding :grounding(abstract)

Comments 43 pages, submitted to arXiv

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11680 2025-05-20 cs.RO 50%

Grounded Task Axes: Zero-Shot Semantic Skill Generalization via Task-Axis Controllers and Visual Foundation Models

M. Yunus Seker, Shobhit Aggarwal, Oliver Kroemer

机构 * Carnegie Mellon University(卡内基梅隆大学) The Robotics Institute(机器人研究所)

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.09823 2025-05-16 cs.HC 50%

WhatsAI: Transforming Meta Ray-Bans into an Extensible Generative AI Platform for Accessibility

Nasif Zaman, Venkatesh Potluri, Brandon Biggs, James M. Coughlan

专题命中 视觉定位与Grounding :visual language model(abstract)

Comments 6 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.06045 2025-05-12 cs.HC 50%

Designing RoutScape: Geospatial Prototyping with XR for Flood Evacuation Planning

Johndayll Lewis Arizala, Joshua Permito, Steven Errol Escopete, John Kovie Niño, Jordan Aiko Deja

专题命中 视觉定位与Grounding :grounding(abstract)

Comments 6 pages, 3 figures

Journal ref Proceedings of the CHIRP 2025: Transforming HCI Research in the Philippines Workshop, May 08, 2025, Baguio, Benguet, Philippines

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.18225 2025-04-28 cs.CL 50%

Even Small Reasoners Should Quote Their Sources: Introducing the Pleias-RAG Model Family

Pierre-Carl Langlais, Pavel Chizhov, Mattia Nee, Carlos Rosas Hinostroza, Matthieu Delsart, Irène Girard, Othman Hicheur, Anastasia Stasenko, Ivan P. Yamshchikov

机构 * PleIAs, Paris, France(巴黎法国PleIAs)

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.17149 2025-04-25 math.HO 50%

Towards a Critical Pragmatic Philosophy of Sustainable Mathematics Education

Dennis Müller

专题命中 视觉定位与Grounding :grounding(abstract)

Comments 22 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09435 2025-04-15 cs.HC 50%

Design Probes for AI-Driven AAC: Addressing Complex Communication Needs in Aphasia

Lei Mao, Jong Ho Lee, Yasmeen Faroqi Shah, Stephanie Valencia

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.03216 2025-04-07 cond-mat.soft 50%

Topological sorting of magnetic colloidal bipeds

Aneena Rinu Perayil, Piotr Kuświk, Maciej Urbaniak, Feliks Stobiecki, Sapida Akhundzada, Arno Ehresmann, Daniel de las Heras, Thomas M. Fischer

专题命中 视觉定位与Grounding :grounding(abstract)

Journal ref Soft Matter, 21, 2716-2722, (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.17703 2025-04-07 cs.RO 50%

RAIDER: Tool-Equipped Large Language Model Agent for Robotic Action Issue Detection, Explanation and Recovery

Silvia Izquierdo-Badiola, Carlos Rizzo, Guillem Alenyà

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.04188 2025-04-04 cs.CL cs.IR 50%

Measuring temporal effects of agent knowledge by date-controlled tool use

R. Patrick Xian, Qiming Cui, Stefan Bauer, Reza Abbasi-Asl

专题命中 视觉定位与Grounding :grounding(abstract)

Comments under review, comments welcome

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.02197 2025-04-04 cs.ET cs.HC 50%

Design and Implementation of the Transparent, Interpretable, and Multimodal (TIM) AR Personal Assistant

Erin McGowan, Joao Rulff, Sonia Castelo, Guande Wu, Shaoyu Chen, Roque Lopez, Bea Steers, Iran R. Roman, Fabio F. Dias, Jing Qian, Parikshit Solunke, Michael Middleton, Ryan McKendrick, Claudio T. Silva

专题命中 视觉定位与Grounding :grounding(abstract)

Comments Copyright 2025 IEEE. All rights reserved, including rights for text and data mining and training of artificial intelligence and similar technologies. Personal use is permitted, but republication/redistribution requires IEEE permission. Article accepted for publication in IEEE Computer Graphics and Applications. This is the author's version, content may change prior to final publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.24225 2025-04-01 physics.flu-dyn math-ph math.AP math.MP 50%

Compressible N-phase fluid mixture models

M. F. P. ten Eikelder, E. H. van Brummelen, D. Schillinger

专题命中 视觉定位与Grounding :grounding(abstract)

Comments preprint, 50 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.23873 2025-04-01 eess.AS cs.SD 50%

Exploring In-Context Learning Capabilities of ChatGPT for Pathological Speech Detection

Mahdi Amiri, Hatef Otroshi Shahreza, Ina Kodrasi

专题命中 视觉定位与Grounding :multimodal large language model(abstract)

Comments submitted to EUSIPCO 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.21833 2025-03-31 cs.CL 50%

Refining Time Series Anomaly Detectors using Large Language Models

Alan Yang, Yulin Chen, Sean Lee, Venus Montes

专题命中 视觉定位与Grounding :multimodal large language model(abstract)

Comments Main content: 4 pages, 1 figure, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.21264 2025-03-28 math.LO 50%

A Direct Characterisation of Logical Grounds and a Decidability Proof

Francesco A. Genco

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.18174 2025-03-25 cs.CL cs.IR 50%

GINGER: Grounded Information Nugget-Based Generation of Responses

Weronika Łajewska, Krisztian Balog

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.16510 2025-03-24 cs.HC cs.CY 50%

Combating the Effects of Cyber-Psychosis: Using Object Security to Facilitate Critical Thinking

Robert H. Thomson, Quan Nguyen, Essien Ayanam, Matthew Canham, Thomas C. Schmidt, Matthias Wählisch, Eric Osterweil

专题命中 视觉定位与Grounding :grounding(abstract)

Comments 13 pages, 3 figures, under submission

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.15682 2025-03-21 cs.CY 50%

Transfeminist AI Governance

Blair Attard-Frost

专题命中 视觉定位与Grounding :grounding(abstract)

Comments 37 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.15212 2025-03-20 eess.IV 50%

Context-Aware Vision Language Foundation Models for Ocular Disease Screening in Retinal Images

Lucie Berger, Mathieu Lamard, Philippe Zhang, Laurent Borderie, Alexandre Le Guilcher, Pascale Massin, Béatrice Cochener, Gwenolé Quellec, Sarah Matta

专题命中 视觉定位与Grounding :vision-language model(abstract)

Comments 4 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.11989 2025-03-20 cs.RO 50%

Dynamic Open-Vocabulary 3D Scene Graphs for Long-term Language-Guided Mobile Manipulation

Zhijie Yan, Shufei Li, Zuoxu Wang, Lixiu Wu, Han Wang, Jun Zhu, Lijiang Chen, Jihong Liu

专题命中 视觉定位与Grounding :vision-language model(abstract)

Comments Accepted by IEEE Robotics and Automation Letters (RA-L), 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.09902 2025-03-14 cs.IR 50%

Conversational Gold: Evaluating Personalized Conversational Search System using Gold Nuggets

Zahra Abbasiantaeb, Simon Lupart, Leif Azzopardi, Jeffery Dalton, Mohammad Aliannejadi

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.09003 2025-03-13 cs.IR cs.CL 50%

Leveraging Retrieval Augmented Generative LLMs For Automated Metadata Description Generation to Enhance Data Catalogs

Mayank Singh, Abhijeet Kumar, Sasidhar Donaparthi, Gayatri Karambelkar

专题命中 视觉定位与Grounding :grounding(abstract)

Comments Presented in 5th International Conference on NLP & Text Mining (NLTM 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏