arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 26465 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7494 篇

2110.11772 2025-05-14 cs.SI 78%

Grounding force-directed network layouts with latent space models

Felix Gaisbauer, Armin Pournaki, Sven Banisch, Eckehard Olbrich

专题命中 视觉定位与Grounding :grounding(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.06030 2025-05-12 cs.AI cs.CV cs.LG 78%

Why Are You Wrong? Counterfactual Explanations for Language Grounding with 3D Objects

Tobias Preintner, Weixuan Yuan, Qi Huang, Adrian König, Thomas Bäck, Elena Raponi, Niki van Stein

机构 * Institute of Advanced Computer Science, Leiden University(先进计算机科学研究所,莱顿大学) BMW Group(宝马集团)

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV、cs.AI、cs.LG

Comments Accepted at IJCNN 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.00002 2025-05-02 cs.CL 78%

Symbol grounding in computational systems: A paradox of intentions

Vincent C. Müller

专题命中 视觉定位与Grounding :grounding(title,abstract)

Journal ref (2009) Minds and Machines, 19 (4), 529-41

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.16502 2025-03-21 cs.CV cs.AI cs.LG cs.RO 78%

GSplatLoc: Grounding Keypoint Descriptors into 3D Gaussian Splatting for Improved Visual Localization

Gennady Sidorov, Malik Mohrat, Denis Gridusov, Ruslan Rakhimov, Sergey Kolyubin

机构 * ITMO University(ITMO大学) Robotics Center(机器人中心) T-Tech

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV、cs.AI、cs.LG

Comments Project website at https://gsplatloc.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.07223 2025-03-13 cs.RO cs.AI cs.CV cs.LG 78%

Grounding Video Models to Actions through Goal Conditioned Exploration

Yunhao Luo, Yilun Du

机构 * Georgia Tech(佐治亚理工学院) Brown(布朗大学) Harvard(哈佛大学)

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV、cs.AI、cs.LG

Comments ICLR 2025 (Spotlight). Project page: https://video-to-action.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.03495 2025-03-06 cs.CL 78%

Deictic Codes, Demonstratives, and Reference: A Step Toward Solving the Grounding Problem

Athanassios Raftopoulos, Vincent C. Müller

专题命中 视觉定位与Grounding :grounding(title,abstract)

Journal ref (2002) in Wayne D. Gray and Christian D. Schunn (eds.), CogSci 2002, 24th annual meeting of the Cognitive Science Society (Mahwah, NY: Lawrence Erlbaum), 762-67

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.01616 2025-03-04 cs.RO 78%

RoboDexVLM: Visual Language Model-Enabled Task Planning and Motion Control for Dexterous Robot Manipulation

Haichao Liu, Sikai Guo, Pengfei Mai, Jiahang Cao, Haoang Li, Jun Ma

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))

专题命中 视觉定位与Grounding :visual language model(title);vision-language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.11144 2025-02-19 cs.HC 78%

CrossA11y: Identifying Video Accessibility Issues via Cross-modal Grounding

Xingyu "Bruce" Liu, Ruolin Wang, Dingzeyu Li, Xiang 'Anthony' Chen, Amy Pavel

专题命中 视觉定位与Grounding :grounding(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.03888 2025-02-18 cs.CL 78%

Multi3Hate: Multimodal, Multilingual, and Multicultural Hate Speech Detection with Vision-Language Models

Minh Duc Bui, Katharina von der Wense, Anne Lauscher

机构 * Johannes Gutenberg University Mainz(美因茨约翰内斯古滕贝格大学) University of Colorado Boulder(科罗拉多大学博尔德分校) University of Hamburg(汉堡大学)

专题命中 视觉定位与Grounding :vision-language model(title,abstract)

Comments Accepted to NAACL 2025 Main (Camera-Ready Version)

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.02243 2025-02-18 cs.CL q-bio.NC 78%

Language Writ Large: LLMs, ChatGPT, Grounding, Meaning and Understanding

Stevan Harnad

专题命中 视觉定位与Grounding :grounding(title,abstract)

Comments 54 pages, 29 references

Journal ref Frontiers in Artificial Intelligence 7: 1490698 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.12812 2025-02-10 cs.CL 78%

Grounding Fallacies Misrepresenting Scientific Publications in Evidence

Max Glockner, Yufang Hou, Preslav Nakov, Iryna Gurevych

机构 * Ubiquitous Knowledge Processing Lab (UKP Lab)(通用知识处理实验室(UKP实验室)) TU Darmstadt(达姆施塔特工业大学) Hessian Center for AI (hessian.AI)(黑森人工智能中心) IT:U Interdisciplinary Transformation University Austria(奥地利跨学科转型大学(IT:U)) IBM Research Ireland(IBM爱尔兰研究院) MBZUAI(穆罕默德·本·扎耶德人工智能大学)

专题命中 视觉定位与Grounding :grounding(title,abstract)

Comments Accepted to NAACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.05753 2025-02-10 cs.LG cs.AI cs.CV 78%

Grounding Continuous Representations in Geometry: Equivariant Neural Fields

David R Wessels, David M Knigge, Samuele Papa, Riccardo Valperga, Sharvaree Vadgama, Efstratios Gavves, Erik J Bekkers

机构 * University of Amsterdam(阿姆斯特丹大学)

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.10491 2025-01-22 math.LO 78%

Denotational semantics for languages of epistemic grounding based on Prawitz's theory of grounds

Antonio Piccolomini d'Aragona

专题命中 视觉定位与Grounding :grounding(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.10490 2025-01-22 math.LO 78%

Calculi of epistemic grounding based on Prawitz's theory of grounds

Antonio Piccolomini d'Aragona

专题命中 视觉定位与Grounding :grounding(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.03200 2025-01-07 cs.CL 78%

The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input

Alon Jacovi, Andrew Wang, Chris Alberti, Connie Tao, Jon Lipovetz, Kate Olszewska, Lukas Haas, Michelle Liu, Nate Keating, Adam Bloniarz, Carl Saroufim, Corey Fry, Dror Marcus, Doron Kukliansky, Gaurav Singh Tomar, James Swirhun, Jinwei Xing, Lily Wang, Madhu Gurumurthy, Michael Aaron, Moran Ambar, Rachana Fellinger, Rui Wang, Zizhao Zhang, Sasha Goldshtein, Dipanjan Das

机构 * Google(谷歌公司)

专题命中 视觉定位与Grounding :grounding(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.06821 2024-12-11 cs.HC 78%

FinFlier: Automating Graphical Overlays for Financial Visualizations with Knowledge-Grounding Large Language Model

Jianing Hao, Manling Yang, Qing Shi, Yuzhe Jiang, Guang Zhang, Wei Zeng

专题命中 视觉定位与Grounding :grounding(title,abstract)

Comments 17 pages, 13 figures, this paper is published on IEEE Transactions on Visualization and Computer Graphics

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.12960 2024-11-21 cs.RO 78%

I Can Tell What I am Doing: Toward Real-World Natural Language Grounding of Robot Experiences

Zihan Wang, Brian Liang, Varad Dhat, Zander Brumbaugh, Nick Walker, Ranjay Krishna, Maya Cakmak

专题命中 视觉定位与Grounding :grounding(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.11323 2024-11-19 cs.RO 78%

SayComply: Grounding Field Robotic Tasks in Operational Compliance through Retrieval-Based Language Models

Muhammad Fadhil Ginting, Dong-Ki Kim, Sung-Kyun Kim, Bandi Jai Krishna, Mykel J. Kochenderfer, Shayegan Omidshafiei, Ali-akbar Agha-mohammadi

专题命中 视觉定位与Grounding :grounding(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.03359 2024-11-07 cs.CV cs.AI cs.LG 78%

Self-Calibrated Tuning of Vision-Language Models for Out-of-Distribution Detection

Geng Yu, Jianing Zhu, Jiangchao Yao, Bo Han

专题命中 视觉定位与Grounding :vision-language model(title);分类 cs.CV、cs.AI、cs.LG

Comments accepted by NeurIPS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.02118 2024-11-05 cs.HC cs.CL 78%

Grounding Emotional Descriptions to Electrovibration Haptic Signals

Guimin Hu, Zirui Zhao, Lukas Heilmann, Yasemin Vardar, Hasti Seifi

专题命中 视觉定位与Grounding :grounding(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.16472 2024-10-23 cs.CL 78%

DocEdit-v2: Document Structure Editing Via Multimodal LLM Grounding

Manan Suri, Puneet Mathur, Franck Dernoncourt, Rajiv Jain, Vlad I Morariu, Ramit Sawhney, Preslav Nakov, Dinesh Manocha

专题命中 视觉定位与Grounding :grounding(title,abstract)

Comments EMNLP 2024 (Main)

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.09106 2024-10-21 cs.CL 78%

"We Demand Justice!": Towards Social Context Grounding of Political Texts

Rajkumar Pujari, Chengfei Wu, Dan Goldwasser

专题命中 视觉定位与Grounding :grounding(title,abstract)

Comments Accepted as an oral at EMNLP 2024 main conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.05663 2024-10-10 cs.RO 78%

Abstract Hardware Grounding towards the Automated Design of Automation Systems

Yu-Zhe Shi, Qiao Xu, Fanxu Meng, Lecheng Ruan, Qining Wang

专题命中 视觉定位与Grounding :grounding(title,abstract)

Comments In International Conference on Intelligent Robotics and Applications (ICIRA'24)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.04819 2024-10-08 cs.CL 78%

MINER: Mining the Underlying Pattern of Modality-Specific Neurons in Multimodal Large Language Models

Kaichen Huang, Jiahao Huo, Yibo Yan, Kun Wang, Yutao Yue, Xuming Hu

专题命中 视觉定位与Grounding :multimodal large language model(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.16671 2024-10-08 cs.CL 78%

StructLM: Towards Building Generalist Models for Structured Knowledge Grounding

Alex Zhuang, Ge Zhang, Tianyu Zheng, Xinrun Du, Junjie Wang, Weiming Ren, Stephen W. Huang, Jie Fu, Xiang Yue, Wenhu Chen

专题命中 视觉定位与Grounding :grounding(title,abstract)

Comments Technical Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.00518 2024-09-30 cs.RO 78%

When Robots Get Chatty: Grounding Multimodal Human-Robot Conversation and Collaboration

Philipp Allgeuer, Hassan Ali, Stefan Wermter

专题命中 视觉定位与Grounding :grounding(title,abstract)

Journal ref International Conference on Artificial Neural Networks, Sep 2024 (pp. 306-321)

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.07980 2024-08-16 cs.LO 78%

Efficiently grounding FOL using bit vectors

Lucas Van Laer, Simon Vandevelde, Joost Vennekens

专题命中 视觉定位与Grounding :grounding(title,abstract)

Comments Short version published at LPNMR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.01139 2024-08-09 cs.CL 78%

It Couldn't Help But Overhear: On the Limits of Modelling Meta-Communicative Grounding Acts with Supervised Learning

Brielen Madureira, David Schlangen

专题命中 视觉定位与Grounding :grounding(title,abstract)

Comments Accepted to SIGdial 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.14321 2024-07-22 cs.CL cs.IR cs.MM 78%

Multimodal Misinformation Detection using Large Vision-Language Models

Sahar Tahmasebi, Eric Müller-Budack, Ralph Ewerth

专题命中 视觉定位与Grounding :vision-language model(title,abstract)

Comments Accepted for publication in: Conference on Information and Knowledge Management (CIKM) 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.12858 2024-07-19 cs.CL cs.AI cs.CV cs.LG 78%

Grounding and Evaluation for Large Language Models: Practical Challenges and Lessons Learned (Survey)

Krishnaram Kenthapadi, Mehrnoosh Sameki, Ankur Taly

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV、cs.AI、cs.LG

Comments Survey Article for the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD 2024) Tutorial

Journal ref Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD 2024)

详情

展开后加载摘要…

URL PDF HTML 收藏