arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 7473 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7473 篇

2503.08593 2025-03-12 cs.RO 50%

Proc4Gem: Foundation models for physical agency through procedural generation

Yixin Lin, Jan Humplik, Sandy H. Huang, Leonard Hasenclever, Francesco Romano, Stefano Saliceti, Daniel Zheng, Jose Enrique Chen, Catarina Barros, Adrian Collister, Matt Young, Adil Dostmohamed, Ben Moran, Ken Caluwaerts, Marissa Giustina, Joss Moore, Kieran Connell, Francesco Nori, Nicolas Heess, Steven Bohez, Arunkumar Byravan

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.05806 2025-03-11 q-bio.NC 50%

I Think, Therefore I Hallucinate: Minds, Machines, and the Art of Being Wrong

Sebastian Barros

专题命中 视觉定位与Grounding :grounding(abstract)

Comments 25 pages, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.05623 2025-03-10 cs.RO 50%

Limits of specifiability for sensor-based robotic planning tasks

Basak Sakcak, Dylan A. Shell, Jason M. O'Kane

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.04421 2025-03-07 cs.CL 50%

Revisiting the Othello World Model Hypothesis

Yifei Yuan, Anders Søgaard

专题命中 视觉定位与Grounding :grounding(abstract)

Comments ICLR World Models Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.02318 2025-03-06 cs.RO 50%

ZeroCAP: Zero-Shot Multi-Robot Context Aware Pattern Formation via Large Language Models

Vishnunandan L. N. Venkatesh, Byung-Cheol Min

专题命中 视觉定位与Grounding :vision-language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.02698 2025-03-05 cs.RO 50%

FlowPlan: Zero-Shot Task Planning with LLM Flow Engineering for Robotic Instruction Following

Zijun Lin, Chao Tang, Hanjing Ye, Hong Zhang

专题命中 视觉定位与Grounding :grounding(abstract)

Comments 8 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.02106 2025-03-05 cs.RO 50%

OVAMOS: A Framework for Open-Vocabulary Multi-Object Search in Unknown Environments

Qianwei Wang, Yifan Xu, Vineet Kamat, Carol Menassa

专题命中 视觉定位与Grounding :VLM(abstract)

Comments 7 pages, 4 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.10352 2025-03-05 cs.CL 50%

Agentic Verification for Ambiguous Query Disambiguation

Youngwon Lee, Seung-won Hwang, Ruofan Wu, Feng Yan, Danmei Xu, Moutasem Akkad, Zhewei Yao, Yuxiong He

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.08168 2025-03-05 cs.CL 50%

SARChat-Bench-2M: A Multi-Task Vision-Language Benchmark for SAR Image Interpretation

Zhiming Ma, Xiayang Xiao, Sihao Dong, Peidong Wang, HaiPeng Wang, Qingyun Pan

专题命中 视觉定位与Grounding :vision language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.16533 2025-03-05 cs.CL 50%

Tool Learning in the Wild: Empowering Language Models as Automatic Tool Agents

Zhengliang Shi, Shen Gao, Lingyong Yan, Yue Feng, Xiuyi Chen, Zhumin Chen, Dawei Yin, Suzan Verberne, Zhaochun Ren

专题命中 视觉定位与Grounding :grounding(abstract)

Comments Accepted by WWW 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.20547 2025-03-03 cs.PL 50%

An Attempt to Catch Up with JIT Compilers: The False Lead of Optimizing Inline Caches

Aurore Poirier, Erven Rohou, Manuel Serrano

专题命中 视觉定位与Grounding :grounding(abstract)

Journal ref The Art, Science, and Engineering of Programming, 2025, Vol. 10, Issue 1, Article 6

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.20542 2025-03-03 cs.PL 50%

Conversational Concurrency with Dataspaces and Facets

Sam Caldwell, Tony Garnock-Jones, Matthias Felleisen

专题命中 视觉定位与Grounding :grounding(abstract)

Journal ref The Art, Science, and Engineering of Programming, 2025, Vol. 10, Issue 1, Article 2

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.20540 2025-03-03 cs.PL 50%

Study of the Use of Property Probes in an Educational Setting

Anton Risberg Alaküla, Niklas Fors, Emma Söderberg

专题命中 视觉定位与Grounding :grounding(abstract)

Journal ref The Art, Science, and Engineering of Programming, 2025, Vol. 10, Issue 1, Article 10

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.20538 2025-03-03 cs.PL 50%

Skitter: A Distributed Stream Processing Framework with Pluggable Distribution Strategies

Mathijs Saey, Joeri De Koster, Wolfgang De Meuter

专题命中 视觉定位与Grounding :grounding(abstract)

Journal ref The Art, Science, and Engineering of Programming, 2025, Vol. 10, Issue 1, Article 4

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.20534 2025-03-03 cs.PL 50%

Consistent Distributed Reactive Programming with Retroactive Computation

Tetsuo Kamina, Tomoyuki Aotani, Hidehiko Masuhara

专题命中 视觉定位与Grounding :grounding(abstract)

Journal ref The Art, Science, and Engineering of Programming, 2025, Vol. 10, Issue 1, Article 11

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.20530 2025-03-03 cs.PL 50%

Evolution Language Framework for Persistent Objects

Tetsuo Kamina, Tomoyuki Aotani, Hidehiko Masuhara

专题命中 视觉定位与Grounding :grounding(abstract)

Journal ref The Art, Science, and Engineering of Programming, 2025, Vol. 10, Issue 1, Article 12

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.20526 2025-03-03 cs.PL 50%

Two Approaches for Programming Education in the Domain of Graphics: An Experiment

Luca Chiodini, Juha Sorva, Arto Hellas, Otto Seppälä, Matthias Hauswirth

专题命中 视觉定位与Grounding :grounding(abstract)

Journal ref The Art, Science, and Engineering of Programming, 2025, Vol. 10, Issue 1, Article 14

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.16383 2025-02-27 cs.HC 50%

Understanding Generative AI Risks for Youth: A Taxonomy Based on Empirical Data

Yaman Yu, Yiren Liu, Jacky Zhang, Yun Huang, Yang Wang

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.18313 2025-02-26 cs.CL 50%

Looking forward: Linguistic theory and methods

John Mansfield, Ethan Gotlieb Wilcox

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.06970 2025-02-26 cs.CL 50%

Can Visual Dialogue Models Do Scorekeeping? Exploring How Dialogue Representations Incrementally Encode Shared Knowledge

Brielen Madureira, David Schlangen

专题命中 视觉定位与Grounding :grounding(abstract)

Comments Accepted at ACL 2022, short paper (v2 fixes labels in Figure 3)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.16528 2025-02-25 cs.RO 50%

OpenVox: Real-time Instance-level Open-vocabulary Probabilistic Voxel Representation

Yinan Deng, Bicheng Yao, Yihang Tang, Yi Yang, Yufeng Yue

专题命中 视觉定位与Grounding :vision-language model(abstract)

Comments Project website: https://open-vox.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.15366 2025-02-24 cs.RO 50%

Rapid Online Learning of Hip Exoskeleton Assistance Preferences

Giulia Ramella, Auke Ijspeert, Mohamed Bouri

专题命中 视觉定位与Grounding :grounding(abstract)

Comments Copyright 2025 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works

Journal ref 2025 IEEE International Conference on Robotics and Automation (ICRA)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.14501 2025-02-21 cs.CL 50%

Towards a Perspectivist Turn in Argument Quality Assessment

Julia Romberg, Maximilian Maurer, Henning Wachsmuth, Gabriella Lapesa

专题命中 视觉定位与Grounding :grounding(abstract)

Comments Accepted to NAACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.02353 2025-02-21 cs.HC 50%

Social-RAG: Retrieving from Group Interactions to Socially Ground AI Generation

Ruotong Wang, Xinyi Zhou, Lin Qiu, Joseph Chee Chang, Jonathan Bragg, Amy X. Zhang

专题命中 视觉定位与Grounding :grounding(abstract)

Comments To appear at CHI2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11890 2025-02-19 cs.CL 50%

Revisiting Classification Taxonomy for Grammatical Errors

Deqing Zou, Jingheng Ye, Yulu Liu, Yu Wu, Zishan Xu, Yinghui Li, Hai-Tao Zheng, Bingxu An, Zhao Wei, Yong Xu

专题命中 视觉定位与Grounding :grounding(abstract)

Comments 26 pages, 4 figures and 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11073 2025-02-18 cs.CL 50%

Demystifying Hateful Content: Leveraging Large Multimodal Models for Hateful Meme Detection with Explainable Decisions

Ming Shan Hee, Roy Ka-Wei Lee

专题命中 视觉定位与Grounding :vision-language model(abstract)

Comments Preprint. Accepted at ICWSM'25

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.14497 2025-02-17 cs.CL 50%

Evaluating and Improving Graph to Text Generation with Large Language Models

Jie He, Yijun Yang, Wanqiu Long, Deyi Xiong, Victor Gutierrez-Basulto, Jeff Z. Pan

专题命中 视觉定位与Grounding :grounding(abstract)

Comments NAACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.09362 2025-02-14 cs.HC 50%

Let's Talk Futures: A Literature Review of HCI's Future-Orientation

Camilo Sanchez, Sui Wang, Kaisa Savolainen, Felix Anand Epp, Antti Salovaara

专题命中 视觉定位与Grounding :grounding(abstract)

Comments CHI Conference on Human Factors in Computing Systems (CHI '25), April 26-May 1, 2025, Yokohama, Japan

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.09236 2025-02-14 cs.LO cs.SE 50%

Early Validation of High-level Requirements on Cyber-Physical Systems

Ondřej Vašíček

专题命中 视觉定位与Grounding :grounding(abstract)

Comments In Proceedings ICLP 2024, arXiv:2502.08453

Journal ref EPTCS 416, 2025, pp. 390-397

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.07544 2025-02-12 cs.CL 50%

Grammar Control in Dialogue Response Generation for Language Learning Chatbots

Dominik Glandorf, Peng Cui, Detmar Meurers, Mrinmaya Sachan

专题命中 视觉定位与Grounding :grounding(abstract)

Comments Accepted to NAACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏