arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 7473 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7473 篇

2510.04049 2025-10-07 cs.PL 50%

Encoding Numeric Computations and Infusing Heuristic Knowledge Using Integrity Constraints in stableKanren

Xiangyu Guo, Ajay Bansal

专题命中 视觉定位与Grounding :grounding(abstract)

Comments 12 pages, 2 figures, ICFP '25 The miniKanren and Relational Programming Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04041 2025-10-07 cs.RO 50%

SITCOM: Scaling Inference-Time COMpute for VLAs

Ayudh Saxena, Harsh Shah, Sandeep Routray, Rishi Rajesh Shah, Esha Pahwa

机构 * Carnegie Mellon University(卡内基梅隆大学)

专题命中 视觉定位与Grounding :grounding(abstract)

Comments Accepted at the NeurIPS 2025 Workshop on Space in Vision, Language, and Embodied AI (SpaVLE). *Equal contribution

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05135 2025-10-07 cs.RO 50%

LERa: Replanning with Visual Feedback in Instruction Following

Svyatoslav Pchelintsev, Maxim Patratskiy, Anatoly Onishchenko, Alexandr Korchemnyi, Aleksandr Medvedev, Uliana Vinogradova, Ilya Galuzinsky, Aleksey Postnikov, Alexey K. Kovalev, Aleksandr I. Panov

机构 * MIPT(莫斯科国立交通大学) Sberbank of Russia, Robotics Center(俄罗斯储蓄银行机器人中心) AIRI

专题命中 视觉定位与Grounding :visual language model(abstract)

Comments Accepted to IROS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02044 2025-10-03 cs.CL cs.SD eess.AS 50%

Stream RAG: Instant and Accurate Spoken Dialogue Systems with Streaming Tool Usage

Siddhant Arora, Haidar Khan, Kai Sun, Xin Luna Dong, Sajal Choudhary, Seungwhan Moon, Xinyuan Zhang, Adithya Sagar, Surya Teja Appini, Kaushik Patnaik, Sanat Sharma, Shinji Watanabe, Anuj Kumar, Ahmed Aly, Yue Liu, Florian Metze, Zhaojiang Lin

机构 * Meta Carnegie Mellon University(卡内基梅隆大学)

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21745 2025-10-02 physics.med-ph 50%

Design and performance of a Toroidal RF Volume Coil with Intrinsic Electromagnetic Interference Rejection for low-field Portable Halbach-Based MRI Systems

Jules Vliem, Najac Chloe, Beatrice Lena, Andrew Webb, Irena Zivkovic

专题命中 视觉定位与Grounding :grounding(abstract)

Comments 20 pages, 6 figures. Magn Reson Med. 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26093 2025-10-02 cs.CL 50%

Reinforced Strategy Optimization for Conversational Recommender Systems via Network-of-Experts

Xiaoyan Zhao, Ming Yan, Yang Zhang, Yang Deng, Jian Wang, Fengbin Zhu, Yilun Qiu, Hong Cheng, Tat-Seng Chua

机构 * The Chinese University of Hong Kong(香港中文大学) University of Science and Technology of China(中国科学技术大学) National University of Singapore(新加坡国立大学) Singapore Management University(新加坡管理学院) The Hong Kong Polytechnic University(香港理工大学)

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24212 2025-09-30 cs.CL 50%

ScenarioBench: Trace-Grounded Compliance Evaluation for Text-to-SQL and RAG

Zahra Atf, Peter R Lewis

机构 * Faculty of Business and Information Technology(商业与信息技术学院) Ontario Tech University(安大略技术大学)

专题命中 视觉定位与Grounding :grounding(abstract)

Comments Accepted for presentation at the LLMs Meet Databases (LMD) Workshop, 35th IEEE International Conference on Collaborative Advances in Software and Computing, 2025. Workshop website: https://sites.google.com/view/lmd2025/home

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06610 2025-09-30 math.GM 50%

Mirror Duality in a Spencer-Type Complex: Analytic and Riemann-Roch Perspectives

Dongzhe Zheng

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21978 2025-09-29 cs.CL 50%

MotivGraph-SoIQ: Integrating Motivational Knowledge Graphs and Socratic Dialogue for Enhanced LLM Ideation

Xinping Lei, Tong Zhou, Yubo Chen, Kang Liu, Jun Zhao

机构 * The Key Laboratory of Cognition and Decision Intelligence for Complex Systems(认知与决策智能复杂系统重点实验室) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Hunan Provincial Key Laboratory of Philosophy and Social Sciences of Artificial Intelligence and Precision International, Hunan Normal University(湖南省人工智能与精准哲学社会科学省级重点实验室,湖南师范大学) Beijing University of Posts and Telecommunications(北京邮电大学)

专题命中 视觉定位与Grounding :grounding(abstract)

Comments EMNLP2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.05926 2025-09-25 cs.CL 50%

LLMs Reproduce Stereotypes of Sexual and Gender Minorities

Ruby Ostrow, Adam Lopez

机构 * University of Edinburgh(爱丁堡大学)

专题命中 视觉定位与Grounding :grounding(abstract)

Comments 13 pages, 5 figures, 9 tables (including bibliography and appendix). Accepted to Findings of EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19573 2025-09-25 cs.RO 50%

Chasing Stability: Humanoid Running via Control Lyapunov Function Guided Reinforcement Learning

Zachary Olkin, Kejun Li, William D. Compton, Aaron D. Ames

机构 * Technology Innovation Institute (TII)(技术创新研究所)

专题命中 视觉定位与Grounding :grounding(abstract)

Comments Submitted to ICRA 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16254 2025-09-25 cs.HC 50%

Reassessing Collaborative Writing Theories and Frameworks in the Age of LLMs: What Still Applies and What We Must Leave Behind

Daisuke Yukita, Tim Miller, Joel Mackenzie

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18509 2025-09-24 cs.CY 50%

Developing a Decolonial Mindset for Indigenising Computing Education (CE)

Jianhua Li, Yin Paradies, Trina Myers, Robin Doss, Armita Zarnegar, Jack Reis

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17727 2025-09-23 cs.CY cs.IT math.IT 50%

Empirical AI Ethics: Reconfiguring Ethics towards a Situated, Plural, and Transformative Approach

Paula Helm, Selin Gerlek

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17523 2025-09-23 cs.CL eess.AS 50%

Leveraging Audio-Visual Data to Reduce the Multilingual Gap in Self-Supervised Speech Models

María Andrea Cruz Blandón, Zakaria Aldeneh, Jie Chi, Maureen de Seyssel

机构 * Tampere University Apple(塔尔库大学苹果)

专题命中 视觉定位与Grounding :grounding(abstract)

Comments 5 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16670 2025-09-23 cs.SD cs.MM eess.AS 50%

Speech-to-See: End-to-End Speech-Driven Open-Set Object Detection

Wenhuan Lu, Xinyue Song, Wenjun Ke, Zhizhi Yu, Wenhao Yang, Jianguo Wei

机构 * College of Intelligence and Computing, Tianjin University, Tianjin, China(智能与计算学院,天津大学,天津,中国) PipeChina Institute of Science and Technology, Tianjin, China(中石油科技研究院,天津,中国)

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.08851 2025-09-23 cs.RO 50%

OTAS: Open-vocabulary Token Alignment for Outdoor Segmentation

Simon Schwaiger, Stefan Thalhammer, Wilfried Wöber, Gerald Steinbauer-Wagner

机构 * Graz University of Technology, Faculty of Computer Science and Biomedical Engineering, Institute of Software Engineering and Artificial Intelligence(格拉茨技术大学,计算机科学与生物医学工程学院,软件工程与人工智能研究所) University of Applied Sciences Technikum Wien, Faculty of Industrial Engineering, Research Group Digital Manufacturing, Automation and Robotics(应用科学大学技术学院,工业工程学院,数字制造、自动化与机器人研究组) University of Natural Resources and Life Sciences, Department of Integrative Biology and Biodiversity Research, Institute for Integrative Nature Conservation Research(自然资源与生命科学大学,整合生物学与生物多样性研究部门,整合自然保护研究 institute)

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17602 2025-09-18 eess.IV 50%

Attention-ResUNet and EfficientSASM-UNet: UNet based frameworks for Lung and Nodule segmentation

Muhammad Abdullah, Furqan Shaukat

专题命中 视觉定位与Grounding :vision language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09281 2025-09-12 cs.HC 50%

Flip Co-op: Cooperative Takeovers in Shared Autonomy

Sandeep Banik, Naira Hovakimyan

专题命中 视觉定位与Grounding :grounding(abstract)

Comments 11 pages and 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13107 2025-09-11 cs.CL cs.IR 50%

All for law and law for all: Adaptive RAG Pipeline for Legal Research

Figarri Keisha, Prince Singh, Pallavi, Dion Fernandes, Aravindh Manivannan, Ilham Wicaksono, Faisal Ahmad, Wiem Ben Rim

机构 * University College London(伦敦大学学院)

专题命中 视觉定位与Grounding :grounding(abstract)

Comments submitted to NLLP 2025 Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02100 2025-09-09 cs.HC cs.CL 50%

E-THER: A Multimodal Dataset for Empathic AI -- Towards Emotional Mismatch Awareness

Sharjeel Tahir, Judith Johnson, Jumana Abu-Khalaf, Syed Afaq Ali Shah

机构 * Centre for AI and ML, Edith Cowan University(人工智能与机器学习中心,埃德温·科温大学) University of Manchester(曼彻斯特大学)

专题命中 视觉定位与Grounding :vision-language model(abstract)

Comments 15 pages, 4 figures. Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11838 2025-09-09 physics.med-ph cond-mat.mtrl-sci physics.app-ph 50%

Deformation Driven Suction Cups: A Mechanics-Based Approach to Wearable Electronics

Seola Lee, Andrew Akerson, Roham Pardakhtim, Ehsan Hajiesmaili, Kevin Rhodes, Zidong Li, Andrew Stanley, Amirhossein Amini, Daniele Piazza, Chiara Daraio, Tianshu Liu

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.18596 2025-09-09 cs.CL 50%

LinkAlign: Scalable Schema Linking for Real-World Large-Scale Multi-Database Text-to-SQL

Yihan Wang, Peiyu Liu, Xin Yang

机构 * China Academy of Information and Communications Technology(中国信息通信技术研究院) Renmin University of China(中国人民大学) University of International Business and Economics(国际经济贸易大学)

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04254 2025-09-05 cs.HC 50%

MuMTAffect: A Multimodal Multitask Affective Framework for Personality and Emotion Recognition from Physiological Signals

Meisam Jamshidi Seikavandi, Fabricio Batista Narcizo, Ted Vucurevich, Andrew Burke Dittberner, Paolo Burelli

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.12065 2025-09-05 cs.CL cs.FL 50%

Autoformalization in the Wild: Assessing LLMs on Real-World Mathematical Definitions

Lan Zhang, Marco Valentino, Andre Freitas

机构 * Department of Computer Science, University of Manchester(曼彻斯特大学计算机科学系) School of Computer Science, University of Sheffield(谢菲尔德大学计算机科学学院) Idiap Research Institute(Idiap研究机构) National Biomarker Centre, CRUK Manchester Institute(国家生物标志物中心、CRUK曼彻斯特研究所)

专题命中 视觉定位与Grounding :grounding(abstract)

Comments EMNLP 2025 Camera-Ready Version

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02425 2025-09-03 cs.RO cs.HC 50%

OpenGuide: Assistive Object Retrieval in Indoor Spaces for Individuals with Visual Impairments

Yifan Xu, Qianwei Wang, Vineet Kamat, Carol Menassa

机构 * Department of Civil and Environmental Engineering, University of Michigan(密歇根大学土木与环境工程系) College of Literature, Science, and the Arts, University of Michigan(密歇根大学文学、科学与艺术学院)

专题命中 视觉定位与Grounding :VLM(abstract)

Comments 32 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02367 2025-09-03 cs.HC 50%

Talking Spell: A Wearable System Enabling Real-Time Anthropomorphic Voice Interaction with Everyday Objects

Xuetong Wang, Ching Christie Pang, Pan Hui

专题命中 视觉定位与Grounding :vision-language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18906 2025-09-03 cs.CL cs.IR 50%

Federated Retrieval-Augmented Generation: A Systematic Mapping Study

Abhijit Chakraborty, Chahana Dahal, Vivek Gupta

机构 * Arizona State University(亚利桑那州立大学) University of Nevada, Las Vegas(内华达大学拉斯维加斯分校)

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00325 2025-09-03 cs.CL cs.IR 50%

GIER: Gap-Driven Self-Refinement for Large Language Models

Rinku Dewri

机构 * University of Denver(丹佛大学)

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18497 2025-08-27 quant-ph 50%

Can Classical Initialization Help Variational Quantum Circuits Escape the Barren Plateau?

Yifeng Peng, Xinyi Li, Zhemin Zhang, Samuel Yen-Chi Chen, Zhiding Liang, Ying Wang

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏