arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 7473 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7473 篇

2510.05387 2025-11-12 cs.CL 50%

Cross-Lingual Mental Health Ontologies for Indian Languages: Bridging Patient Expression and Clinical Understanding through Explainable AI and Human-in-the-Loop Validation

Ananth Kandala, Ratna Kandala, Akshata Kishore Moharir, Niva Manchanda, Sunaina Singh

专题命中 视觉定位与Grounding :grounding(abstract)

Journal ref Integrating NLP and AI for Multilingual and Patient-Centric Healthcare Communication (NLP-AI4Health) 2025 workshop, IJCNLP-AACL

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06892 2025-11-11 cs.RO 50%

Multi-Agent AI Framework for Road Situation Detection and C-ITS Message Generation

Kailin Tong, Selim Solmaz, Kenan Mujkic, Gottfried Allmer, Bo Leng

机构 * Virtual Vehicle Research GmbH(虚拟车辆研究有限公司) ASFINAG Maut Service GmbH(ASFINAG收费服务有限公司) Tongji University(同济大学)

专题命中 视觉定位与Grounding :multimodal large language model(abstract)

Comments submitted to TRA 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15874 2025-11-11 cs.IR cs.CL 50%

Text-to-Pipeline: Bridging Natural Language and Data Preparation Pipelines

Yuhang Ge, Yachuan Liu, Zhangyan Ye, Yuren Mao, Yunjun Gao

机构 * Zhejiang University(浙江大学)

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05383 2025-11-10 cs.CE 50%

Connectomics Informed by Large Language Models

Elinor Thompson, Tiantian He, Anna Schroder, Ahmed Abdulaal, Alec Sargood, Sonja Soskic, Henry F. J. Tregidgo, Daniel C. Alexander

专题命中 视觉定位与Grounding :grounding(abstract)

Comments 35 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03165 2025-11-06 cs.RO 50%

SENT Map -- Semantically Enhanced Topological Maps with Foundation Models

Raj Surya Rajendran Kathirvel, Zach A Chavis, Stephen J. Guy, Karthik Desingh

机构 * Minnesota Robotics Institute (MnRI)(明尼苏达州机器人研究所) Department of Computer Science and Engineering (CS&E)(计算机科学与工程系) University of Minnesota(明尼苏达大学)

专题命中 视觉定位与Grounding :grounding(abstract)

Comments Accepted at ICRA 2025 Workshop on Foundation Models and Neuro-Symbolic AI for Robotics

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.08677 2025-11-04 quant-ph physics.atom-ph 50%

Multiparameter estimation with an array of entangled atomic sensors

Yifan Li, Lex Joosten, Youcef Baamara, Paolo Colciaghi, Alice Sinatra, Philipp Treutlein, Tilman Zibold

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.05049 2025-11-04 cs.SI cs.CY 50%

Uncovering the Sociodemographic Fabric of Reddit

Federico Cinus, Corrado Monti, Paolo Bajardi, Gianmarco De Francisci Morales

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26015 2025-10-31 cs.HC 50%

Designing for Dignity while Driving: Interaction Needs of Blind and Low-Vision Passengers in Fully Automated Vehicles

Zhengtao Ma, Rafael Gomez, Togtokhtur Batbold, Zishuo Zhu, Yueteng Yu, Ronald Schroeter

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24965 2025-10-30 cs.NE 50%

Exponential Dynamic Energy Network for High Capacity Sequence Memory

Arjun Karuvally, Pichsinee Lertsaroj, Terrence J. Sejnowski, Hava T. Siegelmann

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20854 2025-10-27 econ.GN q-fin.EC 50%

Edgeworth's exact and naturally weighted evolutionary utilitarianism and the happiness of Mr. Pongo

Alberto Baccini

专题命中 视觉定位与Grounding :grounding(abstract)

Comments 37 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11514 2025-10-27 cs.CL 50%

LVLMs are Bad at Overhearing Human Referential Communication

Zhengxiang Wang, Weiling Li, Panagiotis Kaliosis, Owen Rambow, Susan E. Brennan

机构 * Department of Linguistics(语言学系) Institute for Advanced Computational Science(先进计算科学研究所) Department of Psychology(心理学系) Department of Computer Science, Stony Brook University(石溪大学计算机科学系)

专题命中 视觉定位与Grounding :vision language model(abstract)

Comments EMNLP 2025 (Main)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20161 2025-10-24 cs.RO 50%

PathFormer: A Transformer with 3D Grid Constraints for Digital Twin Robot-Arm Trajectory Generation

Ahmed Alanazi, Duy Ho, Yugyung Lee

机构 * Department of Computer Science, University of Missouri–Kansas City (UMKC)(密苏里大学哥伦比亚分校计算机科学系) Department of Computer Science, California State University, Fullerton(加州州立大学富尔顿分校计算机科学系)

专题命中 视觉定位与Grounding :grounding(abstract)

Comments 8 pages, 7 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19878 2025-10-23 cs.CL cs.IR 50%

CausalRAG: Integrating Causal Graphs into Retrieval-Augmented Generation

Nengbo Wang, Xiaotian Han, Jagdip Singh, Jing Ma, Vipin Chaudhary

机构 * Department of Computer and Data Sciences, Case Western Reserve University(计算机与数据科学系,凯斯西储大学) Department of Design and Innovation, Case Western Reserve University(设计与创新系,凯斯西储大学)

专题命中 视觉定位与Grounding :grounding(abstract)

Comments Accepted at Findings of ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.15845 2025-10-22 math.ST math.OC stat.TH 50%

On Learning the Optimal Regularization Parameter in Inverse Problems

Jonathan Chirinos Rodriguez, Ernesto De Vito, Cesare Molinari, Lorenzo Rosasco, Silvia Villa

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16500 2025-10-21 cs.RO 50%

Advancing Off-Road Autonomous Driving: The Large-Scale ORAD-3D Dataset and Comprehensive Benchmarks

Chen Min, Jilin Mei, Heng Zhai, Shuai Wang, Tong Sun, Fanjie Kong, Haoyang Li, Fangyuan Mao, Fuyang Liu, Shuo Wang, Yiming Nie, Qi Zhu, Liang Xiao, Dawei Zhao, Yu Hu

机构 * Research Center for Intelligent Computing Systems, SKLP, Institute of Computing Technology, Chinese Academy of Sciences, Beijing, China, 100190(中国科学院计算技术研究所,智能计算系统研究中心,SKLP,北京,中国,100190) Tongji University, Shanghai, China, 200092(同济大学,上海,中国,200092) Xi’an Jiaotong University, Shaanxi, China, 710049(西安交通大学,陕西,中国,710049) Nanchang University, Jiangxi, China, 330047(南昌大学,江西,中国,330047) Defense Innovation Institute, Beijing, China, 100073(国防科技创新院,北京,中国,100073)

专题命中 视觉定位与Grounding :vision-language model(abstract)

Comments Off-road robotics

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.09542 2025-10-21 cs.CL 50%

KG-Infused RAG: Augmenting Corpus-Based RAG with External Knowledge Graphs

Dingjun Wu, Yukun Yan, Zhenghao Liu, Zhiyuan Liu, Maosong Sun

机构 * Tsinghua University(清华大学) Northeastern University(东北大学)

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12384 2025-10-17 cs.DC cs.DB 50%

Exploring Distributed Vector Databases Performance on HPC Platforms: A Study with Qdrant

Seth Ockerman, Amal Gueroudji, Song Young Oh, Robert Underwood, Nicholas Chia, Kyle Chard, Robert Ross, Shivaram Venkataraman

专题命中 视觉定位与Grounding :grounding(abstract)

Comments To appear in the SC'25 Workshop Frontiers in Generative AI for HPC Science and Engineering: Foundations, Challenges, and Opportunities

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12766 2025-10-15 cs.CL 50%

Language Models Model Language

Łukasz Borchmann

机构 * Snowflake AI Research(Snowflake人工智能研究)

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12211 2025-10-15 cs.IR 50%

Reinforced Preference Optimization for Recommendation

Junfei Tan, Yuxin Chen, An Zhang, Junguang Jiang, Bin Liu, Ziru Xu, Han Zhu, Jian Xu, Bo Zheng, Xiang Wang

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11897 2025-10-15 cs.HC 50%

A Longitudinal Study on Different Annotator Feedback Loops in Complex RAG Tasks

Sara Rosenthal, Maeda Hanafi, Yannis Katsis, Lucian Popa, Marina Danilevsky

专题命中 视觉定位与Grounding :grounding(abstract)

Comments 26 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09720 2025-10-14 physics.ed-ph 50%

NotebookLM as a Socratic physics tutor: Design and preliminary observations of a RAG-based tool

Eugenio Tufino

专题命中 视觉定位与Grounding :grounding(abstract)

Comments 9 pages, 5 figures. Revised version accepted for publication in The Physics Educator (World Scientific, in press)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00683 2025-10-14 cs.SD eess.AS 50%

PicoAudio2: Temporal Controllable Text-to-Audio Generation with Natural Language Description

Zihao Zheng, Zeyu Xie, Xuenan Xu, Wen Wu, Chao Zhang, Mengyue Wu

机构 * MoE Key Lab of Artificial Intelligence, X-LANCE Lab, Shanghai Jiao Tong University(人工智能联合实验室、X-LANCE实验室、上海交通大学) Shanghai AI Lab(上海人工智能实验室)

专题命中 视觉定位与Grounding :grounding(abstract)

Comments Demo page: https://HiRookie9.github.io/PicoAudio2-Page

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10168 2025-10-13 stat.ME math.ST stat.TH 50%

Statistical methods: Basic concepts, interpretations, and cautions

Sander Greenland

专题命中 视觉定位与Grounding :grounding(abstract)

Comments 64 pages. For Pigeot I, Ahrens W, eds., Handbook of Epidemiology, 3rd edn. Springer, 2025, Ch. 54-1

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08453 2025-10-10 cs.GT 50%

Extending Games beyond the Finite Horizon

Kiri Sakahara, Takashi Sato

专题命中 视觉定位与Grounding :grounding(abstract)

Comments 34 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08111 2025-10-10 cs.CL cs.CY 50%

Evaluating LLM-Generated Legal Explanations for Regulatory Compliance in Social Media Influencer Marketing

Haoyang Gui, Thales Bertaglia, Taylor Annabell, Catalina Goanta, Tjomme Dooper, Gerasimos Spanakis

机构 * Utrecht University(乌特雷赫大学) Stichting Reclame Code(Reclame Code基金会) Maastricht University(马斯特里赫特大学)

专题命中 视觉定位与Grounding :grounding(abstract)

Comments Accepted for publication at the Natural Legal Language Processing Workshop (NLLP) 2025, co-located with EMNLP

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12637 2025-10-10 cs.CL 50%

How Grounded is Wikipedia? A Study on Structured Evidential Support and Retrieval

William Walden, Kathryn Ricci, Miriam Wanner, Zhengping Jiang, Chandler May, Rongkun Zhou, Benjamin Van Durme

机构 * Johns Hopkins University(约翰霍普金斯大学)

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06411 2025-10-09 cs.CL 50%

Instructional Goal-Aligned Question Generation for Student Evaluation in Virtual Lab Settings: How Closely Do LLMs Actually Align?

R. Alexander Knipper, Indrani Dey, Souvika Sarkar, Hari Narayanan, Sadhana Puntambekar, Santu Karmaker

机构 * Department of EdPsych, University of Wisconsin-Madison(威斯康星大学麦迪逊分校教育心理学系) Department of CS, Wichita State University(威斯康星州立大学Wichita分校计算机科学系) Department of CSSE, Auburn University(阿伯茨罕大学计算机科学与工程系)

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06354 2025-10-09 cs.CL 50%

LLM Bias Detection and Mitigation through the Lens of Desired Distributions

Ingroj Shrestha, Padmini Srinivasan

机构 * University of Iowa(爱荷华大学)

专题命中 视觉定位与Grounding :grounding(abstract)

Comments Accepted to EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05619 2025-10-08 eess.AS 50%

Teaching Machines to Speak Using Articulatory Control

Akshay Anand, Chenxu Guo, Cheol Jun Cho, Jiachen Lian, Gopala Anumanchipalli

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04569 2025-10-07 q-fin.TR 50%

Risk-Sensitive Option Market Making with Arbitrage-Free eSSVI Surfaces: A Constrained RL and Stochastic Control Bridge

Jian'an Zhang

专题命中 视觉定位与Grounding :grounding(abstract)

Comments 34 pages including appendices; figures included. Primary subject class: q-fin.TR. Cross-lists: cs.LG; q-fin.CP

详情

展开后加载摘要…

URL PDF HTML 收藏