arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-07-31 至 2025-07-31 共收录 7 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7 篇

2501.06680 2025-07-31 cs.CV cs.AI cs.LG cs.RO 82%

Application of Vision-Language Model to Pedestrians Behavior and Scene Understanding in Autonomous Driving

Haoxiang Gao, Li Zhang, Yu Zhao, Zhou Yang, Jinghan Cao

机构 * ECE Department(电子工程系) Carnegie Mellon University(卡内基梅隆大学) Computer Science Department(计算机科学系) Columbia University(哥伦比亚大学) Rotman School of Management(罗特曼管理学院) University of Toronto(多伦多大学) Department of Statistics(统计学系) George Washington University(乔治华盛顿大学) Department of Computer Science(计算机科学系) San Francisco State University(旧金山州立大学)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19947 2025-07-31 cs.RO cs.CL cs.IT cs.LG cs.SY eess.SY math.IT 79%

Spatial Language Likelihood Grounding Network for Bayesian Fusion of Human-Robot Observations

Supawich Sitdhipol, Waritwong Sukprasongdee, Ekapol Chuangsuwanich, Rina Tse

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.LG

Comments Accepted to the 2025 IEEE International Conference on Systems, Man, and Cybernetics (SMC); Supplementary video: https://cu-asl.github.io/fp-lgn/

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04047 2025-07-31 cs.CV 79%

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation

Ziyu Zhu, Xilin Wang, Yixuan Li, Zhuofan Zhang, Xiaojian Ma, Yixin Chen, Baoxiong Jia, Wei Liang, Qian Yu, Zhidong Deng, Siyuan Huang, Qing Li

机构 * Tsinghua University(清华大学) Beijing Institute of Technology(北京理工大学) Beihang University(北航) State Key Laboratory of General Artificial Intelligence, BIGAI, China(国家一般人工智能重点实验室, BIGAI, 中国)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Embodied AI; 3D Vision Language Understanding; ICCV 2025 Highlight; https://mtu3d.github.io; Spatial intelligence

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22802 2025-07-31 cs.CV cs.AI 62%

Advancing Fetal Ultrasound Image Quality Assessment in Low-Resource Settings

Dongli He, Hu Wang, Mohammad Yaqub

机构 * Mohamed bin Zayed University of Artificial Intelligence(莫莫德·宾·扎耶德人工智能大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.AI

Comments Accepted to the MICCAI 2025 MIRASOL Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.04858 2025-07-31 cs.AI cs.LG 62%

Don't Lag, RAG: Training-Free Adversarial Detection Using RAG

Roie Kazoom, Raz Lapid, Moshe Sipper, Ofer Hadar

机构 * Electrical and Computer Engineering, Ben Gurion University, Beer Sheba 84105, Israel(电子与计算机工程系,本· Gurion 大学) Computer Science, Ben Gurion University, Beer Sheba 84105, Israel(计算机科学系,本· Gurion 大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.AI、cs.LG

Comments Accepted at VecDB @ ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22617 2025-07-31 cs.CR cs.CV 57%

Hate in Plain Sight: On the Risks of Moderating AI-Generated Hateful Illusions

Yiting Qu, Ziqing Yang, Yihan Ma, Michael Backes, Savvas Zannettou, Yang Zhang

机构 * CISPA Helmholtz Center for Information Security(CISPA 欧洲信息安全研究中心) TU Delft(代尔夫特理工大学)

专题命中 视觉定位与Grounding :vision language model(abstract);分类 cs.CV

Comments Accepted at ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22213 2025-07-31 cs.IR cs.LG 57%

Intent-Aware Neural Query Reformulation for Behavior-Aligned Product Search

Jayanth Yetukuri, Ishita Khan

机构 * eBay Inc(eBay公司)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG

Comments Accepted at SIGIR eCom'25. https://sigir-ecom.github.io/eCom25Papers/paper_23.pdf

详情

展开后加载摘要…

URL PDF HTML 收藏