arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-10-08 至 2025-10-08 共收录 7 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7 篇

2510.05153 2025-10-08 cs.AI cs.IT math.IT 79%

An Algorithmic Information-Theoretic Perspective on the Symbol Grounding Problem

Zhangchi Liu

机构 * Zhangchi Liu(刘志强)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

Comments 7 pages, 1 table (in appendix)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06040 2025-10-08 cs.CV cs.AI 76%

VideoMiner: Iteratively Grounding Key Frames of Hour-Long Videos via Tree-based Group Relative Policy Optimization

Xinye Cao, Hongcan Guo, Jiawen Qian, Guoshun Nan, Chao Wang, Yuqi Pan, Tianhao Hou, Xiaojuan Wang, Yutong Gao

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Minzu University of China(民族大学)

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV、cs.AI

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06056 2025-10-08 cs.AI 57%

Scientific Algorithm Discovery by Augmenting AlphaEvolve with Deep Research

Gang Liu, Yihan Zhu, Jie Chen, Meng Jiang

机构 * University of Notre Dame(诺丁汉大学) MIT-IBM Watson AI Lab, IBM Research(麻省理工-IBM Watson AI实验室,IBM研究)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments 25 pages, 17 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22304 2025-10-08 cs.CV 57%

CADReview: Automatically Reviewing CAD Programs with Error Detection and Correction

Jiali Chen, Xusen Hei, HongFei Liu, Yuancheng Wei, Zikun Deng, Jiayuan Xie, Yi Cai, Li Qing

机构 * School of Software Engineering, South China University of Technology(华南理工大学软件学院) Key Laboratory of Big Data and Intelligent Robot Ministry of Education(教育部大数据与智能机器人重点实验室) The Hong Kong Polytechnic University(香港理工大学)

专题命中 视觉定位与Grounding :multimodal large language model(abstract);分类 cs.CV

Comments ACL 2025 main conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05722 2025-10-08 cs.CV 57%

Data Factory with Minimal Human Effort Using VLMs

Jiaojiao Ye, Jiaxing Zhong, Qian Xie, Yuzhou Zhou, Niki Trigoni, Andrew Markham

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments Tech report

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22688 2025-10-08 cs.CV 57%

Robust Object Detection for Autonomous Driving via Curriculum-Guided Group Relative Policy Optimization

Xu Jia

专题命中 视觉定位与Grounding :multimodal large language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05619 2025-10-08 eess.AS 50%

Teaching Machines to Speak Using Articulatory Control

Akshay Anand, Chenxu Guo, Cheol Jun Cho, Jiachen Lian, Gopala Anumanchipalli

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏