arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-08-11 至 2025-08-11 共收录 5 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 5 篇

2508.06317 2025-08-11 cs.CV 83%

Uncertainty-quantified Rollout Policy Adaptation for Unlabelled Cross-domain Temporal Grounding

Jian Hu, Zixu Cheng, Shaogang Gong, Isabel Guan, Jianye Hao, Jun Wang, Kun Shao

机构 * Queen Mary University of London(伦敦女王学院) Hong Kong University of Science and Technology(香港科学与技术大学) Huawei Noah’s Ark Lab(华为诺亚实验室) University College London(伦敦大学学院)

专题命中 视觉定位与Grounding :grounding(title,abstract);vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05684 2025-08-11 cs.CR cs.LG 79%

MM-FusionNet: Context-Aware Dynamic Fusion for Multi-modal Fake News Detection with Large Vision-Language Models

Junhao He, Tianyu Liu, Jingyuan Zhao, Benjamin Turner

机构 * Huaiyin Institute of Technology(淮阴职业技术学院) Universidad Autónoma de Asunción(阿斯unción自治大学)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05976 2025-08-11 cs.CV cs.RO 77%

PASG: A Closed-Loop Framework for Automated Geometric Primitive Extraction and Semantic Anchoring in Robotic Manipulation

Zhihao Zhu, Yifan Zheng, Siyu Pan, Yaohui Jin, Yao Mu

机构 * MoE key Lab of Artificial Intelligence, AI Institute, Shanghai Jiao Tong University(人工智能教育部重点实验室、人工智能研究院、上海交通大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);grounding(abstract);分类 cs.CV

Comments Accepted to ICCV 2025. 8 pages main paper, 8 figures, plus supplementary material

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.24277 2025-08-11 cs.LG cs.AI 62%

Evaluating and Designing Sparse Autoencoders by Approximating Quasi-Orthogonality

Sewoong Lee, Adam Davies, Marc E. Canby, Julia Hockenmaier

机构 * Siebel School of Computing and Data Science University of Illinois Urbana-Champaign(计算与数据科学学院 耶鲁大学伊利诺伊分校)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05799 2025-08-11 cs.SE cs.AI cs.HC 57%

AI-Guided Exploration of Large-Scale Codebases

Yoseph Berhanu Alebachew

机构 * Department of Computer Science(计算机科学系) Virginia Tech(弗吉尼亚理工大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏