arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-08-27 至 2025-08-27 共收录 2 信号源:cs.CV, cs.AI, cs.LG

1. 其他VLM 2 篇

2506.09284 2025-08-27 cs.RO cs.AI cs.CV 62%

UAD: Unsupervised Affordance Distillation for Generalization in Robotic Manipulation

Yihe Tang, Wenlong Huang, Yingke Wang, Chengshu Li, Roy Yuan, Ruohan Zhang, Jiajun Wu, Li Fei-Fei

机构 * Stanford University(斯坦福大学)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19024 2025-08-27 cs.CV 57%

ProPy: Building Interactive Prompt Pyramids upon CLIP for Partially Relevant Video Retrieval

Yi Pan, Yujia Zhang, Michael Kampffmeyer, Xiaoguang Zhao

机构 * State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences(多模态人工智能系统国家重点实验室,自动化研究所,中国科学院) School of Artificial Intelligence, University of Chinese Academy of Sciences(人工智能学院,中国科学院大学) Department of Physics and Technology, UiT The Arctic University of Norway(物理与技术系,UiT 北极大学)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

Comments Accepted by EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏