arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-08-08 至 2025-08-08 共收录 4 信号源:cs.CV, cs.AI, cs.LG

1. 视觉问答 4 篇

2506.21586 2025-08-08 cs.CL cs.AI cs.CV 81%

Can Vision Language Models Understand Mimed Actions?

Hyundong Cho, Spencer Lin, Tejas Srinivasan, Michael Saxon, Deuksin Kwon, Natali T. Chavez, Jonathan May

机构 * Information Sciences Institute(信息科学研究所) Institute for Creative Technologies(创意技术研究所) Department of Computer Science(计算机科学系) University of Southern California(南加州大学) University of California, Santa Barbara(加州大学圣巴巴拉分校) Aristotle University of Thessaloniki(希腊雅典纳大学)

专题命中 视觉问答 :vision language model(title);vision-language model(abstract);分类 cs.CV、cs.AI

Comments ACL 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.18351 2025-08-08 cs.CL cs.AI 79%

Multi-Agents Based on Large Language Models for Knowledge-based Visual Question Answering

Zhongjian Hu, Peng Yang, Bing Li, Zhenqi Wang

专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.AI

Comments We would like to withdraw this submission due to ongoing internal review and coordination among the author team. Upon the supervisor's recommendation, we have decided to delay public dissemination until the manuscript undergoes further refinement and aligns with our intended academic trajectory

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.16936 2025-08-08 cs.CL cs.AI 79%

Rationale-guided Prompting for Knowledge-based Visual Question Answering

Zhongjian Hu, Peng Yang, Bing Li, Fengyuan Liu

专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.AI

Comments We would like to withdraw this submission due to ongoing internal review and coordination among the author team. Upon the supervisor's recommendation, we have decided to delay public dissemination until the manuscript undergoes further refinement and aligns with our intended academic trajectory

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04895 2025-08-08 cs.SE 71%

Automated Bug Frame Retrieval from Gameplay Videos Using Vision-Language Models

Wentao Lu, Alexander Senchenko, Abram Hindle, Cor-Paul Bezemer

专题命中 视觉问答 :vision-language model(title)

详情

展开后加载摘要…

URL PDF HTML 收藏