arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-08-07 至 2025-08-07 共收录 3 信号源:cs.CV, cs.AI, cs.LG

1. 视觉问答 3 篇

2508.04059 2025-08-07 cs.CV 83%

Beyond the Visible: Benchmarking Occlusion Perception in Multimodal Large Language Models

Zhaochen Liu, Kaiwen Gao, Shuyi Liang, Bin Xiao, Limeng Qiao, Lin Ma, Tingting Jiang

专题命中 视觉问答 :multimodal large language model(title,abstract);visual question answering(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04197 2025-08-07 cs.CV cs.AI 62%

Gather and Trace: Rethinking Video TextVQA from an Instance-oriented Perspective

Yan Zhang, Gangyan Zeng, Daiqing Wu, Huawen Shen, Binbin Li, Yu Zhou, Can Ma, Xiaojun Bi

机构 * Institute of Information Engineering, Chinese Academy of Sciences School of Cyber Security, University of Chinese Academy of Sciences Beijing China School of Cyber Science Engineering, Nanjing University of Science VCIP \& TMCC \& DISSec, College of Computer Science, Nankai University Tianjin China Key Laboratory of Ethnic Language Intelligent Analysis Security Governance of MOE, Minzu University of China Beijing China Institute of Information Engineering, Chinese Academy of Sciences School of Cyber Security, University of Chinese Academy of Sciences VCIP \& TMCC \& DISSec, College of Computer Science, Nankai University Security Governance of MOE, Minzu University of China

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.AI

Comments Accepted by 2025 ACM MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04271 2025-08-07 cs.DC 50%

S2M3: Split-and-Share Multi-Modal Models for Distributed Multi-Task Inference on the Edge

JinYi Yoon, JiHo Lee, Ting He, Nakjung Choi, Bo Ji

专题命中 视觉问答 :visual question answering(abstract)

Comments Accepted at IEEE International Conference on Distributed Computing Systems (ICDCS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏