arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-09-26 至 2025-09-26 共收录 4 信号源:cs.CV, cs.AI, cs.LG

1. 视觉问答 4 篇

2509.20871 2025-09-26 cs.CV cs.AI 84%

SCRA-VQA: Summarized Caption-Rerank for Augmented Large Language Models in Visual Question Answering

Yan Zhang, Jiaqing Lin, Miao Zhang, Kui Xiao, Xiaoju Hou, Yue Zhao, Zhifei Li

机构 * School of Computer Science, Hubei University, Wuhan, China(湖北大学计算机学院) Hubei Key Laboratory of Big Data Intelligent Analysis and Application (Hubei University)(湖北大数据智能分析与应用重点实验室) Key Laboratory of Intelligent Sensing System and Security (Hubei University)(智能传感系统与安全重点实验室) Institute of Vocational Education, Guangdong Industry Polytechnic University(广东行业职业大学职业教育学院) Shandong Police College(山东警察学院)

专题命中 视觉问答 :visual question answering(title,abstract);visual language model(abstract);分类 cs.CV、cs.AI

Comments ACCEPTED as a FULL PAPER for the Research Track at International Conference on Database Systems for Advanced Applications 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20884 2025-09-26 cs.CV cs.AI 81%

Integrating Object Interaction Self-Attention and GAN-Based Debiasing for Visual Question Answering

Zhifei Li, Feng Qiu, Yiran Wang, Yujing Xia, Kui Xiao, Miao Zhang, Yan Zhang

机构 * School of Computer Science, Hubei University(湖北大学计算机学院)

专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV、cs.AI

Comments 14 pages, 6 figures. ACCEPTED for publication as a REGULAR paper in the IEEE Transactions on Multimedia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19973 2025-09-26 cs.CV 70%

OmniScene: Attention-Augmented Multimodal 4D Scene Understanding for Autonomous Driving

Pei Liu, Hongliang Lu, Haichao Liu, Haipeng Liu, Xin Liu, Ruoyu Yao, Shengbo Eben Li, Jun Ma

机构 * The Hong Kong University of Science and Technology(香港科技大学) Li Auto Inc. the School of Vehicle and Mobility, Tsinghua University(清华大学车辆与移动系统学院)

专题命中 视觉问答 :vision-language model(abstract);visual question answering(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20858 2025-09-26 cs.GR cs.CV cs.MM 57%

ArchGPT: Understanding the World's Architectures with Large Multimodal Models

Yuze Wang, Luo Yang, Junyi Wang, Yue Qi

机构 * State Key Laboratory of Virtual Reality Technology and Systems(虚拟现实技术与系统国家重点实验室) School of Computer Science and Engineering(计算机科学与工程学院) Beihang University(北京航空航天大学) School of Computer Science and Technology(计算机科学与技术学院) Shandong University(山东大学)

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏