arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-11-05 至 2025-11-05 共收录 7 信号源:cs.CV, cs.AI, cs.LG

1. 视觉问答 7 篇

2511.02182 2025-11-05 cs.CV 86%

Pinpointing Trigger Moment for Grounded Video QA: Enhancing Spatio-temporal Grounding in Multimodal Large Language Models

Jinhwan Seo, Yoonki Cho, Junhyug Noh, Sung-eui Yoon

机构 * KAIST(韩国科学技术院) Ewha Womans University(成均馆大学)

专题命中 视觉问答 :grounding(title,abstract);multimodal large language model(title);分类 cs.CV

Comments 1st place winner of Grounded Videoqa track at the ICCV2025 Perception Test

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01914 2025-11-05 cs.CV cs.AI cs.RO 73%

iFlyBot-VLA Technical Report

Yuan Zhang, Chenyu Xue, Wenjie Xu, Chao Ji, Jiajia wu, Jia Pan

机构 * iFlyTek Reasearch and Development Group(iFlyTek 研发部)

专题命中 视觉问答 :vision-language model(abstract);VLM(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02046 2025-11-05 cs.CV cs.AI 62%

Text-VQA Aug: Pipelined Harnessing of Large Multimodal Models for Automated Synthesis

Soham Joshi, Shwet Kamal Mishra, Viswanath Gopalakrishnan

机构 * International Institute of Information Technology Bangalore(国际信息科技学院班加罗尔)

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.AI

Comments First two authors contributed equally

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23495 2025-11-05 cs.CL cs.AI cs.LG 62%

Diagnosing and Addressing Pitfalls in KG-RAG Datasets: Toward More Reliable Benchmarking

Liangliang Zhang, Zhuorui Jiang, Hongliang Chi, Haoyang Chen, Mohammed Elkoumy, Fali Wang, Qiong Wu, Zhengyi Zhou, Shirui Pan, Suhang Wang, Yao Ma

机构 * Rensselaer Polytechnic Institute(罗切斯特理工学院) University of Toronto(多伦多大学) Pennsylvania State University(宾夕法尼亚州立大学) AT&T Chief Data Office(AT&T首席数据办公室) Griffith University(格里菲斯大学)

专题命中 视觉问答 :grounding(abstract);分类 cs.AI、cs.LG

Comments Accepted at NeurIPS 2025 Datasets and Benchmarks Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.12926 2025-11-05 eess.SP cs.CV 57%

Task-Oriented Feature Compression for Multimodal Understanding via Device-Edge Co-Inference

Cheng Yuan, Zhening Liu, Jiashu Lv, Jiawei Shao, Yufei Jiang, Jun Zhang, Xuelong Li

机构 * Institute of Artificial Intelligence (TeleAI) of China Telecom(中国电信人工智能研究所) School of Electronic and Information Engineering, Harbin Institute of Technology(哈尔滨工业大学电子与信息工程学院) Department of Electronic and Computer Engineering, Hong Kong University of Science and Technology(香港科技大学电子与计算机工程系) School of Software and Microelectronics, Peking University(北京大学软件与微电子学院)

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

Comments Accepted by IEEE Transactions on Mobile Computing

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.08564 2025-11-05 cs.CV cs.CL 57%

Visual Program Distillation with Template-Based Augmentation

Michal Shlapentokh-Rothman, Yu-Xiong Wang, Derek Hoiem

机构 * University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

Comments EMNLP Camera Ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08039 2025-11-05 cs.SD cs.CL cs.MM eess.AS 50%

Audio-Thinker: Guiding Audio Language Model When and How to Think via Reinforcement Learning

Shu Wu, Chenxing Li, Wenfu Wang, Hao Zhang, Hualei Wang, Meng Yu, Dong Yu

专题命中 视觉问答 :multimodal large language model(abstract)

Comments preprint

详情

展开后加载摘要…

URL PDF HTML 收藏