arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-10-01 至 2025-10-01 共收录 6 信号源:cs.CV, cs.AI, cs.LG

1. 视觉问答 6 篇

2509.25696 2025-10-01 cs.LG cs.CL eess.SP 83%

Can VLM Pseudo-Labels Train a Time-Series QA Model That Outperforms the VLM?

Takuya Fujimura, Kota Dohi, Natsuo Yamashita, Yohei Kawaguchi

机构 * Nagoya University(名古屋大学)

专题命中 视觉问答 :VLM(title,abstract);vision-language model(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.14446 2025-10-01 cs.CV cs.CY cs.LG 73%

Neglected Risks: The Disturbing Reality of Children's Images in Datasets and the Urgent Call for Accountability

Carlos Caetano, Gabriel O. dos Santos, Caio Petrucci, Artur Barros, Camila Laranjeira, Leo S. F. Ribeiro, Júlia F. de Mendonça, Jefersson A. dos Santos, Sandra Avila

机构 * School of Computer Science, University of Sheffield(谢菲尔德大学计算机科学学院)

专题命中 视觉问答 :vision-language model(abstract);visual question answering(abstract);分类 cs.CV、cs.LG

Comments ACM Conference on Fairness, Accountability, and Transparency (FAccT 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26235 2025-10-01 cs.CV 70%

Interpret, prune and distill Donut : towards lightweight VLMs for VQA on document

Adnan Ben Mansour, Ayoub Karine, David Naccache

机构 * DIENS, École Normale Supérieure of Paris, Paris, France(巴黎高等师范学院DIENS) Université Paris Cité, LIPADE, F-75006 Paris, France(巴黎大学) Be-ys Research, Paris, France(Be-ys研究所)

专题命中 视觉问答 :vision-language model(abstract);visual question answering(abstract);分类 cs.CV

Comments Accepted at Workshop on Machine Learning in Document Analysis and Recognition (ICDAR WML 2025), Wuhan, China

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25654 2025-10-01 cs.CV 70%

DescribeEarth: Describe Anything for Remote Sensing Images

Kaiyu Li, Zixuan Jiang, Xiangyong Cao, Jiayu Wang, Yuchen Xiao, Deyu Meng, Zhi Wang

机构 * School of Software Engineering, Xi’an Jiaotong University(西安交通大学软件工程学院) College of Artificial Intelligence, Xi’an Jiaotong University(西安交通大学人工智能学院) School of Computer Science and Technology and Ministry of Education Key Lab For Intelligent Networks and Network Security, Xi’an Jiaotong University(西安交通大学计算机科学与技术学院和教育部智能网络与网络安全重点实验室) School of Mathematics and Statistics and Ministry of Education Key Lab of Intelligent Networks and Network Security, Xi’an Jiaotong University(西安交通大学数学与统计学院和教育部智能网络与网络安全重点实验室) Pazhou Laboratory (Huangpu), Guangzhou, Guangdong, China(琶洲实验室(黄埔),广州,广东,中国)

专题命中 视觉问答 :vision-language model(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04943 2025-10-01 cs.CV cs.CL 70%

ReLoop: "Seeing Twice and Thinking Backwards" via Closed-loop Training to Mitigate Hallucinations in Multimodal understanding

Jianjiang Yang, Yanshu li, Ziyan Huang

机构 * University of Bristol(布里斯托大学) Brown University(布朗大学) South China University of Technology(华南理工大学)

专题命中 视觉问答 :visual question answering(abstract);multimodal large language model(abstract);分类 cs.CV

Comments Accepted by conference EMNLP2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07966 2025-10-01 cs.CV cs.AI cs.CL 62%

Scaling RL to Long Videos

Yukang Chen, Wei Huang, Baifeng Shi, Qinghao Hu, Hanrong Ye, Ligeng Zhu, Zhijian Liu, Pavlo Molchanov, Jan Kautz, Xiaojuan Qi, Sifei Liu, Hongxu Yin, Yao Lu, Song Han

机构 * NVIDIA MIT(麻省理工学院) HKU(香港大学) UC Berkeley(加州大学伯克利分校)

专题命中 视觉问答 :vision-language model(abstract);分类 cs.CV、cs.AI

Comments Accepted by NeurIPS 2025. Code at https://github.com/NVlabs/Long-RL and model at https://huggingface.co/Efficient-Large-Model/LongVILA-R1-7B

详情

展开后加载摘要…

URL PDF HTML 收藏