Can VLM Pseudo-Labels Train a Time-Series QA Model That Outperforms the VLM?
机构 * Nagoya University(名古屋大学)
专题命中 视觉问答 :VLM(title,abstract);vision-language model(abstract);分类 cs.LG
AI 大模型
视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。
机构 * Nagoya University(名古屋大学)
专题命中 视觉问答 :VLM(title,abstract);vision-language model(abstract);分类 cs.LG
机构 * School of Computer Science, University of Sheffield(谢菲尔德大学计算机科学学院)
专题命中 视觉问答 :vision-language model(abstract);visual question answering(abstract);分类 cs.CV、cs.LG
Comments ACM Conference on Fairness, Accountability, and Transparency (FAccT 2025)
机构 * DIENS, École Normale Supérieure of Paris, Paris, France(巴黎高等师范学院DIENS) ; Université Paris Cité, LIPADE, F-75006 Paris, France(巴黎大学) ; Be-ys Research, Paris, France(Be-ys研究所)
专题命中 视觉问答 :vision-language model(abstract);visual question answering(abstract);分类 cs.CV
Comments Accepted at Workshop on Machine Learning in Document Analysis and Recognition (ICDAR WML 2025), Wuhan, China
机构 * School of Software Engineering, Xi’an Jiaotong University(西安交通大学软件工程学院) ; College of Artificial Intelligence, Xi’an Jiaotong University(西安交通大学人工智能学院) ; School of Computer Science and Technology and Ministry of Education Key Lab For Intelligent Networks and Network Security, Xi’an Jiaotong University(西安交通大学计算机科学与技术学院和教育部智能网络与网络安全重点实验室) ; School of Mathematics and Statistics and Ministry of Education Key Lab of Intelligent Networks and Network Security, Xi’an Jiaotong University(西安交通大学数学与统计学院和教育部智能网络与网络安全重点实验室) ; Pazhou Laboratory (Huangpu), Guangzhou, Guangdong, China(琶洲实验室(黄埔),广州,广东,中国)
专题命中 视觉问答 :vision-language model(abstract);MLLM(abstract);分类 cs.CV
机构 * University of Bristol(布里斯托大学) ; Brown University(布朗大学) ; South China University of Technology(华南理工大学)
专题命中 视觉问答 :visual question answering(abstract);multimodal large language model(abstract);分类 cs.CV
Comments Accepted by conference EMNLP2025
机构 * NVIDIA ; MIT(麻省理工学院) ; HKU(香港大学) ; UC Berkeley(加州大学伯克利分校)
专题命中 视觉问答 :vision-language model(abstract);分类 cs.CV、cs.AI
Comments Accepted by NeurIPS 2025. Code at https://github.com/NVlabs/Long-RL and model at https://huggingface.co/Efficient-Large-Model/LongVILA-R1-7B