arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-08-19 至 2025-08-19 共收录 8 信号源:cs.CV, cs.AI, cs.LG

1. 其他VLM 8 篇

2508.12591 2025-08-19 cs.CL cs.AI cs.SD 83%

Beyond Modality Limitations: A Unified MLLM Approach to Automated Speaking Assessment with Effective Curriculum Learning

Yu-Hsuan Fang, Tien-Hong Lo, Yao-Ting Sung, Berlin Chen

机构 * National Taiwan Normal University(台湾国立正常大学)

专题命中 其他VLM :MLLM(title,abstract);multimodal large language model(abstract);分类 cs.AI

Comments Accepted at IEEE ASRU 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.10357 2025-08-19 cs.CV 79%

Optimization of Prompt Learning via Multi-Knowledge Representation for Vision-Language Models

Enming Zhang, Bingke Zhu, Yingying Chen, Qinghai Miao, Ming Tang, Jinqiao Wang

机构 * School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Foundation Model Research Center, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所基础模型研究中心) Wuhan AI Research(武汉人工智能研究所) Peng Cheng Laboratory(鹏城实验室)

专题命中 其他VLM :vision-language model(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19196 2025-08-19 cs.RO cs.CL cs.HC 78%

Towards Multimodal Social Conversations with Robots: Using Vision-Language Models

Ruben Janssens, Tony Belpaeme

机构 * Ghent University–imec(根特大学–imec)

专题命中 其他VLM :vision-language model(title,abstract)

Comments Accepted at the workshop "Human - Foundation Models Interaction: A Focus On Multimodal Information" (FoMo-HRI) at IEEE RO-MAN 2025 (Camera-ready version)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12854 2025-08-19 cs.AI cs.CL cs.CV cs.HC cs.MM 76%

E3RG: Building Explicit Emotion-driven Empathetic Response Generation System with Multimodal Large Language Model

Ronghao Lin, Shuai Shen, Weipeng Hu, Qiaolin He, Aolin Xiong, Li Huang, Haifeng Hu, Yap-peng Tan

机构 * Sun Yat-sen University(中山大学) Nanyang Technological University(南洋理工大学) Desay SV Automotive Co., Ltd(德赛西威汽车有限公司) Pazhou Laboratory(琶洲实验室)

专题命中 其他VLM :multimodal large language model(title);分类 cs.CV、cs.AI

Comments Accepted at ACM MM 2025 Grand Challenge

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12668 2025-08-19 cs.CV 57%

WP-CLIP: Leveraging CLIP to Predict Wölfflin's Principles in Visual Art

Abhijay Ghildyal, Li-Yun Wang, Feng Liu

机构 * Portland State University(波特兰州立大学)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

Comments ICCV 2025 AI4VA workshop (oral), Code: https://github.com/abhijay9/wpclip

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12396 2025-08-19 cs.CV 57%

DeCoT: Decomposing Complex Instructions for Enhanced Text-to-Image Generation with Large Language Models

Xiaochuan Lin, Xiangyong Chen, Xuan Li, Yichen Su

机构 * Henan Polytechnic University(河南理工大学)

专题命中 其他VLM :MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12291 2025-08-19 cs.AI 57%

RadarQA: Multi-modal Quality Analysis of Weather Radar Forecasts

Xuming He, Zhiyuan You, Junchao Gong, Couhua Liu, Xiaoyu Yue, Peiqin Zhuang, Wenlong Zhang, Lei Bai

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) ZheJiang University(浙江大学) The Chinese University of Hong Kong(香港中文大学) Center for Earth System Modeling and Prediction of China Meteorological Administration(中国气象局地球系统模拟与预测中心)

专题命中 其他VLM :MLLM(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11801 2025-08-19 cs.CV cs.CL 57%

VideoAVE: A Multi-Attribute Video-to-Text Attribute Value Extraction Dataset and Benchmark Models

Ming Cheng, Tong Wu, Jiazhen Hu, Jiaying Gong, Hoda Eldardiry

机构 * Virginia Tech(弗吉尼亚理工大学)

专题命中 其他VLM :vision language model(abstract);分类 cs.CV

Comments 5 pages, 2 figures, 5 tables, accepted in CIKM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏