arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-10-28 至 2025-10-28 共收录 16 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 16 篇

2510.23184 2025-10-28 cs.CV 88%

Finding 3D Scene Analogies with Multimodal Foundation Models

Junho Kim, Young Min Kim

机构 * Institute of New Media and Communications(新媒体与通讯研究所) Dept. of Electrical and Computer Engineering(电气与计算机工程系)

专题命中 图文多模态 :multimodal(title,abstract);multimodal foundation model(title,abstract);分类 cs.CV

Comments Accepted to FM4RoboPlan workshop at RSS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15510 2025-10-28 cs.CV cs.CL 84%

Visual Thoughts: A Unified Perspective of Understanding Multimodal Chain-of-Thought

Zihui Cheng, Qiguang Chen, Xiao Xu, Jiaqi Wang, Weiyun Wang, Hao Fei, Yidong Wang, Alex Jinpeng Wang, Zhi Chen, Wanxiang Che, Libo Qin

机构 * School of Computer Science and Engineering, Central South University(中南大学计算机科学与工程学院) Research Center for Social Computing and Interactive Robotics, Harbin Institute of Technology(哈尔滨工业大学社会计算与交互机器人研究中心) Institute of Computing and Intelligence, Harbin Institute of Technology, Shenzhen(哈尔滨工业大学深圳研究院计算与智能研究所) Text Computing and Cognitive Intelligence Ministry of Education Engineering Research Center, Guizhou University(贵州大学文字计算与认知智能教育部工程研究中心) Chinese University of Hong Kong(香港中文大学) Shanghai AI Laboratory(上海人工智能实验室) National University of Singapore(新加坡国立大学) Peking University(北京大学) ByteDance Seed (China)(字节跳动种子(中国))

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CV、cs.CL

Comments Accepted at NeurIPS 2025;

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12116 2025-10-28 cs.CL cs.AI cs.CV 82%

Unsupervised Document and Template Clustering using Multimodal Embeddings

Phillipe R. Sampaio, Helene Maxcici

机构 * BNP Paribas Cardif Nanterre(BNP巴黎银行卡迪夫纳特尔尔分校) BNP Paribas Paris(BNP巴黎银行巴黎)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments 24 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22838 2025-10-28 cs.CV 74%

Semantic-Preserving Cross-Style Visual Reasoning for Robust Multi-Modal Understanding in Large Vision-Language Models

Aya Nakayama, Brian Wong, Yuji Nishimura, Kaito Tanaka

专题命中 图文多模态 :multi-modal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.20625 2025-10-28 cs.CV 74%

T2ICount: Enhancing Cross-modal Understanding for Zero-Shot Counting

Yifei Qian, Zhongliang Guo, Bowen Deng, Chun Tong Lei, Shuai Zhao, Chun Pong Lau, Xiaopeng Hong, Michael P. Pound

机构 * University of Nottingham(诺丁汉大学) University of St Andrews(圣安德鲁大学) City University of Hong Kong(香港城市大学) Nanyang Technology University(南洋理工大学) Harbin Institute of Technology(哈尔滨工业大学)

专题命中 图文多模态 :cross-modal(title);分类 cs.CV

Comments Accepted by CVPR2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21879 2025-10-28 cs.CV cs.AI 73%

TernaryCLIP: Efficiently Compressing Vision-Language Models with Ternary Weights and Distilled Knowledge

Shu-Hao Zhang, Wei-Cheng Tang, Chen Wu, Peng Hu, Nan Li, Liang-Jie Zhang, Qi Zhang, Shao-Qun Zhang

机构 * State Key Laboratory of Novel Software Technology, Nanjing University(新型软件技术国家重点实验室) Microsoft AI(微软人工智能)

专题命中 图文多模态 :multimodal(abstract);image-text(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22785 2025-10-28 cs.CV 70%

Self-Calibrated Consistency can Fight Back for Adversarial Robustness in Vision-Language Models

Jiaxiang Liu, Jiawei Du, Xiao Liu, Prayag Tiwari, Mingkun Xu

机构 * Guangdong Institute of Intelligence Science and Technology(广东智能科学与技术研究院) Agency for Science, Technology and Research(科技研究局) School of Information Technology(信息技术学院)

专题命中 图文多模态 :cross-modal(abstract);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.07019 2025-10-28 cs.CV 70%

A Vision-Language Foundation Model for Leaf Disease Identification

Khang Nguyen Quoc, Lan Le Thi Thu, Luyl-Da Quach

专题命中 图文多模态 :multi-modal(abstract);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.11593 2025-10-28 cs.CV cs.AI cs.CL cs.DB cs.LG 67%

Improving Image Captioning Descriptiveness by Ranking and LLM-based Fusion

Luigi Celona, Simone Bianco, Marco Donzella, Paolo Napoletano

机构 * Department of Informatics, Systems and Communication(信息学、系统与通信系)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV、cs.CL、cs.AI

Comments This manuscript has been accepted for publication in Springer Neural Computing and Applications

Journal ref Neural Computer & Application 37, 27279-27299 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22370 2025-10-28 cs.RO cs.AI cs.CV cs.LG cs.SE 62%

BLIP-FusePPO: A Vision-Language Deep Reinforcement Learning Framework for Lane Keeping in Autonomous Vehicles

Seyed Ahmad Hosseini Miangoleh, Amin Jalal Aghdasian, Farzaneh Abdollahi

机构 * Department of Electrical Engineering, Amirkabir University of Technology (Tehran Polytechnic)(电气工程系,阿米尔卡比尔技术大学(德黑兰理工大学))

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments https://github.com/Amin-A96/BLIP-FusePPO-A-Vision-Language-Deep-Reinforcement-Learning-Framework-for-Lane-Keeping-in-Autonomous.git

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22045 2025-10-28 cs.CV cs.AI 62%

VLM-SlideEval: Evaluating VLMs on Structured Comprehension and Perturbation Sensitivity in PPT

Hyeonsu Kang, Emily Bao, Anjan Goswami

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 39th Conference on Neural Information Processing Systems (NeurIPS 2025) Workshop: Evaluating the Evolving LLM Lifecycle - Benchmarks, Emergent Abilities, and Scaling

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21806 2025-10-28 cs.CV cs.AI 62%

Frame-Difference Guided Dynamic Region Perception for CLIP Adaptation in Text-Video Retrieval

Jiaao Yu, Mingjie Han, Tao Gong, Jian Zhang, Man Lan

机构 * School of Computer Science and Technology, East China Normal University, China(上海师范大学计算机科学与技术学院) School of Information Science and Technology, University of Science and Technology of China(中国科学技术大学信息科学与技术学院)

专题命中 图文多模态 :cross-modal(abstract);分类 cs.CV、cs.AI

Comments 5 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19333 2025-10-28 cs.CV 57%

A Training-Free Framework for Open-Vocabulary Image Segmentation and Recognition with EfficientNet and CLIP

Ying Dai, Wei Yu Chen

专题命中 图文多模态 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05020 2025-10-28 cs.RO cs.AI 57%

Hierarchical Language Models for Semantic Navigation and Manipulation in an Aerial-Ground Robotic System

Haokun Liu, Zhaoqi Ma, Yunong Li, Junichiro Sugihara, Yicheng Chen, Jinjie Li, Moju Zhao

机构 * DRAGON Lab at Department of Mechanical Engineering, The University of Tokyo(东京大学机械工程系DRAGON实验室)

专题命中 图文多模态 :multimodal(abstract);分类 cs.AI

Comments 18 pages, 10 figures

Journal ref Advanced Intelligent Systems, Oct. 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.03610 2025-10-28 cs.CV 57%

Learning Knowledge-based Prompts for Robust 3D Mask Presentation Attack Detection

Fangling Jiang, Qi Li, Bing Liu, Weining Wang, Caifeng Shan, Zhenan Sun, Ming-Hsuan Yang

机构 * School of Computer Science, University of South China(南方大学计算机科学学院) New Laboratory of Pattern Recognition, MAIS, CASIA(模式识别新实验室,MAIS,CASIA) School of Intelligence Science and Technology, Nanjing University(智能科学与技术学院,南京大学) Department of Computer Science and Engineering, University of California, Merced(加州大学默塞德分校计算机科学与工程系) Department of Computer Science and Engineering, Yonsei University(延世大学计算机科学与工程系)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted by TPAMI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13317 2025-10-28 cs.LG 50%

Unlabeled Data vs. Pre-trained Knowledge: Rethinking SSL in the Era of Large Models

Song-Lin Lv, Rui Zhu, Tong Wei, Yu-Feng Li, Lan-Zhe Guo

机构 * School of Intelligence Science and Technology, Nanjing University, China(智能科学与技术学院,南京大学) School of Artificial Intelligence, Nanjing University, China(人工智能学院,南京大学) National Key Laboratory for Novel Software Technology, Nanjing University, China(新型软件技术国家重点实验室,南京大学) School of Computer Science and Engineering, Southeast University, Nanjing, China(计算机科学与工程学院,东南大学)

专题命中 图文多模态 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏