arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-02-17 至 2026-02-17 共收录 9 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 9 篇

2602.13758 2026-02-17 cs.CV cs.AI 86%

OmniScience: A Large-scale Multi-modal Dataset for Scientific Image Understanding

OmniScience: 一个大规模多模态数据集用于科学图像理解

Haoyi Tao, Chaozheng Huang, Nan Wang, Han Lyu, Linfeng Zhang, Guolin Ke, Xi Fang

机构 * DP Technology(DP技术)

专题命中 图文多模态 :multi-modal(title,abstract);multimodal(abstract);image-text(abstract);分类 cs.CV、cs.AI

AI总结 OmniScience是一个大规模多模态数据集,通过动态模型路由生成高信息密度的图像标题,提升多模态模型在科学图像理解上的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.14889 2026-02-17 cs.LG cs.CV cs.ET cs.HC cs.NE 79%

Web-Scale Multimodal Summarization using CLIP-Based Semantic Alignment

基于CLIP的语义对齐的网络级多模态摘要

Mounvik K, N Harshit

机构 * School of Computer Science Engineering(计算机科学与工程学院) VIT-AP University(VIT-AP大学)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出基于CLIP的语义对齐网络级多模态摘要框架,通过结合网络文本和图像数据生成摘要,实现高准确率的多模态对齐。

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.20110 2026-02-17 cs.CV 79%

Cross-Modal Mapping: Mitigating the Modality Gap for Few-Shot Image Classification

跨模态映射:缓解模态差距以实现少样本图像分类

Xi Yang, Pai Peng, Wulin Xie, Xiaohuan Lu, Jie Wen

专题命中 图文多模态 :cross-modal(title,abstract);分类 cs.CV

AI总结 本文提出跨模态映射方法,通过全局对齐和三元组损失优化,缓解模态差距,提升少样本图像分类性能。

Comments The authors request withdrawal of this article. This version was submitted in error. Compared to the intended final version, it contains inaccuracies and fails to accurately reflect the authors' work and conclusions

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.12916 2026-02-17 cs.CV cs.LG 77%

Reliable Thinking with Images

基于图像的可靠思考

Haobin Li, Yutong Yang, Yijie Lin, Xiang Dai, Mouxing Yang, Xi Peng

机构 * College of Computer Science, Sichuan University(四川大学计算机科学学院) Southwest China Institute of Electronic Technology(西南中国电子技术研究所) National Key Laboratory of Fundamental Algorithms(国家基础算法重点实验室) Models for Engineering Simulation, Sichuan University(工程模拟模型研究所,四川大学)

专题命中 图文多模态 :multimodal(abstract);multi-modal(abstract);image-text(abstract);分类 cs.CV

AI总结 本文提出RTWI方法,通过估计视觉提示和文本Co T的可靠性,以解决多模态推理中噪声思考问题,提升多模态大语言模型的性能。

Comments 26 pages, 19 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.14225 2026-02-17 cs.AI 70%

Text Before Vision: Staged Knowledge Injection Matters for Agentic RLVR in Ultra-High-Resolution Remote Sensing Understanding

文本优先于视觉:针对超高清遥感理解的代理强化学习在超高清遥感理解中的知识注入至关重要

Fengxiang Wang, Mingshuo Chen, Yueying Li, Yajie Yang, Yuhao Zhou, Di Wang, Yifan Zhang, Haoyu Wang, Haiyan Zhao, Hongda Sun, Long Lan, Jun Song, Yulin Wang, Jing Zhang, Wenlong Zhang, Bo Du

机构 * National University of Defense Technology, China(国防科技大学) Beijing University of Posts and Telecommunications, China(北京邮电大学) University of the Chinese Academy of Sciences, China(中国科学院大学) Sichuan University, China(四川大学) Wuhan University, China(武汉大学) Chinese Academy of Science, China(中国科学院) Tsinghua University, China(清华大学) Shanghai Artificial Intelligence Laboratory, China(上海人工智能实验室) Renmin University of China, China(中国人民大学)

专题命中 图文多模态 :multimodal(abstract);image-text(abstract);分类 cs.AI

AI总结 本文提出分阶段知识注入方法,利用文本引导提升超高清遥感理解的视觉推理能力,实现XLRS-Bench上的高准确率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13352 2026-02-17 cs.CV cs.AI cs.CL 67%

Using Deep Learning to Generate Semantically Correct Hindi Captions

使用深度学习生成语义正确的印地语描述

Wasim Akram Khan, Anil Kumar Vuppala

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 本研究利用深度学习生成印地语图像描述,采用VGG16和双向LSTM结合注意力机制,实现语义准确的描述生成。

Comments 34 pages, 12 figures, 3 tables. Master's thesis, Liverpool John Moores University, November 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.14425 2026-02-17 cs.CV 57%

Hierarchical Vision-Language Interaction for Facial Action Unit Detection

层次化视觉-语言交互用于面部动作单元检测

Yong Li, Yi Ren, Yizhe Zhang, Wenhua Zhang, Tianyi Zhang, Muyun Jiang, Guo-Sen Xie, Cuntai Guan

机构 * Key Laboratory of Child Development and Learning Science (Ministry of Education), School of Biological Sciences and Medical Engineering, Southeast University(儿童发展与学习科学重点实验室(教育部),生物科学与医学工程学院,东南大学) School of Computer Science and Engineering, Nanjing University of Science and Technology(计算机科学与工程学院,南京理工大学) School of Computer Science and Engineering, Nanyang Technological University(计算机科学与工程学院,南洋理工大学)

专题命中 图文多模态 :cross-modal(abstract);分类 cs.CV

AI总结 HiVA通过层次化视觉-语言交互方法,利用文本描述和多模态注意力机制提升面部动作单元检测的鲁棒性和语义丰富性。

Comments Accepted to IEEE Transaction on Affective Computing 2026

Journal ref IEEE Transaction on Affective Computing 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15704 2026-02-17 cs.CV 57%

Pyramid Token Pruning for High-Resolution Large Vision-Language Models via Region, Token, and Instruction-Guided Importance

通过区域、令牌和指令引导的重要性进行高分辨率大视觉-语言模型的金字塔令牌修剪

Yuxuan Liang, Xu Li, Xiaolei Chen, Yi Zheng, Haotian Chen, Bin Li, Xiangyang Xue

机构 * Shanghai Key Laboratory of Intelligent Information Processing(上海智能信息处理重点实验室) College of Computer Science and Artificial Intelligence(计算机科学与人工智能学院)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

AI总结 通过区域、令牌和指令引导的重要性进行高分辨率大视觉-语言模型的金字塔令牌修剪,有效降低计算和内存开销,同时保持性能

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.13157 2026-02-17 eess.SP 50%

Seeing Radio: From Zero RF Priors to Explainable Modulation Recognition with Vision Language Models

看见无线电:从零RF先验到可解释的调制识别与视觉语言模型

Hang Zou, Bohao Wang, Yu Tian, Lina Bariah, Chongwen Huang, Samson Lasaulce, Mérouane Debbah

专题命中 图文多模态 :multimodal(abstract)

AI总结 本文提出利用视觉语言模型直接感知无线电波信号并推断调制模式,通过转换IQ流为图像数据,提升模型准确性至90%,并实现可解释的调制识别。

详情

展开后加载摘要…

URL PDF HTML 收藏