arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-01-21 至 2026-01-21 共收录 12 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 12 篇

2601.04720 2026-01-21 cs.CL 83%

Qwen3-VL-Embedding and Qwen3-VL-Reranker: A Unified Framework for State-of-the-Art Multimodal Retrieval and Ranking

Qwen3-VL-Embedding 和 Qwen3-VL-Reranker:面向最新多模态检索与排序的统一框架

Mingxin Li, Yanzhao Zhang, Dingkun Long, Keqin Chen, Sibo Song, Shuai Bai, Zhibo Yang, Pengjun Xie, An Yang, Dayiheng Liu, Jingren Zhou, Junyang Lin

专题命中 跨模态检索 :multimodal(title,abstract);image-text(abstract);分类 cs.CL

AI总结 Qwen3-VL-Embedding和Qwen3-VL-Reranker通过统一框架实现多模态检索与排序的高精度性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05014 2026-01-21 cs.AI cs.LG 83%

Think Then Embed: Generative Context Improves Multimodal Embedding

思考后再嵌入:生成性上下文提升多模态嵌入

Xuanming Cui, Jianpeng Cheng, Hong-you Chen, Satya Narayan Shukla, Abhijeet Awasthi, Xichen Pan, Chaitanya Ahuja, Shlok Kumar Mishra, Yonghuan Yang, Jun Xiao, Qi Guo, Ser-Nam Lim, Aashu Singh, Xiangjun Fan

机构 * AI at Meta(Meta人工智能部门) University of Central Florida(中央佛罗里达大学) New York University(纽约大学)

专题命中 跨模态检索 :multimodal(title,abstract);MLLM(abstract);分类 cs.AI

AI总结 本研究提出Think-Then-Embed框架,通过生成性推理提升多模态嵌入性能,实现更高效的复杂指令处理。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07780 2026-01-21 cs.CV cs.AI 81%

Semantic-Consistent Bidirectional Contrastive Hashing for Noisy Multi-Label Cross-Modal Retrieval

语义一致的双向对比哈希用于噪声多标签跨模态检索

Likang Peng, Chao Su, Wenyuan Wu, Yuan Sun, Dezhong Peng, Xi Peng, Xu Wang

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV、cs.AI

AI总结 本文提出语义一致的双向对比哈希方法,通过跨模态语义一致性分类和双向软对比哈希模块,有效应对多标签数据中的噪声问题,提升跨模态检索的鲁棒性与性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11970 2026-01-21 cs.CV 79%

Real-Time Multi-Modal Embedded Vision Framework for Object Detection Facial Emotion Recognition and Biometric Identification on Low-Power Edge Platforms

面向低功耗边缘平台的实时多模态嵌入视觉框架:用于目标检测、面部情绪识别和生物识别

S. M. Khalid Bin Zahid, Md. Rakibul Hasan Nishat, Abdul Hasib, Md. Rakibul Hasan, Md. Ashiqussalehin, Md. Sahadat Hossen Sajib, A. S. M. Ahsanul Sarkar Akib

机构 * 1,2Department of Mechatronics Engineering, Rajshahi University of Engineering \& Technology 3 ,4,5Department of Internet of Things Robotics Engineering, University of Frontier Technology, Bangladesh 6Department of Computer Science Engineering, Varendra University, Rajshahi, Bangladesh 7Department of Robotics, Robo Tech Valley, Dhaka, Bangladesh Emails: , , 3

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CV

AI总结 本文提出了一种面向低功耗边缘平台的实时多模态视觉框架,通过自适应调度机制整合目标检测、面部识别和情绪检测,提升边缘设备上的智能感知效率与隐私保护。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12486 2026-01-21 cs.HC 78%

A Multimodal Assistive System for Product Localization and Retrieval for People who are Blind or have Low Vision

一种多模态辅助系统,用于帮助盲人或低视力者进行产品定位与检索

Ligao Ruan, Giles Hamilton-Fletcher, Mahya Beheshti, Todd E Hudson, Maurizio Porfiri, John-Ross Rizzo

专题命中 跨模态检索 :multimodal(title,abstract)

AI总结 本文提出一种多模态辅助系统,通过目标检测与视觉-语言模型结合,帮助盲人或低视力者自主完成产品定位与检索,提升其自主性和控制感。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14154 2026-01-21 cs.CV cs.AI 76%

LLM Augmented Intervenable Multimodal Adaptor for Post-operative Complication Prediction in Lung Cancer Surgery

基于大语言模型的可干预多模态适配器用于肺癌手术后并发症预测

Shubham Pandey, Bhavin Jawade, Srirangaraj Setlur, Venu Govindaraju, Kenneth Seastedt

机构 * University at Buffalo(布法罗大学) Roswell Park Comprehensive Cancer Center(罗斯威尔帕克综合癌症中心)

专题命中 跨模态检索 :multimodal(title);分类 cs.CV、cs.AI

AI总结 MIRACLE通过整合术前临床和放射学数据,利用超球面嵌入空间和干预式深度学习模块,实现肺癌手术后并发症风险的预测与可解释性管理。

Comments Accepted to P2P-CV @ WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14259 2026-01-21 cs.CV 70%

ManipShield: A Unified Framework for Image Manipulation Detection, Localization and Explanation

ManipShield: 一种用于图像篡改检测、定位和解释的统一框架

Zitong Xu, Huiyu Duan, Xiaoyu Wang, Zhaolin Cai, Kaiwei Zhang, Qiang Hu, Jing Liu, Xiongkuo Min, Guangtao Zhai

机构 * Institute of Image Communication and Network Engineering, Shanghai Jiao Tong University(上海交通大学图像通信与网络工程研究所) University of Electronic and Science Technology of China(电子科技大学) Tianjin University(天津大学)

专题命中 跨模态检索 :multimodal(abstract);MLLM(abstract);分类 cs.CV

AI总结 ManipShield基于多模态大语言模型,通过对比学习LoRA微调和任务特定解码器,实现图像篡改的统一检测、定位和解释。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14121 2026-01-21 cs.CL 57%

NewsRECON: News article REtrieval for image CONtextualization

NewsRECON: 用于图像上下文化的新闻文章检索

Jonathan Tonglet, Iryna Gurevych, Tinne Tuytelaars, Marie-Francine Moens

机构 * Ubiquitous Knowledge Processing Lab (UKP Lab), Department of Computer Science, TU Darmstadt and National Research Center for Applied Cybersecurity ATHENE(通用知识处理实验室(UKP实验室)、计算机科学系、图恩大学(TU Darmstadt)和应用网络安全国家研究中心ATHENE) Department of Electrical Engineering, KU Leuven(电气工程系、鲁汶大学) Department of Computer Science, KU Leuven(计算机科学系、鲁汶大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL

AI总结 NewsRECON通过链接图像与相关新闻文章,利用元数据推断图像日期和位置,优于现有方法并结合多模态大语言模型实现新的SOTA结果。

Comments Preprint under review. Code available at https://github.com/jtonglet/arxiv2025-newsrecon

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14060 2026-01-21 cs.CV 57%

Fine-Grained Zero-Shot Composed Image Retrieval with Complementary Visual-Semantic Integration

细粒度零样本组合图像检索与互补视觉-语义整合

Yongcong Ye, Kai Zhang, Yanghai Zhang, Enhong Chen, Longfei Li, Jun Zhou

机构 * State Key Laboratory of Cognitive Intelligence, University of Science and Technology of China(认知智能国家重点实验室,中国科学技术大学) Zhejiang University(浙江大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

AI总结 本文提出CVSI方法,通过互补视觉-语义整合提升细粒度零样本组合图像检索性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.13476 2026-01-21 cs.LG cs.AI 57%

A Unified Variational Imputation Framework for Electric Vehicle Charging Data Using Retrieval-Augmented Language Model

基于检索增强语言模型的统一变分填补框架用于电动汽车充电数据

Jinhao Li, Hao Wang

机构 * Department of Data Science and AI, Faculty of IT and Monash Energy Institute, Monash University(数据科学与人工智能系,信息科技学院和莫纳什能源研究所,莫纳什大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.AI

AI总结 本文提出PRAIM框架,利用检索增强语言模型和变分神经架构,提升电动汽车充电数据填补的准确性和统计分布保留能力,从而提高预测性能。

Comments 15 pages

Journal ref IEEE Transactions on Smart Grid, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23770 2026-01-21 cs.CV 57%

GenView++: Unifying Adaptive Generative Augmentation and Quality-Driven Supervision for Contrastive Representation Learning

GenView++:统一自适应生成增强和质量驱动监督以实现对比表征学习

Xiaojie Li, Bei Wang, Wei Liu, Jianlong Wu, Yue Yu, Liqiang Nie, Min Zhang

机构 * Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳)) Peng Cheng Laboratory(鹏城实验室)

专题命中 跨模态检索 :image-text(abstract);分类 cs.CV

AI总结 GenView++通过自适应生成增强和质量驱动监督,提升对比学习在视觉和视觉-语言任务中的性能。

Comments The code is available at \url{https://github.com/xiaojieli0903/GenViewPlusPlus}

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12010 2026-01-21 cs.CV 57%

SMc2f: Robust Scenario Mining for Robotic Autonomy from Coarse to Fine

SMc2f: 从粗到细的机器人自主性鲁棒场景挖掘

Yifei Chen, Ross Greer

机构 * Department of Computer Science at Xi’an University of Technology(西安理工大学计算机科学系) department of Computer Science & Engineering at the University of California, Merced(加州大学默塞德分校计算机科学与工程系)

专题命中 跨模态检索 :image-text(abstract);分类 cs.CV

AI总结 SMc2f通过从粗到细的流程,利用视觉语言模型和文本-轨迹对比学习提升机器人场景挖掘的鲁棒性和效率。

详情

展开后加载摘要…

URL PDF HTML 收藏