arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-03-03 至 2026-03-03 共收录 206 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 30 篇

2410.16953 2026-03-03 cs.CV 57%

Towards Real Zero-Shot Camouflaged Object Segmentation without Camouflaged Annotations

面向无遮蔽标注的零样本遮蔽物分割

Cheng Lei, Jie Fan, Xinran Li, Tianzhu Xiang, Ao Li, Ce Zhu, Le Zhang

机构 * University of Electronic Science and Technology of China(电子科学与技术大学) Space42

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

AI总结 本文提出了一种无需遮蔽标注的零样本遮蔽物分割框架,通过结合MIM、M-LLM和MFA机制,实现高效分割与快速推理。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00793 2026-03-03 cs.CV 57%

Neural Functional Alignment Space: Brain-Referenced Representation of Artificial Neural Networks

神经功能对齐空间:人工神经网络的脑参考表示

Ruiyu Yan, Hanqi Jiang, Yi Pan, Xiaobo Li, Tianming Liu, Xi Jiang, Lin Zhao

机构 * Tandon School of Engineering, New York University(纽约大学工程学院) School of Computing, University of Georgia(佐治亚大学计算机学院) Department of Biomedical Engineering, New Jersey Institute of Technology(新泽西理工学院生物医学工程系) The Clinical Hospital of Chengdu Brain Science Institute, MOE Key Laboratory for NeuroInformation,School of Life Science and Technology, University of Electronic Science and Technology of China(成都脑科学研究所临床医院,国家神经信息重点实验室,电子科技大学生命科学与技术学院)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

AI总结 本文提出神经功能对齐空间,通过建模神经网络的动态表示,揭示了在脑参考空间中的结构化组织,包括模态特定聚类和跨模态收敛。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00609 2026-03-03 cs.CV 57%

Linking Modality Isolation in Heterogeneous Collaborative Perception

异构协作感知中的模态隔离问题

Changxing Liu, Zichen Chao, Siheng Chen

机构 * Shanghai Jiao Tong University(上海交通大学) Nanjing University of Science and Technology(南京理工大学)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

AI总结 CodeAlign通过跨模态特征-代码-特征翻译有效解决异构协作感知中的模态隔离问题,显著降低训练参数和通信负载,提升感知性能。

Comments Accepted by CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.19534 2026-03-03 cs.RO cs.AI 57%

Large Language Model-Assisted UAV Operations and Communications: A Multifaceted Survey and Tutorial

大型语言模型辅助的无人机操作与通信:多方面的综述与教程

Yousef Emami, Hao Zhou, Radha Reddy, Atefeh Hajijamali Arani, Biliang Wang, Kai Li, Luis Almeida, Zhu Han

机构 * IEEE

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

AI总结 本文综述了大型语言模型在无人机操作与通信中的应用,探讨了LLMs在提升UAV智能方面的多方面技术与未来研究方向。

Comments 40 pages, 10 figures, 13 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.15663 2026-03-03 cs.CV 57%

MSSPlace: Multi-Sensor Place Recognition with Visual and Text Semantics

MSSPlace: 多传感器位置识别与视觉和文本语义

Alexander Melekhin, Dmitry Yudin, Ilia Petryashin, Vitaly Bezuglyj

机构 * Intelligent Transport Laboratory, Moscow Institute of Physics and Technology(智能交通实验室,莫斯科物理技术学院) Artificial Intelligence Research Institute (AIRI)(人工智能研究机构(AIRI))

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

AI总结 MSSPlace通过整合多传感器数据和视觉文本语义,提升位置识别性能,达到最先进的效果。

Comments This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00156 2026-03-03 cs.CV 57%

BiCLIP: Bidirectional and Consistent Language-Image Processing for Robust Medical Image Segmentation

BiCLIP: 用于鲁棒医学图像分割的双向和一致语言-图像处理

Saivan Talaei, Fatemeh Daneshfar, Abdulhady Abas Abdullah, Mustaqeem Khan

机构 * Department of Computer Engineering, University of Kurdistan, Iran(伊朗库尔德大学计算机工程系) Artificial Intelligence and Innovation Centre, University of Kurdistan, Erbil, Iraq(伊拉克埃尔比尔库尔德大学人工智能与创新中心) College of Information Technology, United Arab Emirates University, UAE(阿联酋大学信息科技学院)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

AI总结 BiCLIP通过双向多模态融合和一致性目标提升医学图像分割的鲁棒性,有效应对标注稀少和临床伪影挑战。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21028 2026-03-03 cs.LG 50%

TRIDENT: Tri-Modal Molecular Representation Learning with Taxonomic Annotations and Local Correspondence

TRIDENT:结合分类注释和局部对应关系的三模态分子表示学习

Feng Jiang, Mangal Prakash, Hehuan Ma, Jianyuan Deng, Yuzhi Guo, Amina Mollaysa, Tommaso Mansi, Rui Liao, Junzhou Huang

机构 * University of Texas at Arlington(德克萨斯大学阿灵顿分校) Johnson & Johnson Innovative Medicine(强生创新医学)

专题命中 多模态训练与对齐 :multimodal(abstract)

AI总结 TRIDENT通过整合SMILES、文本和分类功能注释,学习丰富的分子表示,从而在多个下游任务中取得最佳性能。

Comments Accepted to NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 其他多模态 19 篇

2509.03113 2026-03-03 cs.CV cs.CL 86%

Mitigating Multimodal Hallucinations via Gradient-based Self-Reflection

通过基于梯度的自我反思缓解多模态幻觉

Shan Wang, Maying Shen, Nadine Chang, Chuong Nguyen, Hongdong Li, Jose M. Alvarez

机构 * NVIDIA Australian National University(澳大利亚国立大学) Data61, CSIRO(Data61,CSIRO)

专题命中 其他多模态 :multimodal(title,abstract);MLLM(abstract);cross-modal(abstract);分类 cs.CV、cs.CL

AI总结 本文提出GACD方法,通过梯度分析缓解多模态模型的幻觉问题,提升输出的视觉基础性。

Comments CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22171 2026-03-03 cs.HC 82%

A Taxonomy of Human--MLLM Interaction in Early-Stage Sketch-Based Design Ideation

早期阶段基于草图的设计构想中人类与大语言模型交互的分类

Weiyan Shi, Kenny Tsu Wei Choo

专题命中 其他多模态 :MLLM(title,abstract);multimodal(abstract)

AI总结 本文提出了一种分类方法,用于描述人类与大语言模型在早期阶段基于草图的设计构想中的交互模式,揭示了人类与AI角色的动态变化。

Comments Accepted at CHI 2026 Posters

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.12880 2026-03-03 cs.AI cs.MM 81%

Has Multimodal Learning Delivered Universal Intelligence in Healthcare? A Comprehensive Survey

多模态学习是否在医疗领域实现了通用智能?一项全面的综述

Qika Lin, Yifan Zhu, Xin Mei, Ling Huang, Jingying Ma, Kai He, Zhen Peng, Erik Cambria, Mengling Feng

机构 * Saw Swee Hock School of Public Health, National University of Singapore(新加坡国立大学公共健康学院) School of Computer Science, Beijing University of Posts and Telecommunications(北京邮电大学计算机学院) School of Automation, Northwestern Polytechnical University(西北工业大学自动化学院) School of Computer Science and Technology, Xi’an Jiaotong University(西安交通大学计算机科学与技术学院)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.AI、cs.MM

AI总结 本文通过全面调查,指出当前多模态学习在医疗领域尚未实现通用智能,并提出十个潜在研究方向。

Comments 21 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01106 2026-03-03 cs.AI 79%

DIVA-GRPO: Enhancing Multimodal Reasoning through Difficulty-Adaptive Variant Advantage

DIVA-GRPO:通过难度自适应变体优势增强多模态推理

Haowen Gao, Zhenyu Zhang, Liang Pang, Fangda Guo, Hongjian Dou, Guannan Lv, Shaoguo Liu, Tingting Gao, Huawei Shen, Xueqi Cheng

机构 * State Key Laboratory of AI Safety, Institute of Computing Technology, CAS, Beijing, China(人工智能安全国家重点实验室,计算技术研究所,中国科学院,北京,中国) University of Chinese Academy of Sciences, Beijing, China(中国科学院大学,北京,中国) Kuaishou Technology, Beijing, China(快手科技,北京,中国)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.AI

AI总结 DIVA-GRPO通过难度自适应变体优势方法提升多模态推理能力,解决GRPO在困难问题上的奖励稀疏性和优势消失问题,提升训练稳定性与推理性能。

Comments Accepted to ICLR 2026. Code and models are available at https://github.com/Siaaaaaa1/DIVA-GRPO

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00289 2026-03-03 cs.CV 79%

Seeking Necessary and Sufficient Information from Multimodal Medical Data

从多模态医学数据中寻求必要和充分的信息

Boyu Chen, Weiye Bao, Junjie Liu, Michael Shen, Bo Peng, Paul Taylor, Zhu Li, Mengyue Yang

机构 * University College London, London, UK(伦敦大学学院) Imperial College London, London, UK(伦敦帝国学院) Mingdu Tech, China(明都科技) University of Bristol, Bristol, UK(布里斯托大学)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出通过概率必要性和充分性学习多模态医学数据中的必要和充分特征,以提升模型性能和鲁棒性。

Comments 11 pages, 1 figure. Submitted to MICCAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.27492 2026-03-03 cs.CV 79%

ThinkMorph: Emergent Properties in Multimodal Interleaved Chain-of-Thought Reasoning

ThinkMorph:多模态交错链式推理中的涌现特性

Jiawei Gu, Yunzhuo Hao, Huichen Will Wang, Linjie Li, Michael Qizhe Shieh, Yejin Choi, Ranjay Krishna, Yu Cheng

机构 * National University of Singapore(新加坡国立大学) Zhejiang University(浙江大学) University of Washington(华盛顿大学) Stanford University(斯坦福大学) absolute AI The Chinese University of Hong Kong(香港中文大学)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 ThinkMorph通过统一模型提升多模态推理性能,展现视觉操控与模式切换等新兴能力。

Comments project page: https://thinkmorph.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.02339 2026-03-03 cs.CL 79%

AStar: Boosting Multimodal Reasoning with Automated Structured Thinking

AStar: 通过自动化结构化思维提升多模态推理

Jinyang Wu, Mingkuan Feng, Guocheng Zhai, Shuai Zhang, Zheng Lian, Fangrui Lv, Pengpeng Shao, Ruihan Jin, Zhengqi Wen, Jianhua Tao

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CL

AI总结 AStar通过自动化结构化思维提升多模态推理效率,实现更高准确率和更强迁移能力。

Comments Accepted by AAAI 2026 Oral

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02109 2026-03-03 eess.SP cs.LG 78%

Orchestrating Multimodal DNN Workloads in Wireless Neural Processing

在无线神经处理中协调多模态DNN工作负载

Sai Xu, Kai-Kit Wong, Yanan Du, Hyundong Shin

机构 * Department of Electronic and Electrical Engineering, University College London(电子与电气工程系,伦敦大学学院) Department of Electronic Engineering, Kyung Hee University(电子工程系,庆熙大学) School of Electrical and Electronic Engineering, the University of Sheffield(电气与电子工程学院,谢菲尔德大学) Department of Electronics and Information Convergence Engineering, Kyung Hee University(电子与信息融合工程系,庆熙大学)

专题命中 其他多模态 :multimodal(title,abstract)

AI总结 本文提出O-WiN框架和PACS算法,通过通信-计算流水线优化无线神经处理中多模态DNN工作负载,提升执行效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00565 2026-03-03 cs.CV cs.AI cs.CR 62%

MIDAS: Multi-Image Dispersion and Semantic Reconstruction for Jailbreaking MLLMs

MIDAS: 多图像分散与语义重建用于对抗多模态大语言模型

Yilian Liu, Xiaojun Jia, Guoshun Nan, Jiuyang Lyu, Zhican Chen, Tao Guan, Shuyuan Luo, Zhongyi Zhai, Yang Liu

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Nanyang Technological University(南洋理工大学) Guilin University of Electronic Technology(桂林电子科技大学)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 MIDAS通过多图像分散与语义重建技术,提升对抗多模态大语言模型的劫持性能,达到81.46%的平均攻击成功率。

Journal ref The Fourteenth International Conference on Learning Representations(2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01948 2026-03-03 cs.CV 57%

PreSight: Preoperative Outcome Prediction for Parkinson's Disease via Region-Prior Morphometry and Patient-Specific Weighting

PreSight:通过区域先验形态学和患者特异性加权进行帕金森病术前预后预测

Yand Wang, Chen Zhang, Lanyun Zhu, Yixin Chen, Qunbo Wang, Yutong Bai, Jurgen Germann, Yinghong Wen, Shuai Shao

机构 * Beijing Jiaotong University(北京交通大学) Nanyang Technological University(南洋理工大学) Institute of Medical Technology, Peking University(北京大学医学技术研究院) Beijing Tiantan Hospital, Capital Medical University(北京天坛医院) University Health Network, University of Toronto(多伦多大学健康网络) Suzhou Institute for Advanced Research, University of Science and Technology of China(中国科学技术大学苏州研究院)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV

AI总结 PreSight通过结合临床先验与区域自适应形态学,实现帕金森病术前预后预测,提升术后运动改善预测的准确性和临床实用性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.23348 2026-03-03 cs.RO cs.CV 57%

Physically Ground Commonsense Knowledge for Articulated Object Manipulation with Analytic Concepts

为拟合物体操纵的物理常识知识而引入分析概念

Jiude Wei, Yuxuan Li, Cewu Lu, Jianhua Sun

机构 * School of Artificial Intelligence, Shanghai Jiao Tong University(上海交通大学人工智能学院) School of Computer Science, Shanghai Jiao Tong University(上海交通大学计算机学院) Shanghai Innovation Institute(上海创新研究院)

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV

AI总结 本文提出通过引入分析概念,将语义级常识知识接地到物理世界,以提升机器人对关节物体的通用精确操纵能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.04490 2026-03-03 cs.CL q-bio.GN 57%

Large Language Models in Bioinformatics: A Survey

大语言模型在生物信息学中的应用:综述

Zhenyu Wang, Zikang Wang, Jiyue Jiang, Pengan Chen, Xiangyu Shi, Yu Li

机构 * The Chinese University of Hong Kong(香港中文大学) Peking University Third Hospital(北京大学第三医院) The Hong Kong Polytechnic University(香港理工大学) The University of Hong Kong(香港大学)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CL

AI总结 本文综述了大语言模型在生物信息学中的应用,涵盖基因组序列建模、RNA结构预测等核心方法,并探讨了数据稀缺和跨组学整合等挑战及未来发展方向。

Comments Accepted by ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01026 2026-03-03 cs.CV 57%

RaUF: Learning the Spatial Uncertainty Field of Radar

RaUF: 学习雷达的时空不确定性场

Shengpeng Wang, Kuangyu Wang, Wei Wang

机构 * Huazhong University of Science and Technology(华中科技大学) Wuhan University(武汉大学)

专题命中 其他多模态 :cross-modal(abstract);分类 cs.CV

AI总结 RaUF通过学习雷达测量的各向异性特性,解决方位模糊和虚假回波问题,提升空间检测的可靠性与不确定性校准。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00988 2026-03-03 cs.CV cs.SE 57%

Foundation Models in Remote Sensing: Evolving from Unimodality to Multimodality

遥感中的基础模型:从单模态到多模态的演变

Danfeng Hong, Chenyu Li, Xuyang Li, Gustau Camps-Valls, Jocelyn Chanussot

机构 * School of Automation, Southeast University(自动化学院,东南大学) School of Mathematics, Southeast University(数学学院,东南大学) Aerospace Information Research Institute, Chinese Academy of Sciences(航天信息研究所,中国科学院) Image Processing Laboratory (IPL), Universitat de València(图像处理实验室(IPL),瓦伦西亚大学) Univ. Grenoble Alpes, INRIA, CNRS, Grenoble INP, LJK(格勒诺布尔阿尔卑斯大学,INRIA,CNRS,格勒诺布尔INP,LJK)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV

AI总结 本文探讨了遥感中基础模型从单模态到多模态的演变,旨在为研究人员提供深入理解与应用指导。

Comments Accepted by IEEE GRSM

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00436 2026-03-03 cs.LG cs.AI 57%

ROKA: Robust Knowledge Unlearning against Adversaries

ROKA: 面对对抗者的鲁棒知识反学习

Jinmyeong Shin, Joshua Tapia, Nicholas Ferreira, Gabriel Diaz, Moayed Daneshyari, Hyeran Jeon

机构 * University of California, Merced(加州大学梅尔塞德斯分校) California State University, East Bay(加州州立大学东湾分校)

专题命中 其他多模态 :multi-modal(abstract);分类 cs.AI

AI总结 ROKA通过神经愈合机制实现鲁棒的知识反学习,有效对抗间接反学习攻击,同时保护保留数据的准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22753 2026-03-03 math.OC math.PR 50%

Enhancing Exploration in Global Optimization by Noise Injection in the Probability Measures Space

通过在概率测度空间中注入噪声增强全局优化的探索

Gaëtan Serré, Pierre Germain, Samuel Gruffaz, Argyris Kalogeratos

专题命中 其他多模态 :multimodal(abstract)

AI总结 本文通过在概率测度空间中注入噪声,提升全局优化中探索和收敛能力,适用于多种动态配置。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18460 2026-03-03 cs.LG 50%

Learning Boltzmann Generators via Constrained Mass Transport

通过约束质量传输学习Boltzmann生成器

Christopher von Klitzing, Denis Blessing, Henrik Schopmans, Pascal Friederich, Gerhard Neumann

机构 * Autonomous Learning Robots, Karlsruhe Institute of Technology(自动化学习机器人,卡尔斯鲁厄大学技术学院) Artificial Intelligence for Materials Sciences, Karlsruhe Institute of Technology(材料科学人工智能,卡尔斯鲁厄大学技术学院)

专题命中 其他多模态 :multimodal(abstract)

AI总结 本研究提出约束质量传输框架,通过约束KL散度和熵衰减来提升Boltzmann生成器的采样效果,有效避免模式崩溃并提高样本效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00313 2026-03-03 nlin.AO physics.comp-ph 50%

Synchronization, Collective Oscillations, and Information Flow in Duplex Networks

同步、集体振荡与双网络中的信息流

Ali Seif, Mina Zarei

专题命中 其他多模态 :multimodal(abstract)

AI总结 研究双网络中部分同步与集体振荡的机制,揭示多模式动态的形成原理。

Comments 26 pages (21 main and 5 Supplementary), 12 figures (8 main and 4 Supplementary)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21029 2026-03-03 cs.LG 50%

FORCE: Transferable Visual Jailbreaking Attacks via Feature Over-Reliance CorrEction

FORCE:通过特征过度依赖校正实现可转移的视觉劫持攻击

Runqi Lin, Alasdair Paren, Suqin Yuan, Muyang Li, Philip Torr, Adel Bibi, Tongliang Liu

机构 * Sydney AI Centre, The University of Sydney(悉尼人工智能中心,悉尼大学) Department of Engineering Science, University of Oxford(工程科学系,牛津大学)

专题命中 其他多模态 :multimodal(abstract)

AI总结 FORCE方法通过校正特征过度依赖,提升视觉劫持攻击的跨模型可转移性。

Comments Accepted by CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏