arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6903 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6903 篇

1901.00877 2019-01-07 cs.LG q-bio.QM stat.ML 78%

A Network-based Multimodal Data Fusion Approach for Characterizing Dynamic Multimodal Physiological Patterns

Miaolin Fan, Chun-An Chou, Sheng-Che Yen, Yingzi Lin

专题命中 多模态训练与对齐 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
1701.03619 2018-11-21 stat.ML cs.LG physics.data-an 78%

Diffusion-based nonlinear filtering for multimodal data fusion with application to sleep stage assessment

Ori Katz, Ronen Talmon, Yu-Lun Lo, Hau-Tieng Wu

专题命中 多模态训练与对齐 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
1807.07611 2018-08-21 physics.flu-dyn 78%

Multimodal Functions as Flow Signatures in Complex Porous Media

Branko Bijeljic, Ali Q. Raeini, Qingyang Lin, Martin J. Blunt

专题命中 多模态训练与对齐 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
1807.01298 2018-07-04 cs.LG stat.ML 78%

Generalized Bilinear Deep Convolutional Neural Networks for Multimodal Biometric Identification

Sobhan Soleymani, Amirsina Torfi, Jeremy Dawson, Nasser M. Nasrabadi

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Comments Accepted in 2018 IEEE International Conference on Image Processing

详情

展开后加载摘要…

URL PDF HTML 收藏
1701.01352 2018-02-14 stat.AP cs.IT math.IT 78%

Compressive Sensing-Based Detection with Multimodal Dependent Data

Thakshila Wimalajeewa, Pramod K. Varshney

专题命中 多模态训练与对齐 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
1702.04815 2017-11-10 cs.IR 78%

Multimodal Content Representation and Similarity Ranking of Movies

Konstantinos Bougiatiotis, Theodore Giannakopoulos

专题命中 多模态训练与对齐 :multimodal(title);multi-modal(abstract)

Comments Preliminary work

详情

展开后加载摘要…

URL PDF HTML 收藏
1711.00614 2017-11-03 cs.RO cs.LG 78%

A Multimodal Anomaly Detector for Robot-Assisted Feeding Using an LSTM-based Variational Autoencoder

Daehyung Park, Yuuna Hoshi, Charles C. Kemp

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Comments 8 pages, under review

详情

展开后加载摘要…

URL PDF HTML 收藏
1710.08306 2017-10-24 cs.NI cs.LG 78%

CollabLoc: Privacy-Preserving Multi-Modal Localization via Collaborative Information Fusion

Vidyasagar Sadhu, Dario Pompili, Saman Zonouz, Vincent Sritapan

专题命中 多模态训练与对齐 :multi-modal(title,abstract)

Comments 9 pages, 26th International Conference on Computer Communication and Networks (ICCCN), Vancouver, BC, Canada, 2017, pp. 1-9

详情

展开后加载摘要…

URL PDF HTML 收藏
1705.09406 2017-08-02 cs.LG 78%

Multimodal Machine Learning: A Survey and Taxonomy

Tadas Baltrušaitis, Chaitanya Ahuja, Louis-Philippe Morency

专题命中 多模态训练与对齐 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
1707.00750 2017-07-05 cs.NE 78%

Structure Optimization for Deep Multimodal Fusion Networks using Graph-Induced Kernels

Dhanesh Ramachandram, Michal Lisicki, Timothy J. Shields, Mohamed R. Amer, Graham W. Taylor

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Comments Proceedings of the 25th European Symposium on Artificial Neural Networks, Computational Intelligence and Machine Learning, April 2017, Bruges, Belgium

详情

展开后加载摘要…

URL PDF HTML 收藏
1703.10511 2017-03-31 cs.SI 78%

Multimodal Network Alignment

Huda Nassar, David F. Gleich

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Comments 14 pages, 6 figures, Siam Data Mining 2017

详情

展开后加载摘要…

URL PDF HTML 收藏
1703.08970 2017-03-28 cs.LG 78%

Multimodal deep learning approach for joint EEG-EMG data compression and classification

Ahmed Ben Said, Amr Mohamed, Tarek Elfouly, Khaled Harras, Z. Jane Wang

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Comments IEEE Wireless Communications and Networking Conference (WCNC), 2017

详情

展开后加载摘要…

URL PDF HTML 收藏
1702.01992 2017-02-08 stat.ML cs.LG 78%

Gated Multimodal Units for Information Fusion

John Arevalo, Thamar Solorio, Manuel Montes-y-Gómez, Fabio A. González

专题命中 多模态训练与对齐 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
1611.01891 2016-11-08 stat.ML cs.LG 78%

Joint Multimodal Learning with Deep Generative Models

Masahiro Suzuki, Kotaro Nakayama, Yutaka Matsuo

专题命中 多模态训练与对齐 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
1609.05111 2016-09-19 cs.IT math.IT 78%

Detection with Multimodal Dependent Data Using Low Dimensional Random Projections

Thakshila Wimalajeewa, Pramod K. Varshney

专题命中 多模态训练与对齐 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
1605.02710 2016-05-26 cs.SI 78%

Tracking Illicit Drug Dealing and Abuse on Instagram using Multimodal Analysis

Xitong Yang, Jiebo Luo

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Comments 5 pages, 5 figures, under review

详情

展开后加载摘要…

URL PDF HTML 收藏
1605.00766 2016-05-04 cs.CR 78%

Walk-Unlock: Zero-Interaction Authentication Protected with Multi-Modal Gait Biometrics

Babins Shrestha, Manar Mohamed, Nitesh Saxena

专题命中 多模态训练与对齐 :multi-modal(title,abstract)

Comments 20 pages, 4 figures, under submission

详情

展开后加载摘要…

URL PDF HTML 收藏
1509.08095 2015-09-29 physics.soc-ph cs.SI physics.data-an 78%

User-based representation of time-resolved multimodal public transportation networks

Laura Alessandretti, Márton Karsai, Laetitia Gauvin

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Comments 24 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.11788 2024-02-20 cs.CV cs.AI 77%

MM-SurvNet: Deep Learning-Based Survival Risk Stratification in Breast Cancer Through Multimodal Data Fusion

Raktim Kumar Mondol, Ewan K. A. Millar, Arcot Sowmya, Erik Meijering

专题命中 多模态训练与对齐 :multimodal(title,comments);分类 cs.CV、cs.AI

Comments Keywords: Multimodal Fusion, Breast Cancer, Whole Slide Images, Survival Prediction

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09105 2026-08-18 cs.AI 版本更新 77%

SMA: Who Said That? Auditing Membership Leakage in Semi-Black-box RAG Controlling

SMA:谁说的?半黑盒RAG控制中的成员泄露审计

Shixuan Sun, Siyuan Liang, Jianjie Huang, Jingzhi Li, Xiaochun Cao

机构 * Sun Yat-Sen University(孙中山大学) Nanyang Technological University(南洋理工大学) University of Chinese Academy of Science(中国科学院大学) Zhongguancun Academy(中关村学院) Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所)

专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract_cn);cross-modal(abstract);分类 cs.AI

AI总结 本研究针对半黑盒RAG控制中的成员泄露问题,提出首个源感知成员审计SMA,通过零阶优化归因估计机制与跨模态归因技术,实现生成内容的细粒度来源归因,为复杂生成系统的数据来源审计提供新视角。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.13980 2026-08-17 cs.CV 新提交 77%

FIRM: Fine-Grained Intra-Token Representation of Masks for Remote Sensing Reasoning Segmentation

FIRM:面向遥感推理分割的掩码细粒度令牌内表示

Weidong Tang, Kaiyu Li, Yikai Wang, Yanan Wu, Haotian Gan, Shihong Wang, Xiangyong Cao

专题命中 多模态训练与对齐 :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV

AI总结 本文提出FIRM模型,通过预测视觉令牌内的r×r二值子单元掩码代码结合轻量连续渲染器优化边界,在5个遥感推理分割基准上取得领先结果,提升了细粒度分割性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.09227 2026-08-11 cs.AI 新提交 77%

Omni2LoRA: Coherence-Preserving Parametric Memory for Efficient Omni Language Models

Omni2LoRA:用于高效全模态语言模型的保持一致性的参数化存储器

Puneet Mathur, Manan Suri, Dinesh Manocha

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);audio-visual(abstract);分类 cs.AI

AI总结 Omni2LoRA是一种保持一致性的参数化存储器压缩框架,通过优化秩分配策略提升全模态语言模型效率,在视听问答任务中优于多种基线,大幅缩短推理时间。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.06948 2026-08-10 cs.AI 新提交 77%

LMM Modality Transfer: A Pre-requisite for Autonomous GIS Agents

多模态大模型模态迁移:自主地理信息系统智能体的先决条件

Ivan Majic, Zexian Huang, Franziska Hübl, Krzysztof Janowicz, Meilin Shi, Mina Karimi, Zilong Liu, Alexandra Fortacz-Lazan

机构 * Graz University of Technology(格拉茨工业大学) University of Vienna(维也纳大学) University of Liverpool(利物浦大学)

专题命中 多模态训练与对齐 :multimodal(abstract,abstract_cn);multi-modal(abstract);分类 cs.AI

AI总结 本文针对自主GIS智能体的先决条件,提出LMM模态迁移任务,发现现有LMM在图像与文本模态间传递空间信息的能力不足,需加强多模态对齐以实现地理空间理解。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.01635 2026-08-04 cs.CV 新提交 77%

Mitigating Visual Degradation in MLLMs via Spatial-Spectral Visual Anchor Learning

通过空间-光谱视觉锚学习缓解多模态大语言模型(MLLMs)中的视觉退化

Qianlong Yang, Bowen Ye, Xianda Guo, Yanlun Peng, Wenke Huang, Hongyuan Zhang, Yulei Jia

机构 * China University of Petroleum (East China)(中国石油大学(华东)) Shanghai Jiao Tong University(上海交通大学) Wuhan University(武汉大学) Great Wall Motor(长城汽车) Nanyang Technological University(南洋理工大学) The University of Hong Kong(香港大学)

专题命中 多模态训练与对齐 :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV

AI总结 针对MLLMs推理时的视觉表示退化问题,提出SSVAL方法,通过VAPI及辅助对齐损失实现稳定视觉锚,性能优于现有方法。

Comments This paper has been accepted by ACM MM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.11581 2026-07-14 cs.CV 新提交 77%

Actor as Its Own Critic: Unifying Region Understanding and Localization via CycleGRPO

作为自身评论家的智能体:通过循环组相对策略优化统一区域理解与定位

Xin Zhang, Haochen Wang, Yikang Zhou, Jason Li, Robby T. Tan

机构 * National University of Singapore(新加坡国立大学) University of Chinese Academy of Sciences(中国科学院大学) Nanyang Technological University(南洋理工大学) Wuhan University(武汉大学)

专题命中 多模态训练与对齐 :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV

AI总结 研究针对多模态大语言模型的区域理解与定位问题,提出循环组相对策略优化框架CycleGRPO,利用任务对偶性构建自我评估范式,仅需区域输入,通过质量感知奖励评估字幕,在多基准测试中提升能力,为推进MLLMs像素级能力提供新途径。

Comments Accepted to ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.09488 2026-07-13 cs.CV 新提交 77%

SigLIP-HD by Fine-to-Coarse Supervision

通过从细到粗的监督实现SigLIP-HD

Lihe Yang, Zhen Zhao, Hengshuang Zhao

机构 * The University of Hong Kong(香港大学) Shanghai AI Laboratory(上海人工智能实验室)

专题命中 多模态训练与对齐 :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV

AI总结 研究如何在低成本下实现精细视觉感知,提出SigLIP-HD,采用从细到粗监督设计,基于SigLIP 2模型构建,在相同推理预算下能产生更好视觉令牌,在多基准测试中结果优于基线模型。

Comments ICLR 2026. Code and model: https://github.com/LiheYoung/SigLIP-HD

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.21664 2026-07-08 cs.CV 版本更新 77%

HumanOmni-Speaker: Identifying Who said What and When

HumanOmni-Speaker: 识别说话者说了什么以及何时

Detao Bai, Zhiheng Ma, Xihan Wei

机构 * Tongyi Lab Alibaba Group(阿里云实验室) Shenzhen University of Advanced Technology(深圳先进技术大学)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);omni-modal(abstract);分类 cs.CV

AI总结 本文提出HumanOmni-Speaker,通过视觉注册说话人辨识与识别方法,解决多人物对话中说话者身份与时间的识别问题,实现端到端的多模态协同。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.27947 2026-06-29 cs.CV 新提交 77%

Understanding How MLLMs Describe Artworks Using Token Activation Maps

理解多模态大语言模型如何通过Token激活图描述艺术品

Nicola Fanelli, Pasquale De Marinis, Raffaele Scaringi, Eva Cetinic, Gennaro Vessio, Giovanna Castellano

机构 * University of Bari Aldo Moro(巴里阿尔多莫罗大学) University of Zurich(苏黎世大学)

专题命中 多模态训练与对齐 :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV

AI总结 使用Token激活图(TAM)分析MLLMs在描述艺术品时对视觉区域的依赖,发现不同语义类别的token在视觉基础上有显著差异,并比较了TAM与SAM~3开放词汇分割。

Comments Accepted at PRESTIGE workshop at ICPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.19915 2026-06-19 cs.CV 新提交 77%

SpatialSV: Internalizing Interpretable 3D Spatial Awareness in MLLMs via Task-Oriented Visual Supervision

SpatialSV: 通过任务导向的视觉监督在多模态大语言模型中内化可解释的3D空间感知

Jiayu Tang, Yuchen Zhou, Chao Gou

机构 * School of Intelligent Systems Engineering, Sun Yat-sen University(中山大学智能工程学院)

专题命中 多模态训练与对齐 :MLLM(summary_cn);multimodal(abstract);分类 cs.CV

AI总结 提出SpatialSV框架,通过任务导向的视觉监督将MLLM的2D特征提升为显式3D表示(深度图、相机姿态、点云),实现可解释的3D空间感知内化,无需外部工具,并在半监督设置中展现强泛化能力。

Comments Accepted by IJCAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.26656 2026-05-27 cs.CV 77%

DV-SFT: Direct Vision Supervision for Fine-Grained Visual Understanding

DV-SFT: 直接视觉监督用于细粒度视觉理解

Jianfei Zhao, Feng Zhang, Xin Sun, Chong Feng, Bing Wang, Zhixing Tan

机构 * School of Computer Science and Technology, Beijing Institute of Technology(北京理工大学计算机科学与技术学院) Zhongguancun Academy(中关村学院) Beihang University(北航) Zhongguancun Laboratory(中关村实验室) Southeast Academy of Information Technology, Beijing Institute of Technology(北京理工大学信息科学技术东南学院)

专题命中 多模态训练与对齐 :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV

AI总结 提出DV-SFT方法,通过为视觉令牌构建显式令牌级监督信号,利用OCR场景中的直接视觉-文本对应关系,在不修改模型架构或增加前向传播的情况下,显著提升多模态大语言模型的细粒度视觉理解能力。

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏