arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-02-03 至 2026-02-03 共收录 11 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 其他多模态 11 篇

2512.11325 2026-02-03 cs.CV cs.AI 81%

Robust MLLM Unlearning via Visual Knowledge Distillation

通过视觉知识蒸馏实现鲁棒的多模态大语言模型去学习

Yuhang Wang, Zhenxing Niu, Haoxuan Ji, Guangyu He, Haichang Gao, Gang Hua

机构 * Xidian University, China(西安电子科技大学) XJTU University, China(西安交通大学) Amazon.com, USA(亚马逊公司)

专题命中 其他多模态 :MLLM(title,abstract);分类 cs.CV、cs.AI

AI总结 本文提出通过视觉知识蒸馏实现多模态大语言模型的鲁棒去学习,有效保留文本知识并提升效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01884 2026-02-03 cs.AI cs.LG 79%

Entropy-Guided Data-Efficient Training for Multimodal Reasoning Reward Models

熵引导的数据高效训练用于多模态推理奖励模型

Shidong Yang, Tongwen Huang, Hao Wen, Yong Wang, Li Chen, Xiangxiang Chu

机构 * School of Software, Tsinghua University(清华大学软件学院) AMAP, Alibaba Group(阿里巴巴集团AMAP)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.AI

AI总结 本文提出熵引导训练方法,通过熵指导数据筛选和训练策略提升多模态推理奖励模型的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00590 2026-02-03 cond-mat.mtrl-sci cond-mat.soft cs.AI cs.LG physics.data-an 79%

Multimodal Machine Learning for Integrating Heterogeneous Analytical Systems

多模态机器学习用于整合异质分析系统

Shun Muroga, Hideaki Nakajima, Taiyo Shimizu, Kazufumi Kobashi, Kenji Hata

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.AI

AI总结 本文提出多模态机器学习框架,通过整合多种分析数据,提升复杂材料的表征精度与解释性。

Comments 12 pages, 4 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00067 2026-02-03 cs.LG cs.AI 79%

Modality as Heterogeneity: Node Splitting and Graph Rewiring for Multimodal Graph Learning

模态作为异质性:用于多模态图学习的节点分裂与图重 wiring

Yihan Zhang, Ercan E. Kuruoglu

机构 * Institute of Data and Information, Shenzhen International Graduate School, Tsinghua University(数据与信息研究所,深圳国际研究生院,清华大学)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.AI

AI总结 NSG-MoE通过节点分裂与图重 wiring机制,结合结构化MoE架构,有效解决多模态图学习中的模态混淆问题,提升模型的结构信息保留与泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00360 2026-02-03 cs.LG 78%

Leveraging Textual-Cues for Enhancing Multimodal Sentiment Analysis by Object Recognition

利用文本提示增强多模态情感分析的物体识别

Sumana Biswas, Karen Young, Josephine Griffith

专题命中 其他多模态 :multimodal(title,abstract)

AI总结 本文提出TEMSA方法,通过结合物体识别和文本信息提升多模态情感分析的准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00171 2026-02-03 stat.ML cs.LG 78%

Uncertainty-Aware Multimodal Learning via Conformal Shapley Intervals

通过置信区间实现的不确定性感知多模态学习

Mathew Chandy, Michael Johnson, Judong Shen, Devan V. Mehrotra, Hua Zhou, Jin Zhou, Xiaowu Dai

机构 * Department of Statistics and Data Science, UCLA(统计与数据科学系, UCLA) Merck & Co(默克公司) Department of Biostatistics, UCLA(生物统计学系, UCLA) Departments of Statistics and Data Science, and of Biostatistics, UCLA(统计与数据科学系以及生物统计学系, UCLA)

专题命中 其他多模态 :multimodal(title,abstract)

AI总结 本文提出了一种结合Shapley值与符合推断的框架,用于构建不确定性感知的重要性区间,从而实现更可靠和可解释的多模态学习。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00597 2026-02-03 cs.CL cs.AI 76%

Hermes the Polyglot: A Unified Framework to Enhance Expressiveness for Multimodal Interlingual Subtitling

赫мес:一种增强多模态跨语言字幕表达力的统一框架

Chaoqun Cui, Shijing Wang, Liangbin Huang, Qingqing Gu, Zhaolong Huang, Xiao Zeng, Wenji Mao

机构 * MAIS, Institute of Automation, Chinese Academy of Sciences Beijing China School of AI, University of\ Academy of Sciences Beijing China Beijing Jiaotong University Beijing China Geely AI lab Ningbo Zhejiang China MAIS, Institute of Automation, Chinese Academy of Sciences School of AI, University of\ Academy of Sciences Beijing Jiaotong University Geely AI lab

专题命中 其他多模态 :multimodal(title);分类 cs.CL、cs.AI

AI总结 赫мес通过整合说话人分离、术语识别和表达力增强模块,提升了多模态跨语言字幕的表达力和连贯性,实现了最先进的字幕生成性能。

Comments Accepted to The Web Conference (WWW) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01034 2026-02-03 cs.AI cs.CL 62%

Discovering Process-Outcome Credit in Multi-Step LLM Reasoning

在多步骤LLM推理中发现过程-结果信用

Xiangwei Wang, Wei Wang, Ken Chen, Nanduni Nimalsiri, Saman Halgamuge

机构 * The University of Melbourne(墨尔本大学)

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CL、cs.AI

AI总结 本文提出了一种新的框架,通过分步边际信息增益机制和解耦掩码策略,提升多步骤LLM推理的样本效率和准确性,并增强模型的分布外鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10040 2026-02-03 cs.SI cs.AI 57%

Exploring Content and Social Connections of Fake News with Explainable Text and Graph Learning

探索虚假新闻的内容与社会联系:可解释文本与图学习

Vítor N. Lourenço, Aline Paes, Tillman Weyde

机构 * Universidade Federal Fluminense, Niter\' o i, Rio de Janeiro, Brazil , City St George's, University of London, United Kingdom

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

AI总结 本文提出结合内容、社交媒体和图学习的可解释框架,通过多模态信息提升事实核查性能,并评估其解释性、可信度和鲁棒性。

Comments Accepted to publication at the 35th Brazilian Conference on Intelligent Systems, BRACIS 2025. -- This submitted manuscript has not undergone any post-submission improvements or corrections. The Version of Record of this contribution will be provided when available

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01924 2026-02-03 cs.LG cs.CY 50%

Bayesian Integration of Nonlinear Incomplete Clinical Data

贝叶斯非线性不完整临床数据整合

Lucía González-Zamorano, Nuria Balbás-Esteban, Vanessa Gómez-Verdejo, Albert Belenguer-Llorens, Carlos Sevilla-Salcedo

机构 * Department of Signal Theory and Communications, Universidad Carlos III de Madrid(信号理论与通信系,卡洛斯三世大学)

专题命中 其他多模态 :multimodal(abstract)

AI总结 BIONIC通过联合生成-判别潜在架构整合异质多模态临床数据,在缺失性下实现稳健学习和高判别性能,同时提供内在可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00835 2026-02-03 stat.ML cs.LG 50%

Score-based Metropolis-Hastings for Fractional Langevin Algorithms

基于分数朗之万算法的评分法梅特罗波利斯-哈斯廷斯方法

Ahmed Aloui, Junyi Liao, Ali Hasan, Jose Blanchet, Vahid Tarokh

机构 * Duke University(杜克大学) Morgan Stanley(摩根士丹利) Stanford University(斯坦福大学)

专题命中 其他多模态 :multimodal(abstract)

AI总结 本文提出MAFLA算法,通过基于评分的校正机制改进分数朗之万算法,提升在重尾分布采样中的精度和稳定性。

详情

展开后加载摘要…

URL PDF HTML 收藏