arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-02-10 至 2026-02-10 共收录 13 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 13 篇

2601.21648 2026-02-10 cs.CV cs.CY cs.HC 88%

CAF-Mamba: Mamba-Based Cross-Modal Adaptive Attention Fusion for Multimodal Depression Detection

CAF-Mamba:基于Mamba的跨模态自适应注意力融合用于多模态抑郁症检测

Bowen Zhou, Marc-André Fiedler, Ayoub Al-Hamadi

机构 * Neuro-Information Technology Group (NIT) IIKT, Otto von Guericke University Magdeburg Magdeburg, Germany(奥托·冯·格里克大学马格德堡分校神经信息技术小组(NIT)IIKT)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(title,abstract);分类 cs.CV

AI总结 CAF-Mamba通过基于Mamba的跨模态自适应注意力融合框架,提升多模态抑郁症检测的性能。

Comments The paper contains a total of 5 pages and 3 figures. This paper has been accepted for publication in the proceedings of 2026 IEEE ICASSP Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.08282 2026-02-10 cs.CV cs.AI 81%

Tighnari v2: Mitigating Label Noise and Distribution Shift in Multimodal Plant Distribution Prediction via Mixture of Experts and Weakly Supervised Learning

Tighnari v2: 通过专家混合和弱监督学习缓解多模态植物分布预测中的标签噪声和分布偏移

Haixu Liu, Yufei Wang, Tianxiang Xu, Chuancheng Shi, Hongsheng Xing

机构 * The University of Sydney, Sydney, New South Wales, Australia(悉尼大学) The University of New South Wales, Sydney, New South Wales, Australia(新南威尔士大学) School of Software and Microelectronics, Peking University, Beijing, China(北京大学软件与微电子学院) Shandong University of Technology, Zibo, Shandong, China(山东科技大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 Tighnari v2通过专家混合和弱监督学习缓解多模态植物分布预测中的标签噪声和分布偏移,提升预测性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.08713 2026-02-10 cs.CV cs.LG 79%

Towards Understanding Multimodal Fine-Tuning: Spatial Features

迈向多模态微调的理解:空间特征

Lachin Naghashyar, Hunar Batra, Ashkan Khakzar, Philip Torr, Ronald Clark, Christian Schroeder de Witt, Constantin Venhoff

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 本文通过分阶段模型差分技术,揭示了多模态微调过程中视觉特征如何形成及空间关系编码,提升了多模态训练的可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.08077 2026-02-10 cs.LG cs.AI 79%

Multimodal normative modeling in Alzheimers Disease with introspective variational autoencoders

在阿尔茨海默病中使用反思变分自编码器的多模态规范建模

Sayantan Kumar, Peijie Qiu, Aristeidis Sotiras

机构 * Washington University in St Louis(华盛顿大学圣路易斯分校) Washington University in St Louis School of Medicine(华盛顿大学圣路易斯医学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

AI总结 本文提出mmSIVAE,通过结合MOPOE聚合提升多模态数据的规范建模效果,提高参考分布保真度和多模态整合能力,用于阿尔茨海默病的偏差分析。

Comments Conference on Health, Inference, and Learning (CHIL)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.07465 2026-02-10 cs.LG 78%

Multi-Modal Data Fusion for Moisture Content Prediction in Apple Drying

多模态数据融合用于苹果干燥中的含水率预测

Shichen Li, Chenhui Shao

机构 * Department of Mechanical Science and Engineering, University of Illinois at Urbana-Champaign, Urbana, IL 61801, USA(伊利诺伊大学厄巴纳-香槟分校机械科学与工程系) Department of Mechanical Engineering, University of Michigan, Ann Arbor, MI 48109, USA(密歇根大学安娜堡分校机械工程系)

专题命中 多模态训练与对齐 :multi-modal(title,abstract)

AI总结 本文提出多模态数据融合框架,通过融合表格数据和图像数据提高苹果干燥含水率预测的准确性,有效降低误差并增强鲁棒性。

Comments Accepted for publication in the Proceedings of the 53rd North American Manufacturing Research Conference (NAMRC 53), to appear in Manufacturing Letters

Journal ref Manufacturing Letters Volume 44, Supplement, August 2025, Pages 1316-1325

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.08559 2026-02-10 cs.IR 71%

QARM V2: Quantitative Alignment Multi-Modal Recommendation for Reasoning User Sequence Modeling

QARM V2:量化对齐多模态推荐用于推理用户序列建模

Tian Xia, Jiaqi Zhang, Yueyang Liu, Hongjian Dou, Tingya Yin, Jiangxia Cao, Xulei Liang, Tianlu Xie, Lihao Liu, Xiang Chen, Shen Wang, Changxin Lao, Haixiang Gan, Jinkai Yu, Keting Cen, Lu Hao, Xu Zhang, Qiqiang Zhong, Zhongbo Sun, Yiyu Wang, Shuang Yang, Mingxin Wen, Xiangyu Wu, Shaoguo Liu, Tingting Gao, Zhaojie Liu, Han Li, Kun Gai

专题命中 多模态训练与对齐 :multi-modal(title)

AI总结 QARM V2通过量化对齐多模态推荐方法,解决推荐系统中用户序列建模的语义理解与业务需求不匹配问题。

Comments Work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.06973 2026-02-10 cs.CL cs.AI cs.LG 62%

Does Visual Rendering Bypass Tokenization? Investigating Script-Tokenizer Misalignment in Pixel-Based Language Models

视觉渲染能否绕过分词?探究基于像素的语言模型中脚本-分词器不一致问题

Lucky Susanto, Musa Izzanardi Wijanarko, Khumaisa Nur'aini, Farid Adilazuarda, Alham Fikri Aji, Derry Tanti Wijaya

机构 * Monash University Indonesia(墨尔本大学印尼分校) MBZUAI Boston University(波士顿大学) University of Edinburgh(爱丁堡大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CL、cs.AI

AI总结 本研究探讨了基于像素的语言模型中视觉渲染是否能绕过分词约束,发现重新引入文本分词器加剧了分词不一致问题,自定义分词器在性能上表现更优。

Comments Submitted to ARR January

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.08208 2026-02-10 cs.CL cs.HC 57%

LLMs and people both learn to form conventions -- just not with each other

大语言模型和人类都学会形成惯例——但并不是彼此之间

Cameron R. Jones, Agnese Lombardi, Kyle Mahowald, Benjamin K. Bergen

机构 * Department of Psychology, Stony Brook University(心理学系,石溪大学) Department of Cognitive Science, University of California San Diego(认知科学系,加州圣地亚哥大学) Department of Philology, Literature, and Linguistics, University of Pisa(philology、文学与语言学系,比萨大学) Department of Linguistics, University of Texas at Austin(语言学系,德克萨斯大学奥斯汀分校)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CL

AI总结 研究发现人类和AI在同类型对话中能形成惯例,但人机对话效果较差,表明对话协调需要共同的解释偏见。

Comments 10 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.07540 2026-02-10 cs.CV cs.LG 57%

LLM-Guided Diagnostic Evidence Alignment for Medical Vision-Language Pretraining under Limited Pairing

基于LLM的诊断证据对齐用于有限配对情况下的医学视觉-语言预训练

Huimin Yan, Liang Bai, Xian Yang, Long Chen

机构 * Institute of Intelligent Information Processing, Shanxi University(山西大学智能信息处理研究所) Alliance Manchester Business School, The University of Manchester(曼彻斯特大学阿利安斯商学院) Department of Computer Science and Engineering, The Hong Kong University of Science and Technology(香港科技大学计算机科学与工程系)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

AI总结 本文提出LLM引导的诊断证据对齐方法,旨在解决有限配对数据下医学视觉-语言预训练的诊断表示学习问题,通过提取关键诊断证据提升跨模态对齐效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.01212 2026-02-10 cs.CV 57%

TSJNet: A Multi-modality Target and Semantic Awareness Joint-driven Image Fusion Network

TSJNet: 一种多模态目标和语义意识联合驱动的图像融合网络

Yuchan Jie, Yushen Xu, Xiaosong Li, Huafeng Li, Haishu Tan, Feiping Nie

机构 * School of Physics and Optoelectronic Engineering, Foshan University(物理与光电工程学院,佛山大学) School of Information Engineering and Automation, Kunming University of Science and Technology(信息工程与自动化学院,昆明理工大学) School of Artificial Intelligence, Optics and Electronics (i0PEN), School of Computer Science, Northwestern Polytechnical University(人工智能、光学与电子学院(i0PEN),计算机科学学院,西北工业大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

AI总结 TSJNet通过联合驱动的多模态图像融合网络提升目标检测与语义分割的性能,实现7.97%和10.88%的精度提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.07154 2026-02-10 cs.LG cs.AI 57%

Beyond Pooling: Matching for Robust Generalization under Data Heterogeneity

超越池化:在数据异质性下实现鲁棒泛化的匹配

Ayush Roy, Rudrasis Chakraborty, Lav Varshney, Vishnu Suresh Lokhande

机构 * SUNY Buffalo(纽约州立大学布法罗分校) Lawrence Livermore National Lab(劳伦斯利弗莫尔国家实验室) SUNY Stony Brook(纽约州立大学石溪分校)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

AI总结 本文提出一种匹配框架,通过自适应质心选择样本并迭代优化表示分布,以在数据异质性下实现更稳健的泛化,尤其在零样本医学异常检测中取得显著提升。

Comments AISTATS 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.05825 2026-02-10 cs.HC 50%

ToMigo: Interpretable Design Concept Graphs for Aligning Generative AI with Creative Intent

ToMigo:可解释的设计概念图用于对齐生成式AI与创意意图

Lena Hegemann, Xinyi Wen, Michael A. Hedderich, Tarmo Nurmi, Hariharan Subramonyam

专题命中 多模态训练与对齐 :multimodal(abstract)

AI总结 ToMigo通过设计概念图实现生成式AI与创意意图的对齐,提供可解释的交互方式提升用户控制力。

Comments 18 pages, 10 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.07706 2026-02-10 cs.LG 50%

Dense Feature Learning via Linear Structure Preservation in Medical Data

通过线性结构保持实现密集特征学习:医学数据

Yuanyun Zhang, Mingxuan Zhang, Siyuan Li, Zihan Wang, Haoran Chen, Wenbo Zhou, Shi Li

机构 * Independent Researcher(独立研究者) The Chinese University of Hong Kong(香港中文大学) Columbia University(哥伦比亚大学)

专题命中 多模态训练与对齐 :multimodal(abstract)

AI总结 本文提出密集特征学习方法,通过线性结构保持提升医学数据表示的稳定性与可解释性,实现更优的下游性能。

Comments ICLR Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏