arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4884 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 其他多模态 4884 篇

2604.09076 2026-04-13 cs.CV 80%

Cross-Modal Knowledge Distillation from Spatial Transcriptomics to Histology

从空间转录组到组织学的跨模态知识蒸馏

Arbel Hizmi, Artemii Bakulin, Shai Bagon, Nir Yosef

机构 * Weizmann Institute of Science(魏茨曼科学研究所) Reichman University(赖希曼大学)

专题命中 其他多模态 :cross-modal(title,abstract);分类 cs.CV

AI总结 本文提出利用空间转录组与H&E数据进行跨模态蒸馏,将转录组学的 niches 结构转移到仅依赖组织学的模型中,提升与转录组学 niches 结构的一致性,并通过细胞类型分析验证生物意义的邻近组成。

Comments Accepted to the CVMI Workshop at CVPR 2026. Project page: https://cross-modal-distillation.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.14879 2025-10-31 cs.HC cs.AI 80%

Vital Insight: Assisting Experts' Context-Driven Sensemaking of Multi-modal Personal Tracking Data Using Visualization and Human-In-The-Loop LLM

Jiachen Li, Xiwen Li, Justin Steinberg, Akshat Choube, Bingsheng Yao, Xuhai Xu, Dakuo Wang, Elizabeth Mynatt, Varun Mishra

机构 * Northeastern University(东北大学) Columbia University(哥伦比亚大学)

专题命中 其他多模态 :multi-modal(title,abstract);分类 cs.AI

Comments Jiachen Li, Xiwen Li, Justin Steinberg, Akshat Choube, Bingsheng Yao, Xuhai Xu, Dakuo Wang, Elizabeth Mynatt, and Varun Mishra. 2025. Vital Insight: Assisting Experts' Context-Driven Sensemaking of Multi-modal Personal Tracking Data Using Visualization and Human-in-the-Loop LLM. Proc. ACM Interact. Mob. Wearable Ubiquitous Technol. 9, 3, Article 101 (September 2025), 37 pages

Journal ref Proc. ACM Interact. Mob. Wearable Ubiquitous Technol. 9 (2025) 101:1-37

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21812 2025-10-28 cs.IR cs.AI cs.LG 80%

Unifying Inductive, Cross-Domain, and Multimodal Learning for Robust and Generalizable Recommendation

Chanyoung Chung, Kyeongryul Lee, Sunbin Park, Joyce Jiyoung Whang

机构 * KAIST(韩国科学技术院)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.AI

Comments 7 pages, 3 figures, and 4 tables. International Workshop on Multimodal Generative Search and Recommendation (MMGenSR) at The 34th ACM International Conference on Information and Knowledge Management (CIKM 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2006.05327 2025-08-29 cs.CV cs.HC 80%

mEBAL: A Multimodal Database for Eye Blink Detection and Attention Level Estimation

Roberto Daza, Aythami Morales, Julian Fierrez, Ruben Tolosana

机构 * Biometrics and Data Pattern Analytics, BiDA-Lab, Universidad Autonoma de Madrid, Spain(生物特征与数据模式分析,BiDA实验室,马德里自治大学,西班牙)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV

Journal ref International Conference on Multimodal Interaction (ICMI), 2020, pp. 32-36

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.06024 2024-03-12 cs.CV cs.ET cs.LG 80%

Semi-Supervised Multimodal Multi-Instance Learning for Aortic Stenosis Diagnosis

Zhe Huang, Xiaowei Yu, Benjamin S. Wessler, Michael C. Hughes

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV

Comments Echocardiography; Multimodal; Semi-supervised Learning; Multiple-Instance Learning

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.10454 2023-08-22 cs.AI cs.CY cs.HC 80%

Elucidating STEM Concepts through Generative AI: A Multi-modal Exploration of Analogical Reasoning

Chen Cao, Zijian Ding, Gyeong-Geon Lee, Jiajun Jiao, Jionghao Lin, Xiaoming Zhai

专题命中 其他多模态 :multi-modal(title,abstract);分类 cs.AI;multimodal(journal_ref)

Journal ref IJCAI2023 Symposium on Multimodal Reasoning with LLM

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.02521 2023-07-18 cs.CV cs.LG 80%

Exploiting Partial Common Information Microstructure for Multi-Modal Brain Tumor Segmentation

Yongsheng Mei, Guru Venkataramani, Tian Lan

专题命中 其他多模态 :multi-modal(title,abstract);分类 cs.CV;multimodal(comments)

Comments 2023 ICML Workshop on Machine Learning for Multimodal Healthcare Data (ML4MHD)

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.02555 2022-08-05 eess.IV cs.CV 80%

Multi-modal volumetric concept activation to explain detection and classification of metastatic prostate cancer on PSMA-PET/CT

Rosa C. J. Kraaijveld, Marielle E. P. Philippens, Wietse S. C. Eppinga, Ina M. Jürgenliemk-Schulz, Kenneth G. A. Gilhuijs, Petra S. Kroon, Bas H. M. van der Velden

专题命中 其他多模态 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted as: Kraaijveld, R.C.J., Philippens, M.E.P., Eppinga, W.S.C., Jürgenliemk-Schulz, I.M., Gilhuijs, K.G.A., Kroon, P.S., van der Velden, B.H.M. "Multi-modal volumetric concept activation to explain detection and classification of metastatic prostate cancer on PSMA-PET/CT." MICCAI workshop on Interpretability of Machine Intelligence in Medical Image Computing (iMIMIC), 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.09242 2022-06-22 cs.CV cs.LG 80%

GaLeNet: Multimodal Learning for Disaster Prediction, Management and Relief

Rohit Saha, Mengyi Fang, Angeline Yasodhara, Kyryl Truskovskyi, Azin Asgarian, Daniel Homola, Raahil Shah, Frederik Dieleman, Jack Weatheritt, Thomas Rogers

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV

Comments Accepted to CVPR 2022 Workshop on Multimodal Learning for Earth and Environment

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.03759 2021-10-11 cs.AI cs.HC cs.LG 80%

Explanation as a process: user-centric construction of multi-level and multi-modal explanations

Bettina Finzel, David E. Tafler, Stephan Scheele, Ute Schmid

专题命中 其他多模态 :multi-modal(title,abstract);分类 cs.AI

Comments 14 pages, 5 figures; Camera-ready submission to KI2021; The final authenticated publication is available online at https://link.springer.com/chapter/10.1007/978-3-030-87626-5_7 ; code is available at https://gitlab.rz.uni-bamberg.de/cogsys/public/multi-level-multi-modal-explanation

详情

展开后加载摘要…

URL PDF HTML 收藏
2011.02840 2021-05-05 eess.IV cs.CV cs.LG 80%

DR-Unet104 for Multimodal MRI brain tumor segmentation

Jordan Colman, Lei Zhang, Wenting Duan, Xujiong Ye

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV

Comments Part of the Multimodal Brain Tumor Segmentation 2020 Challenge conference proceedings

Journal ref BrainLes 2020. Lecture Notes in Computer Science, vol 12659, pp 410-419

详情

展开后加载摘要…

URL PDF HTML 收藏
2009.11195 2021-01-13 cs.HC cs.AI 80%

Studying Person-Specific Pointing and Gaze Behavior for Multimodal Referencing of Outside Objects from a Moving Vehicle

Amr Gomaa, Guillermo Reyes, Alexandra Alles, Lydia Rupp, Michael Feld

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.AI

Journal ref In Proceedings of the 2020 International Conference on Multimodal Interaction, pp. 501-509. 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2004.08388 2020-04-21 cs.CV 80%

Multi-Modal Face Anti-Spoofing Based on Central Difference Networks

Zitong Yu, Yunxiao Qin, Xiaobai Li, Zezheng Wang, Chenxu Zhao, Zhen Lei, Guoying Zhao

专题命中 其他多模态 :multi-modal(title,abstract);分类 cs.CV

Comments 1st place in "Track Multi-Modal" of ChaLearn Face Anti-spoofing Attack Detection Challenge@CVPR2020; Accepted by CVPR2020 Media Forensics Workshop. arXiv admin note: text overlap with arXiv:2003.04092

详情

展开后加载摘要…

URL PDF HTML 收藏
1803.02124 2018-03-07 cs.AI cs.HC 80%

MIRIAM: A Multimodal Chat-Based Interface for Autonomous Systems

Helen Hastie, Francisco J. Chiyah Garcia, David A. Robb, Pedro Patron, Atanas Laskov

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.AI

Comments 2 pages, ICMI'17, 19th ACM International Conference on Multimodal Interaction, November 13-17 2017, Glasgow, UK

详情

展开后加载摘要…

URL PDF HTML 收藏
1710.04782 2017-10-16 cs.CV 80%

Multimodal and Multiscale Deep Neural Networks for the Early Diagnosis of Alzheimer's Disease using structural MR and FDG-PET images

Donghuan Lu, Karteek Popuri, Weiguang Ding, Rakesh Balachandar, Mirza Faisal Beg

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV

Comments 12 pages, 4 figures, Alzheimer's disease, deep learning, multimodal, early diagnosis, multiscale

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.15781 2026-04-20 cs.HC 80%

ReVis: Towards Reusable Image-Based Visualizations with MLLMs

ReVis:面向基于图像的可视化可重用性的方法

Xiaolin Wen, Changlin Li, Manusha Karunathilaka, Can Liu, Fangzhuo Jin, Yong Wang

专题命中 其他多模态 :MLLM(summary_cn,abstract)

AI总结 ReVis通过引入领域特定语言和MLLM管道,实现复杂可视化的灵活重用,支持分解与再现,并通过交互界面实现数据更新与编码定制。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.12556 2026-08-19 cs.CV 79%

M2Retinexformer: Multi-Modal Retinexformer for Low-Light Image Enhancement

M2Retinexformer:多模态Retinexformer用于低光图像增强

Youssef Aboelwafa, Hicham G. Elmongui, Marwan Torki

机构 * Alexandria University, Egypt(亚历山大大学,埃及)

专题命中 其他多模态 :multi-modal(title,abstract);分类 cs.CV

AI总结 本文提出M2Retinexformer,通过融合深度线索、亮度先验和语义特征,改进低光图像增强效果,实验表明优于现有方法。

Comments Accepted at 2026 IEEE International Conference on Image Processing (ICIP)

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.28336 2026-08-17 cs.AI 版本更新 79%

Correcting What You Cannot See: Credit Assignment for Perception Distillation in Multimodal Reasoners

修正不可见之缺陷:多模态推理器中感知蒸馏的信用分配

Feng Xiong, Leyan Xue, Hongyu Lin

机构 * Thinking Machines Lab(思维机器实验室) DeepSeek-AI(深度求索公司)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.AI

AI总结 该研究针对多模态推理器的感知蒸馏信用分配问题,提出感知修正蒸馏(PCD)方法,经8个基准测试,可提升不同规模模型的宏平均和结果,消融实验验证其有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.13892 2026-08-14 cs.CL cs.IR 79%

Like Article, Like Audience: Enforcing Multimodal Correlations for Disinformation Detection

Liesbeth Allein, Marie-Francine Moens, Domenico Perrotta

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CL

Journal ref Allein, Liesbeth, Marie-Francine Moens, and Domenico Perrotta. "Preventing profiling for ethical fake news detection." Information Processing & Management 60.2 (2023): 103206

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.07759 2026-08-11 eess.SP cs.AI 新提交 79%

CFD-Guided Detection of Concept Drift in Multimodal Physiologic Signals

CFD引导的多模态生理信号概念漂移检测

Farouk Ganiyu Adewumi, Timothy Oladunni, Rochak Ghimire, Kosisochukwu Ogbuanya, Sanaa Reeves, Sandy Akoy

专题命中 其他多模态 :multimodal(title);cross-modal(abstract);分类 cs.AI

AI总结 本研究提出PECS生理稳定性框架,将模型内部变化与信号可测量变化对比,在多生理信号数据集上验证其性能优于基线,可用于可穿戴心血管AI的概念漂移监测。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.28652 2026-08-03 q-bio.OT cs.AI stat.AP 新提交 79%

HenTwin: A Multimodal Digital Twin Framework for Longitudinal Biological State Monitoring in Laying Hens

HenTwin:用于蛋鸡纵向生物状态监测的多模态数字孪生框架

Yashan Dhaliwal, Shreya Rao, Suresh Neethirajan

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.AI

AI总结 该研究提出HenTwin多模态数字孪生框架,通过五层物联网架构整合多模态数据构建蛋鸡生物状态模型,验证了其稳定性与跨房间适用性,为精准畜牧养殖提供了状态感知的数字孪生解决方案。

Comments 24 pages, 18 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.27286 2026-07-31 eess.IV cs.CV 新提交 79%

Toward Multi-Modal Deep Learning for Pulmonary Disease Classification: A Texture-Based Machine Learning Pilot Study on Public Chest X-Ray Data

面向肺部疾病分类的多模态深度学习:基于公开胸部X线数据的纹理机器学习试点研究

Yogisri Pujitha Chinthoti

专题命中 其他多模态 :multi-modal(title,abstract);分类 cs.CV

AI总结 本研究通过公开胸部X线数据开展试点,用HOG、GLCM等特征结合经典分类器区分COVID-19与其他肺炎,为未来多模态深度学习架构研究提供方向。

Comments 6 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.26909 2026-07-30 cs.CL 新提交 79%

Dual-Path LLM Reasoning for Multimodal Few-Shot Knowledge Graph Completion

用于多模态少样本知识图谱补全的双路径大语言模型推理

Jinlan Liu, Zhiying Tu, Yongchao Xing, Yicheng Liu, Bolin Zhang, Dianbo Sui, Dianhui Chu, Hongliang Sun

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CL

AI总结 该研究针对多模态少样本知识图谱补全难题,提出双路径大语言模型推理框架DuPLeR,结合多模态LLM先验与事实支撑构建校准关系图,经实验验证其在数据稀缺场景下性能稳健。

Comments 10 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.25021 2026-07-29 cs.AI cs.HC cs.MA cs.SE 新提交 79%

Chart-Supported or Model-Supplied? Examining MLLM-Generated Claims for Accessible Visualization

图表支持还是模型提供?审视多模态大语言模型生成的可访问可视化声明

Ishrat Jahan Eliza, Md Dilshadur Rahman

专题命中 其他多模态 :MLLM(title);multimodal(abstract);分类 cs.AI

AI总结 研究多模态大语言模型生成可视化声明的证据基础,通过对多种来源、模型及输入条件的探索性研究,分析模型标签与数值一致性,发现可访问图表上下文有作用,添加图像及无上下文提示效果不佳,推动可区分证据支持声明与模型解释的描述系统发展。

Comments Submitted to the 3rd Workshop on Accessible Data Visualization, IEEE VIS 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.24539 2026-07-28 cs.AI cs.SY eess.SY 新提交 79%

Task-Conditional Faithfulness Auditing of Multimodal LLMs for Grid Diagnosis

用于电网诊断的多模态大语言模型的任务条件忠实性审计

Tianqiao Zhao, Meng Yue, Jianhui Wang

机构 * The University of Texas at Arlington(德克萨斯大学阿灵顿分校) Brookhaven National Laboratory(布鲁克海文国家实验室) Southern Methodist University(南卫理公会大学)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.AI

AI总结 研究针对多模态大语言模型用于电网诊断时答案准确性与证据使用的问题,提出通用框架进行任务条件忠实性审计,通过比较多种依赖并设计校正和重新审计机制,经案例研究验证了框架检测、诊断及纠正相关失败的能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.23700 2026-07-28 cs.AI 新提交 79%

Offline-Online Curriculum RL for Multimodal Reasoning

用于多模态推理的离线-在线课程强化学习

Wendi Deng, Hang Du, Guoshun Nan, Haokun Tian, Jiaqi Yu, Xinlei Cao, Jaile Li, Jingfeng Chen, Ling Deng, Ting Li, Hao Yang, Jun Liu, Xudong Jiang, Sicong Leng

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) BUPT Shenzhen Institute(北京邮电大学深圳研究院) Carnegie Mellon University(卡内基梅隆大学) Lancaster University(兰卡斯特大学) Nanyang Technological University(南洋理工大学) China Telecom Corporation Limited Sichuan Branch(中国电信股份有限公司四川分公司) Changsha University of Science & Technology(长沙理工大学)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.AI

AI总结 研究多模态推理中模型中间步骤有缺陷的问题,提出$O^2$-CritiCuRL框架,通过离线多步展开分析和在线渐进式强化学习策略,区分关键与冗余步骤,提升模型性能及训练和推理效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.07632 2026-07-28 cs.LG cs.AI 版本更新 79%

Sheaf-Laplacian Obstruction and Projection Hardness for Cross-Modal Compatibility on a Modality-Independent Site

sheaf-Laplacian 障碍与投影难度分析跨模态兼容性在模态无关站点

Tibor Sloboda

机构 * Faculty of Informatics and Information Technology, Slovak Technical University(斯洛伐克技术大学信息学与信息技术学院) aleph0 s. r. o.

专题命中 其他多模态 :cross-modal(title,abstract);分类 cs.AI

AI总结 本文提出统一框架分析学习表示中的跨模态兼容性,通过模态无关邻域站点和有限维实内积空间细胞sheaf,定义投影难度和sheaf-Laplacian障碍,分析两种失败模式并推导稳定性与误差界。

Comments 31 pages, 7 figures, submitted to Annals of Mathematics and Artificial Intelligence of Springer Nature, post-major revision

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.21085 2026-07-24 cs.CV 新提交 79%

Geo3R: Mitigating Spatial Reasoning Hallucination in Multimodal Large Language Models

Geo3R:减轻多模态大语言模型中的空间推理幻觉

Mingyu Wang, Weilin Jin, Wenbo Li, Haoyang Huang, Tong Jia, Ying Li

机构 * Peking University(北京大学) Joy Future Academy(京东探索研究院)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 多模态大语言模型推理空间关系易产生幻觉,现有方法效果有限。本文定义空间推理幻觉及常见场景,提出无训练的Geo3R框架,结合几何证据和结构化3D推理减轻幻觉,实验证明该框架显著减少幻觉,优于现有模型和方法。

Comments Accepted by ACM MM 2026. This is the arXiv preprint version, not the camera-ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.18153 2026-07-21 cs.CV 新提交 79%

Robust Multimodal Dynamic Object Segmentation

鲁棒多模态动态目标分割

Zhe Xin, Hanzhi Chang, Penghui Huang, Yinian Mao, Guoquan Huang

机构 * Meituan UAV(美团无人机) University of Delaware(特拉华大学)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 研究针对动态目标分割难题,提出整合多模态线索的框架,设计结合Transformer与特征聚类模块的网络进行分类,引入新后处理方法,在动态目标分割和静态场景重建任务中取得最优性能。

Comments Accepted by IEEE International Conference on Robotics & Automation ICRA 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.17262 2026-07-21 cs.CL 新提交 79%

Should Missing Modalities Always Be Necessary to Repair for Multi-modal Sentiment Analysis?

多模态情感分析中缺失模态总是需要修复吗?

Yubo Gao, Haotian Wu, Xiaoyu Xu, Yibo Yan, Hong Chen, Ruoshui Peng, Fei Pan, Puay Siew Tan, Zhuoran Gao, Yonghua Hei, Jie Zhang, Xuming Hu

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) The Hong Kong University of Science and Technology(香港科技大学) Nanyang Technological University(南洋理工大学) Singapore Institute of Manufacturing Technology, A*STAR(新加坡制造技术研究所,新加坡科技研究局) Lingnan University(岭南大学)

专题命中 其他多模态 :multi-modal(title);multimodal(abstract);分类 cs.CL

AI总结 研究多模态情感分析中缺失模态是否总需修复,提出SIEVE方法,通过比较直接预测与修复分支,从样本损失差距得经验充分性信号,经证据门路由输入,与修复无关,实验证明其能改进修复主干并接近最优值。

详情

展开后加载摘要…

URL PDF HTML 收藏