arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-11-06 至 2025-11-06 共收录 37 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6 篇

2409.15132 2025-11-06 cs.CV eess.IV 57%

FusionRF: High-Fidelity Satellite Neural Radiance Fields from Multispectral and Panchromatic Acquisitions

Michael Sprintson, Rama Chellappa, Cheng Peng

机构 * Johns Hopkins University(约翰霍普金斯大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03488 2025-11-06 cs.LG 50%

NAP: Attention-Based Late Fusion for Automatic Sleep Staging

Alvise Dei Rossi, Julia van der Meer, Markus H. Schmidt, Claudio L. A. Bassetti, Luigi Fiorillo, Francesca Faraci

机构 * Faculty of Informatics, Università della Svizzera Italiana(信息学院,瑞士意大利大学) Inst. of Digital Tech. for Personalized Healthcare, SUPSI(个性化医疗数字技术研究所,SUPSI) Sleep Wake Epilepsy Center, Department of Neurology, Inselspital University of Bern(睡眠觉醒癫痫中心,神经科,伯恩大学医院) Faculty of Informatics Università della Svizzera Italiana(信息学院,瑞士意大利大学)

专题命中 多模态训练与对齐 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 其他多模态 5 篇

2511.03328 2025-11-06 cs.CL cs.AI cs.CV cs.LG 82%

Benchmarking the Thinking Mode of Multimodal Large Language Models in Clinical Tasks

Jindong Hong, Tianjie Chen, Lingjie Luo, Chuanyang Zheng, Ting Xu, Haibao Yu, Jianing Qiu, Qianzhong Chen, Suning Huang, Yan Xu, Yong Gui, Yijun He, Jiankai Sun

机构 * Bytedance(字节跳动) Peking University(北京大学) The Chinese University of Hong Kong(香港中文大学) The University of Hong Kong(香港大学) Mohamed bin Zayed University of Artificial Intelligence(马尔代夫人工智能大学) Stanford University(斯坦福大学) University of Michigan(密歇根大学)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03206 2025-11-06 cs.CV cs.AI cs.LG 81%

QG-CoC: Question-Guided Chain-of-Captions for Large Multimodal Models

Kuei-Chun Kao, Hsu Tzu-Yin, Yunqi Hong, Ruochen Wang, Cho-Jui Hsieh

机构 * Department of Computer Science, University of California, Los Angeles(计算机科学系,加州大学洛杉矶分校)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 16 pages

Journal ref EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03617 2025-11-06 cs.GR cs.AI 79%

Visualization Biases MLLM's Decision Making in Network Data Tasks

Timo Brand, Henry Förster, Stephen G. Kobourov, Jacob Miller

机构 * Technical University of Munich, Heilbronn, Germany(慕尼黑技术大学)

专题命中 其他多模态 :MLLM(title,abstract);分类 cs.AI

Comments This manuscript was presented at VIS x GenAI, a workshop co-located with IEEE VIS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03478 2025-11-06 cs.HC 78%

SVG Decomposition for Enhancing Large Multimodal Models Visualization Comprehension: A Study with Floor Plans

Jeongah Lee, Ali Sarvghad

专题命中 其他多模态 :multimodal(title,abstract)

Comments 10 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.13765 2025-11-06 cs.HC cs.CL 57%

SciDaSynth: Interactive Structured Data Extraction from Scientific Literature with Large Language Model

Xingbo Wang, Samantha L. Huey, Rui Sheng, Saurabh Mehta, Fei Wang

专题命中 其他多模态 :multimodal(abstract);分类 cs.CL

Comments Preprint version of the paper accepted to Campbell Systematic Reviews. Code is available at https://github.com/xingbow/SciDaEx

Journal ref Campbell Systematic Reviews 21 (2025): 1-16

详情

展开后加载摘要…

URL PDF HTML 收藏