arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-11-24 至 2025-11-24 共收录 43 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 8 篇

2511.16853 2025-11-24 cs.CV 57%

Towards Unified Vision Language Models for Forest Ecological Analysis in Earth Observation

面向地球观测森林生态分析的统一视觉语言模型

Xizhe Xue, Xiao Xiang Zhu

机构 * Xizhe Xue 1, 2 Xiao Xiang Zhu 1, 2

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

AI总结 REO-Instruct提出首个面向地球观测森林生态分析的统一视觉语言模型基准,旨在解决描述与回归任务中的多模态感知与生物物理变量对齐问题。

Comments AAAI2026 AI for Environmental Science Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04079 2025-11-24 cs.CL 57%

Improving the Performance of Radiology Report De-identification with Large-Scale Training and Benchmarking Against Cloud Vendor Methods

通过大规模训练和与云服务提供商方法的基准测试来改进放射学报告去标识化性能

Eva Prakash, Maayane Attias, Pierre Chambon, Justin Xu, Steven Truong, Jean-Benoit Delbrouck, Tessa Cook, Curtis Langlotz

机构 * Stanford University(斯坦福大学) JP Morgan Chase & Co(摩根大通公司) Sorbonne University(索邦大学) University of Oxford(牛津大学) NVIDIA(英伟达) HOPPR University of Pennsylvania(宾夕法尼亚大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CL

AI总结 本文提出了一种基于变压器的去标识化模型,通过大规模训练和与商业系统的基准测试,实现了在放射学报告中更高效的PHI检测性能。

Comments In submission to JAMIA

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 多模态Agent 3 篇

2511.17335 2025-11-24 cs.RO cs.CL cs.CV cs.SD eess.AS 82%

Robot Confirmation Generation and Action Planning Using Long-context Q-Former Integrated with Multimodal LLM

基于长上下文Q-Former与多模态大语言模型的机器人确认生成与动作规划

Chiori Hori, Yoshiki Masuyama, Siddarth Jain, Radu Corcodel, Devesh Jha, Diego Romeres, Jonathan Le Roux

机构 * Mitsubishi Electric Research Laboratories (MERL), Cambridge, MA, USA(三菱电机研究实验室(MERL),马萨诸塞州剑桥市)

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV、cs.CL、eess.AS

AI总结 本文提出基于长上下文Q-former与多模态大语言模型的机器人确认生成与动作规划方法,通过整合视频上下文信息提升动作规划性能。

Comments Accepted to ASRU 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21651 2025-11-24 cs.AI 57%

Can AI Perceive Physical Danger and Intervene?

AI能否感知物理危险并干预?

Abhishek Jindal, Dmitry Kalashnikov, R. Alex Hofer, Oscar Chang, Divya Garikapati, Anirudha Majumdar, Pierre Sermanet, Vikas Sindhwani

机构 * Google DeepMind Robotics(谷歌深Mind机器人技术)

专题命中 多模态Agent :multi-modal(abstract);分类 cs.AI

AI总结 本文提出了一种用于评估具身体验AI系统物理安全性的基准测试方法,通过生成逼真图像和视频来测试模型对安全风险的理解和干预能力,并开发了训练后范式以提升模型的安全推理能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.23463 2025-11-24 cs.CV 57%

OpenDriveVLA: Towards End-to-end Autonomous Driving with Large Vision Language Action Model

OpenDriveVLA: 向端到端自动驾驶迈进的大型视觉语言动作模型

Xingcheng Zhou, Xuyuan Han, Feng Yang, Yunpu Ma, Volker Tresp, Alois Knoll

机构 * Technical University of Munich(慕尼黑技术大学) Ludwig Maximilian University of Munich(慕尼黑路德维希-马克西米利安大学)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV

AI总结 OpenDriveVLA基于开源大语言模型,通过多模态输入和分层视觉语言对齐,实现端到端自动驾驶中的高精度轨迹规划和驾驶任务回答。

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 多模态训练与对齐 6 篇

2511.16990 2025-11-24 cs.HC 82%

Senti-iFusion: An Integrity-centered Hierarchical Fusion Framework for Multimodal Sentiment Analysis under Uncertain Modality Missingness

Senti-iFusion: 一种以完整性为中心的多模态情感分析多模态融合框架,用于在不确定模态缺失情况下

Liling Li, Guoyang Xu, Xiongri Shen, Zhifei Xu, Yanbo Zhang, Zhiguo Zhang, Zhenxi Song

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract)

AI总结 Senti-iFusion提出了一种以完整性为中心的多模态融合框架,通过分层结构处理模态缺失问题,提升多模态情感分析的鲁棒性和准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17308 2025-11-24 cs.CV 79%

SpatialGeo:Boosting Spatial Reasoning in Multimodal LLMs via Geometry-Semantics Fusion

SpatialGeo:通过几何-语义融合提升多模态大语言模型的空间推理能力

Jiajie Guo, Qingpeng Zhu, Jin Zeng, Xiaolong Wu, Changyong He, Weida Wang

机构 * School of Computer Science and Technology, Tongji University, Shanghai, China(计算机科学与技术学院,同济大学,上海,中国)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 SpatialGeo通过几何-语义融合提升多模态大语言模型的空间推理能力,实验表明在空间推理任务中准确率提升8.0%且内存消耗减少50%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16816 2025-11-24 stat.AP stat.CO 71%

Trust-Aware Multimodal Data Fusion for Yield Estimation: A Case Study of the 2020 Beirut Explosion

具有信任意识的多模态数据融合用于产量估计:贝鲁特2020爆炸的案例研究

Lekha Patel, Craig Ulmer, Stephen J. Verzi, Daniel J. Krofcheck, Indu Manickam, Asmeret Naugle, Jaideep Ray

专题命中 多模态训练与对齐 :multimodal(title)

AI总结 本文提出一种基于贝叶斯分数后验框架的多模态数据融合方法,用于估计爆炸产量,通过信任权重校准不同观测数据,提升不确定性量化和抗偏差能力。

Comments 19 pages, 4 figures, supplementary material, journal

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16991 2025-11-24 cs.CV 70%

DReX: Pure Vision Fusion of Self-Supervised and Convolutional Representations for Image Complexity Prediction

DReX:纯视觉融合自监督和卷积表示以预测图像复杂度

Jonathan Skaza, Parsa Madinei, Ziqi Wen, Miguel Eckstein

机构 * University of California, Santa Barbara(加州大学圣芭芭拉分校)

专题命中 多模态训练与对齐 :multimodal(abstract);image-text(abstract);分类 cs.CV

AI总结 DReX通过融合自监督和卷积表示,实现图像复杂度预测,取得最佳性能并减少参数量。

Comments 8 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17358 2025-11-24 cs.CL 57%

Don't Learn, Ground: A Case for Natural Language Inference with Visual Grounding

不要学习,而是依托:自然语言推理与视觉依托的案例

Daniil Ignatev, Ayman Santeer, Albert Gatt, Denis Paperno

机构 * Utrecht University(乌特勒支大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CL

AI总结 本文提出了一种基于视觉依托的零样本自然语言推理方法,通过生成视觉表示并比较与假设的相似度,实现高精度推理,展示了对文本偏见的鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10518 2025-11-24 cs.LG 50%

Holographic Knowledge Manifolds: A Novel Pipeline for Continual Learning Without Catastrophic Forgetting in Large Language Models

全息知识流形:一种实现大语言模型持续学习中无灾难性遗忘的新型流程

Justin Arndt

专题命中 多模态训练与对齐 :multimodal(abstract)

AI总结 全息知识流形通过分形量化等技术实现大语言模型持续学习中零灾难性遗忘,提升知识压缩效率并降低训练成本。

Comments This paper includes significant errors discovered post publication by the author

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 其他多模态 2 篇

2511.17455 2025-11-24 cs.CV 74%

Improving Multimodal Distillation for 3D Semantic Segmentation under Domain Shift

改进多模态蒸馏以应对激光雷达语义分割中的域移位

Björn Michele, Alexandre Boulch, Gilles Puy, Tuan-Hung Vu, Renaud Marlet, Nicolas Courty

机构 * CNRS, IRISA, Univ. Bretagne Sud(CNRS、IRISA、布列塔尼大学) LIGM, Ecole des Ponts, Univ Gustave Eiffel, CNRS(LIGM、巴黎理工学院、古斯塔夫·埃菲尔大学、CNRS)

专题命中 其他多模态 :multimodal(title);分类 cs.CV

AI总结 本研究提出改进多模态蒸馏方法,通过冻结预训练主干网络并训练MLP头,提升激光雷达语义分割在域移位下的性能。

Comments Accepted at BMVC 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17158 2025-11-24 physics.med-ph cs.CV 57%

Exploring the added value of pretherapeutic MR descriptors in predicting breast cancer pathologic complete response to neoadjuvant chemotherapy

探讨术前MRI描述符在预测乳腺癌新辅助化疗病理完全缓解中的附加价值

Caroline Malhaire, Fatine Selhane, Marie-Judith Saint-Martin, Vincent Cockenpot, Pia Akl, Enora Laas, Audrey Bellesoeur, Catherine Ala Eddine, Melodie Bereby-Kahane, Julie Manceau, Delphine Sebbag-Sfez, Jean-Yves Pierga, Fabien Reyal, Anne Vincent-Salomon, Herve Brisse, Frederique Frouin

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV

AI总结 本研究探讨术前MRI特征在预测乳腺癌新辅助化疗病理完全缓解中的作用,发现非分叶边缘和单发性是独立预测因素,可提高预测模型性能。

Journal ref European Radiology, 2023, 33 (11), pp.8142-8154

详情

展开后加载摘要…

URL PDF HTML 收藏