arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-10-01 至 2025-10-01 共收录 12 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 12 篇

2501.17391 2025-10-01 cs.CV cs.AI cs.CL 85%

LFTR: Learning-Free Token Reduction for Multimodal Large Language Models

Zihui Zhao, Yingxin Li, Yang Li

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04943 2025-10-01 cs.CV cs.CL 84%

ReLoop: "Seeing Twice and Thinking Backwards" via Closed-loop Training to Mitigate Hallucinations in Multimodal understanding

Jianjiang Yang, Yanshu li, Ziyan Huang

机构 * University of Bristol(布里斯托大学) Brown University(布朗大学) South China University of Technology(华南理工大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.CL

Comments Accepted by conference EMNLP2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25278 2025-10-01 cs.LG stat.ML 82%

MAESTRO : Adaptive Sparse Attention and Robust Learning for Multimodal Dynamic Time Series

Payal Mohapatra, Yueyuan Sui, Akash Pandey, Stephen Xia, Qi Zhu

机构 * Northwestern University(西北大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract)

Comments Accepted to Neurips 2025 (Spotlight)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12299 2025-10-01 cs.CR cs.AI 79%

QGuard:Question-based Zero-shot Guard for Multi-modal LLM Safety

Taegyeong Lee, Jeonghwa Yoo, Hyoungseo Cho, Soo Yong Kim, Yunho Maeng

机构 * FnGuide Inc.(FnGuide公司) Safe Generative AI Lab, MODULABS(MODULABS安全生成AI实验室) A.I.MATICS Inc.(A.I.MATICS公司) Ewha Womans University(成均馆大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.AI

Comments Accept to ACLW 2025 (WOAH); fix typo

Journal ref ACL Workshop 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25525 2025-10-01 cs.CR cs.LG 78%

Defeating Cerberus: Concept-Guided Privacy-Leakage Mitigation in Multimodal Language Models

Boyang Zhang, Istemi Ekin Akkus, Ruichuan Chen, Alice Dethise, Klaus Satzke, Ivica Rimac, Yang Zhang

机构 * CISPA Helmholtz Center for Information Security(CISPA赫尔姆霍茨信息安全中心) Nokia Bell Labs(诺基亚贝尔实验室)

专题命中 多模态训练与对齐 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26625 2025-10-01 cs.LG cs.AI cs.CV cs.MM 75%

Learning to See Before Seeing: Demystifying LLM Visual Priors from Language Pre-training

Junlin Han, Shengbang Tong, David Fan, Yufan Ren, Koustuv Sinha, Philip Torr, Filippos Kokkinos

机构 * Meta Superintelligence Labs(Meta 超智能实验室) University of Oxford(牛津大学)

专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.AI、cs.MM

Comments Project page: https://junlinhan.github.io/projects/lsbs/

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26223 2025-10-01 q-bio.GN 71%

Nephrobase Cell+: Multimodal Single-Cell Foundation Model for Decoding Kidney Biology

Chenyu Li, Elias Ziyadeh, Yash Sharma, Bernhard Dumoulin, Jonathan Levinsohn, Eunji Ha, Siyu Pan, Vishwanatha Rao, Madhav Subramaniyam, Mario Szegedy, Nancy Zhang, Katalin Susztak

专题命中 多模态训练与对齐 :multimodal(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26016 2025-10-01 cs.CV 70%

GeoLink: Empowering Remote Sensing Foundation Model with OpenStreetMap Data

Lubian Bai, Xiuyuan Zhang, Siqi Zhang, Zepeng Zhang, Haoyu Wang, Wei Qin, Shihong Du

机构 * School of Earth and Space Sciences, Peking University(地球与空间科学学院,北京大学) College of Urban and Environmental Sciences, Peking University(城市与环境科学学院,北京大学) State Key Laboratory of Multimodal Artificial Intelligence Systems Institute of Automation, CAS(多模态人工智能系统国家重点实验室,中国科学院自动化研究所) Intelligent Maintenance and Operations Systems Lab(智能维护与运营系统实验室)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25712 2025-10-01 cs.LG 67%

Expert Merging: Model Merging with Unsupervised Expert Alignment and Importance-Guided Layer Chunking

Dengming Zhang, Xiaowen Ma, Zhenliang Ni, Zhenkai Wu, Han Shu, Xin Jiang, Xinghao Chen

机构 * Zhejiang University(浙江大学) Huawei Noah’s Ark Lab(华为诺亚实验室)

专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.05728 2025-10-01 cs.CV cs.AI cs.LG cs.RO 62%

LiDAR-BIND-T: Improved and Temporally Consistent Sensor Modality Translation and Fusion for Robotic Applications

Niels Balemans, Ali Anwar, Jan Steckel, Siegfried Mercelis

机构 * IDLab - Faculty of Applied Engineering, University of Antwerp - imec(IDLab - 应用工程学院,安特卫普大学 - imec) Cosys-Lab - Faculty of Applied Engineering, University of Antwerp(Cosys-Lab - 应用工程学院,安特卫普大学) Flanders Make Strategic Research Centre(弗拉芒制造战略研究中心)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26039 2025-10-01 cs.CV 57%

SGS: Segmentation-Guided Scoring for Global Scene Inconsistencies

Gagandeep Singh, Samudi Amarsinghe, Urawee Thani, Ki Fung Wong, Priyanka Singh, Xue Li

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments 6 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.12737 2025-10-01 cs.CV 57%

PolSAM: Polarimetric Scattering Mechanism Informed Segment Anything Model

Yuqing Wang, Zhongling Huang, Shuxin Yang, Hao Tang, Xiaolan Qiu, Junwei Han, Dingwen Zhang

机构 * the BRain and Artificial INtelligence Lab (BRAIN LAB), School of Automation, Northwestern Polytechnical University(人工智能实验室(BRAIN LAB)、自动化学院、西北工业大学) the Shenzhen Research Institute of Northwestern Polytechnical University(西北工业大学深圳研究院) the National Key Laboratory for Multimedia Information Processing, School of Computer Science, Peking University(多媒体信息处理国家重点实验室、计算机科学学院、北京大学) the National Key Laboratory of Microwave Imaging Technology, Chinese Academy of Sciences(微波成像技术国家重点实验室、中国科学院) the Aerospace Information Research Institute, Chinese Academy of Sciences(航空信息研究所、中国科学院)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments The manuscript is 15 pages long, includes 13 figures and 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏