arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-12-04 至 2025-12-04 共收录 5 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 5 篇

2502.00530 2025-12-04 cs.LG cs.AI cs.SI 79%

Generic Multimodal Spatially Graph Network for Spatially Embedded Network Representation Learning

通用多模态空间图网络用于空间嵌入网络表示学习

Xudong Fan, Jürgen Hackl

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

AI总结 本文提出通用多模态空间图卷积网络,通过多模态特征提升空间嵌入网络的表示准确性,实验显示在电力网络中边存在预测任务的准确率提高37.1%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03404 2025-12-04 cs.CV 79%

MOS: Mitigating Optical-SAR Modality Gap for Cross-Modal Ship Re-Identification

MOS:缓解光学-合成孔径雷达模态差距以实现跨模态舰船重识别

Yujian Zhao, Hankun Liu, Guanglin Niu

机构 * School of Artificial Intelligence, Beihang University(北京航空航天大学人工智能学院) School of Computer Science and Engineering, Beihang University(北京航空航天大学计算机科学与工程学院)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

AI总结 MOS通过模态一致表示学习和跨模态数据生成与融合,有效缓解光学与SAR图像间的模态差距,提升舰船跨模态重识别的准确率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.01513 2025-12-04 cs.CR cs.CV 79%

SafePTR: Token-Level Jailbreak Defense in Multimodal LLMs via Prune-then-Restore Mechanism

SafePTR: 通过剪枝-恢复机制实现多模态大语言模型的令牌级 Jailbreak 防御

Beitao Chen, Xinyu Lyu, Lianli Gao, Jingkuan Song, Heng Tao Shen

机构 * Shenzhen Institute for Advanced Study, University of Electronic Science and Technology of China(电子科技大学深圳研究院) Southwestern University of Finance and Economics(西南财经大学) Engineering Research Center of Intelligent Finance, Ministry of Education(教育部智能金融工程研究中心) Tongji University(同济大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 SafePTR 提出一种无需训练的多模态大语言模型防御机制,通过剪枝有害令牌并恢复良性特征,有效提升安全性并保持效率。

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.04005 2025-12-04 cs.CV cs.LG cs.RO 70%

LargeAD: Large-Scale Cross-Sensor Data Pretraining for Autonomous Driving

LargeAD: 大规模跨传感器数据预训练用于自动驾驶

Lingdong Kong, Xiang Xu, Youquan Liu, Jun Cen, Runnan Chen, Wenwei Zhang, Liang Pan, Kai Chen, Ziwei Liu

机构 * WorldBench Team(WorldBench团队)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 LargeAD通过跨传感器数据预训练提升自动驾驶中的三维场景理解,结合多模态对比学习和时间一致性,实现更鲁棒的感知性能。

Comments IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03747 2025-12-04 cs.CV 57%

LoRA Patching: Exposing the Fragility of Proactive Defenses against Deepfakes

LoRA Patching: 暴露对抗深度伪造的主动防御的脆弱性

Zuomin Qu, Yimao Guo, Qianyue Hu, Wei Lu

机构 * School of Computer Science and Engineering, Sun Yat-sen University(中山大学计算机科学与工程学院) Electric Power Research Institute, China Southern Power Grid Company Ltd.(中国南方电网有限责任公司电力科学研究院)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

AI总结 LoRA Patching通过注入LoRA修补块绕过深度伪造防御,并引入MMFA损失提升对抗性输出特征对齐,揭示现有防御的脆弱性。

详情

展开后加载摘要…

URL PDF HTML 收藏