arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6929 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6929 篇

2511.21582 2026-01-06 cs.CV 79%

Data-Augmented Multimodal Feature Fusion for Multiclass Visual Recognition of Oral Cancer Lesions

数据增强多模态特征融合用于口腔癌病变的多类视觉识别

Joy Naoum, Revana Salama, Ali Hamdi

机构 * MSA University(MSA大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出一种数据增强驱动的多模态特征融合框架,用于提升口腔癌病变的多类视觉识别性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.00273 2026-01-05 eess.IV cs.CV 79%

UKAN-EP: Enhancing U-KAN with Efficient Attention and Pyramid Aggregation for 3D Multi-Modal MRI Brain Tumor Segmentation

UKAN-EP: 通过高效的注意力和金字塔聚合增强U-KAN以实现3D多模态MRI脑肿瘤分割

Yanbing Chen, Tianze Tang, Taehyo Kim, Hai Shu

机构 * Department of Biostatistics, School of Global Public Health, New York University(生物统计学系,全球公共卫生学院,纽约大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

AI总结 UKAN-EP通过引入高效的通道注意力和金字塔特征聚合模块,提升3D多模态MRI脑肿瘤分割的准确性和效率。

Journal ref BMC Medical Imaging, Volume 25, article number 517, (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22748 2026-01-01 cs.CV 79%

TrimTokenator-LC: Towards Adaptive Visual Token Pruning for Large Multimodal Models with Long Contexts

TrimTokenator-LC: 向大型多模态模型长上下文的自适应视觉标记修剪迈进

Hao Zhang, Mengsi Lyu, Bo Huang, Yulong Ao, Yonghua Lin

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 TrimTokenator-LC通过自适应视觉标记修剪方法,在长上下文和多图像场景中有效减少视觉标记数量,同时保持性能。

Comments 17 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22503 2025-12-30 cs.CV 79%

SCAFusion: A Multimodal 3D Detection Framework for Small Object Detection in Lunar Surface Exploration

SCAFusion: 一种针对月球表面探索的小目标多模态3D检测框架

Xin Chen, Kang Luo, Yangyi Xiao, Hesheng Wang

机构 * Department of Automation, Key Laboratory of System Control and Information Processing of Ministry of Education, Key Laboratory of Marine Intelligent Equipment and System of Ministry of Education, Shanghai Engineering Research Center of Intelligent Control and Management, Shanghai Jiao Tong University(自动化系、教育部系统控制与信息处理重点实验室、教育部海洋智能装备与系统重点实验室、上海智能控制与管理工程研究中心、上海交通大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 SCAFusion提出了一种针对月球表面探索的小目标多模态3D检测框架,通过改进的特征对齐和坐标注意机制提升小目标检测性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21916 2025-12-29 cs.CV 79%

Patch as Node: Human-Centric Graph Representation Learning for Multimodal Action Recognition

补丁作为节点:面向多模态动作识别的人本图表示学习

Zeyu Liang, Hailun Xia, Naichuan Zheng

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出PAN框架,通过人本图表示学习提升多模态动作识别性能,结合双路径和统一网络结构实现高效融合。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21897 2025-12-29 cs.LG cs.AI 79%

MMCTOP: A Multimodal Textualization and Mixture-of-Experts Framework for Clinical Trial Outcome Prediction

MMCTOP: 一种用于临床试验结果预测的多模态文本化与专家混合框架

Carolina Aparício, Qi Shi, Bo Wen, Tesfaye Yadete, Qiwei Han

机构 * Nova School of Business and Economics(诺瓦商学院) Hogarthian Technologies(霍加斯技术) School of Medicine(医学院) Oregon Health & Science University(俄勒冈健康与科学大学) Cleveland Clinic(克利夫兰诊所) IBM Research(IBM研究院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

AI总结 MMCTOP通过多模态文本化与专家混合框架,提升临床试验结果预测的精度与稳定性。

Comments 15 pages, 3 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.19443 2025-12-29 cs.CV 79%

D2Pruner: Debiased Importance and Structural Diversity for MLLM Token Pruning

D2Pruner: 用于MLLM标记剪枝的去偏重要与结构多样性

Evelyn Zhang, Fufu Yu, Aoqi Wu, Zichen Wen, Ke Yan, Shouhong Ding, Biqing Qi, Linfeng Zhang

机构 * Tencent YouTu Lab(腾讯YouTu实验室)

专题命中 多模态训练与对齐 :MLLM(title);multimodal(abstract);分类 cs.CV

AI总结 D2Pruner通过结合去偏重要与结构剪枝机制,有效提升MLLM标记剪枝的效率和保真度,尤其在细粒度定位任务中表现突出。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.09626 2025-12-29 cs.SI cs.AI cs.LG 79%

Certainly Bot Or Not? Trustworthy Social Bot Detection via Robust Multi-Modal Neural Processes

确定是机器人还是不是?通过鲁棒多模态神经过程进行可信的社交机器人检测

Qi Wu, Yingguang Yang, hao liu, Hao Peng, Buyun He, Yutong Xia, Yong Liao

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.AI

AI总结 本研究提出鲁棒多模态神经过程框架,通过增强多模态神经过程的鲁棒性来检测社交机器人,同时提升不确定性估计能力。

Comments We withdraw this paper due to an error identified in the experimental setup. Specifically, the evaluation protocol described in Section 4 does not correctly reflect the intended experimental design, which may affect the validity of the reported results. To avoid potential misunderstanding by readers, we choose to withdraw this version and revise the work before resubmission

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20146 2025-12-29 cs.CV 79%

AlignFreeNet: Is Cross-Modal Pre-Alignment Necessary? An End-to-End Alignment-Free Lightweight Network for Visible-Infrared Object Detection

AlignFreeNet: 跨模态预对齐是否必要?一种端到端无对齐的轻量级网络用于可见-红外目标检测

Dingkun Zhu, Haote Zhang, Lipeng Gu, Wuzhou Quan, Fu Lee Wang, Honghui Fan, Jiali Tang, Haoran Xie, Xiaoping Zhang, Mingqiang Wei

机构 * School of Computer Science, Jiangsu University of Technology(江苏科技大学计算机科学学院) School of Computer Science and Technology, Nanjing University of Aeronautics and Astronautics(南京航空航天大学计算机科学与技术学院) School of Science and Technology, Hong Kong Metropolitan University(香港都会大学科技学院) School of Data Science, Lingnan University(岭南大学数据科学学院) Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

AI总结 AlignFreeNet通过无对齐融合范式,提出VCC和FCF模块,有效缓解可见-红外目标检测中的跨模态错位问题,实现端到端轻量级网络的高鲁棒性与泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.20084 2025-12-24 cs.LG cs.AI 79%

QE-Catalytic: A Graph-Language Multimodal Base Model for Relaxed-Energy Prediction in Catalytic Adsorption

QE-Catalytic: 一种图-语言多模态基础模型,用于催化吸附中放松能量的预测

Yanjie Li, Jian Xu, Xueqing Chen, Lina Yu, Shiming Xiang, Weijun Li, Cheng-lin Liu

机构 * AnnLab(安实验室) Institute of Semiconductors, Chinese Academy of Sciences(半导体研究所,中国科学院) Zhongguancun Academy(中关村学院) State Key Laboratory of Multimodal Artificial Intelligence Systems(多模态人工智能系统国家重点实验室) Institute of Automation, Chinese Academy of Sciences(自动化研究所,中国科学院) University of Chinese Academy of Sciences(中国科学院大学) Computer Network Information Center(计算机网络信息中心)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

AI总结 QE-Catalytic结合语言模型与图Transformer,实现高精度催化吸附能量预测及逆向设计

Comments 25 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.20026 2025-12-24 cs.CV 79%

MAPI-GNN: Multi-Activation Plane Interaction Graph Neural Network for Multimodal Medical Diagnosis

MAPI-GNN:多激活平面交互图神经网络用于多模态医学诊断

Ziwei Qin, Xuhui Song, Deqing Huang, Na Qin, Jun Li

机构 * Ziwei Qin(独立研究者) Xuhui Song(独立研究者) Deqing Huang(独立研究者) Na Qin(独立研究者) Jun Li(独立研究者)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 MAPI-GNN通过多激活平面交互机制,有效建模患者特异性病理关系,提升多模态医学诊断的准确性。

Comments Accepted by Proceedings of the AAAI Conference on Artificial Intelligence 40 (AAAI-26)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.19213 2025-12-23 cs.CV 79%

InvCoSS: Inversion-driven Continual Self-supervised Learning in Medical Multi-modal Image Pre-training

InvCoSS: 基于反向驱动的医学多模态图像预训练中的连续自监督学习

Zihao Luo, Shaohao Rui, Zhenyu Tang, Guotai Wang, Xiaosong Wang

机构 * University of Electronic Science and Technology of China(电子科技大学) Shanghai Innovation Institute(上海创新研究院) Shanghai Jiao Tong University(上海交通大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Brain-Computer Interface & Brain-Inspired Intelligence Key Laboratory of Sichuan Province(四川省脑机接口与脑启发智能重点实验室)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

AI总结 InvCoSS通过反向生成合成图像和多尺度融合网络,实现连续自监督学习,减少存储需求并保护数据隐私。

Comments 16 pages, 10 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.14972 2025-12-23 cs.CL 79%

Multimodal Cultural Safety: Evaluation Framework and Alignment Strategies

多模态文化安全:评估框架与对齐策略

Haoyi Qiu, Kung-Hsiang Huang, Ruichen Zheng, Jiao Sun, Nanyun Peng

机构 * University of California, Los Angeles(加州大学洛杉矶分校) Salesforce AI Research(Salesforce人工智能研究) Google DeepMind(谷歌DeepMind)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

AI总结 本文提出CROSS基准和CROSS-Eval框架,评估多模态模型的文化安全能力,发现提升推理能力可改善文化对齐,但需结合监督微调和偏好微调策略以增强文化合规性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11361 2025-12-23 cs.CL 79%

VLDBench Evaluating Multimodal Disinformation with Regulatory Alignment

VLDBench:评估具有监管对齐的多模态虚假信息

Shaina Raza, Ashmal Vayani, Aditya Jain, Aravind Narayanan, Vahid Reza Khazaie, Syed Raza Bashir, Elham Dolatabadi, Gias Uddin, Christos Emmanouilidis, Rizwan Qureshi, Mubarak Shah

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

AI总结 VLDBench 是首个多模态虚假信息检测基准,通过大规模标注数据提升检测准确率,支持 AI 管治框架下的可信虚假信息分析。

Comments Accepted in Information Fusion Journal

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.17227 2025-12-22 cs.CV 79%

Learning When to Look: A Disentangled Curriculum for Strategic Perception in Multimodal Reasoning

学习何时观察:一种解耦的课程学习框架,用于多模态推理中的战略感知

Siqi Yang, Zilve Gao, Haibo Qiu, Fanfan Liu, Peng Shi, Zhixiong Zeng, Qingmin Liao, Lin Ma

机构 * Meituan(美团) Tsinghua University(清华大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出了一种解耦课程学习框架,通过分阶段训练提升多模态推理中的抽象推理与战略视觉感知能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.20494 2025-12-22 cs.LG cs.MM 79%

Multimodal Representation Learning and Fusion

多模态表示学习与融合

Qihang Jin, Enze Ge, Yuhang Xie, Hongying Luo, Junhao Song, Ziqian Bi, Chia Xin Liang, Jibin Guan, Joe Yeong, Xinyuan Song, Junfeng Hao

机构 * AI Agent Lab, Vokram Group, United Kingdom(AI代理实验室,Vokram集团,英国) University of Bologna, Italy(博洛尼亚大学,意大利) University of Minnesota, United States(明尼苏达大学,美国) Singapore General Hospital, Singapore(新加坡中央医院,新加坡) Emory University, United States(埃默里大学,美国)

专题命中 多模态训练与对齐 :multimodal(title);multi-modal(abstract);分类 cs.MM

AI总结 多模态学习通过融合多种数据模态提升AI系统的理解和决策能力,旨在解决数据格式差异、输入缺失及对抗攻击等问题,推动计算机视觉、自然语言处理等领域的发展。

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.11160 2025-12-17 cs.CV 79%

A Unified Framework with Multimodal Fine-tuning for Remote Sensing Semantic Segmentation

多模态微调的统一框架用于遥感语义分割

Xianping Ma, Xiaokang Zhang, Man-On Pun, Bo Huang

机构 * School of Science and Engineering, The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)科学与工程学院) School of Information Science and Engineering, Wuhan University of Science and Technology(武汉科技大学信息科学与工程学院) Department of Geography, The University of Hong Kong(香港大学地理系)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出了一种多模态微调的统一框架,用于改进遥感语义分割的性能,通过引入新的MFNet和DFM模块,显著提升了多模态数据的分割效果。

Comments 15 pages, 11 figures

Journal ref IEEE Transactions on Geoscience and Remote Sensing, vol. 63, pp. 1-15, 2025, Art no. 5405015

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12657 2025-12-16 cs.CV 79%

Cross-modal Fundus Image Registration under Large FoV Disparity

跨模态视网膜图像在大视野差异下的配准

Hongyang Li, Junyi Tao, Qijie Wei, Ningzhi Yang, Meng Wang, Weihong Yu, Xirong Li

机构 * Renmin University of China, Beijing, China(中国人民大学) Peking Union Medical College Hospital, Beijing, China(北京友谊医院)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

AI总结 本文提出CARe方法,用于解决大视野差异下的跨模态视网膜图像配准问题,通过裁剪和双拟合对齐改进配准效果。

Comments Accepted as a regular paper at MMM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.20991 2025-12-16 cs.CV 79%

MR-COSMO: Visual-Text Memory Recall and Direct CrOSs-MOdal Alignment Method for Query-Driven 3D Segmentation

MR-COSMO:一种用于查询驱动3D分割的视觉-文本记忆召回与直接跨模态对齐方法

Chade Li, Pengju Zhang, Yihong Wu

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

AI总结 MR-COSMO通过视觉-文本记忆召回与直接跨模态对齐方法,在查询驱动的3D分割中实现几何与语义特征的精确融合,提升点云分割性能。

Comments Accepted by AAAI 2026. Copyright (c) 2026, Association for the Advancement of Artificial Intelligence (www.aaai.org). All rights reserved

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.04519 2025-12-12 cs.CV 79%

l0-Regularized Sparse Coding-based Interpretable Network for Multi-Modal Image Fusion

基于l0正则化稀疏编码的可解释网络用于多模态图像融合

Gargi Panda, Soumitra Kundu, Saumik Bhattacharya, Aurobinda Routray

机构 * Department of EE, IIT Kharagpur, India(印度IIT Kharagpur电子工程系) Rekhi Centre of Excellence for the Science of Happiness, IIT Kharagpur, India(印度IIT Kharagpur幸福科学卓越中心) Department of E&ECE, IIT Kharagpur, India(印度IIT Kharagpur电子工程与电子学系)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

AI总结 本文提出基于l0正则化稀疏编码的可解释网络FNet,用于多模态图像融合,通过分离独特和共同特征提升融合质量并增强下游任务性能。

Comments Accetped by IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.09311 2025-12-11 cs.CV cs.CR 79%

Transformer-Driven Multimodal Fusion for Explainable Suspiciousness Estimation in Visual Surveillance

基于Transformer的多模态融合用于视觉监控中的可解释性可疑性估计

Kuldeep Singh Yadav, Lalan Kumar

机构 * Big Data Research and Supercomputing Division, CSIR Fourth Paradigm Institute(CSIR第四范式研究所大数据研究与超级计算部门) Department of Electrical Engineering, Bharti School of Telecommunication, Yardi School of Artificial Intelligence, IIT Delhi(电信学院电子工程系、Yardi人工智能学院、德里理工学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出基于Transformer的多模态融合框架DeepUSEvision,结合YOLOv12、双深度卷积网络和Transformer判别器,实现高准确率和可解释性的可疑性估计。

Comments 12 pages, 10 figures, IEEE Transaction on Image Processing

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.08135 2025-12-10 cs.CV 79%

CVP: Central-Peripheral Vision-Inspired Multimodal Model for Spatial Reasoning

CVP:基于中央-外围视觉的多模态模型用于空间推理

Zeyuan Chen, Xiang Zhang, Haiyang Xu, Jianwen Xie, Zhuowen Tu

机构 * UC San Diego(圣迭戈大学) Lambda, Inc(Lambda公司)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 CVP通过结合中央视觉和外围视觉的启发,提出了一种多模态模型,以提升复杂3D环境的空间推理能力。

Comments Accepted to WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05996 2025-12-09 cs.CV cs.CY cs.RO eess.IV 79%

FishDetector-R1: Unified MLLM-Based Framework with Reinforcement Fine-Tuning for Weakly Supervised Fish Detection, Segmentation, and Counting

FishDetector-R1: 基于统一MLLM框架的弱监督鱼类检测、分割与计数强化微调方法

Yi Liu, Jingyu Song, Vedanth Kallakuri, Katherine A. Skinner

机构 * University of Michigan(密歇根大学)

专题命中 多模态训练与对齐 :MLLM(title,abstract);分类 cs.CV

AI总结 FishDetector-R1通过统一MLLM框架和强化学习微调,实现了弱监督下的鱼类检测、分割与计数的高效准确提升。

Comments 18 pages, under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10573 2025-12-09 cs.CV 79%

Improving Medical Visual Representation Learning with Pathological-level Cross-Modal Alignment and Correlation Exploration

通过病理级跨模态对齐和相关性探索提升医学视觉表示学习

Jun Wang, Lixing Zhu, Xiaohan Yu, Abhir Bhalerao, Yulan He

机构 * Department of Computer Science, University of Warwick(沃里克大学计算机科学系) Department of Informatics, King’s College London(伦敦国王学院信息学系) School of Computing, Macquarie University(麦考瑞大学计算机科学学院) Alan Turing Institute, UK(英国艾伦·图灵研究所)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

AI总结 本文提出PLACE框架,通过病理级跨模态对齐和相关性探索提升医学视觉表示学习,实现多下游任务的性能提升。

Comments Accepted to IEEE Journal of Biomedical and Health Informatics (JBHI).Code: https://github.com/Markin-Wang/PLACE

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.04943 2025-12-05 cs.CV 79%

Towards Adaptive Fusion of Multimodal Deep Networks for Human Action Recognition

面向人类动作识别的多模态深度网络自适应融合

Novanto Yudistira

机构 * Departemen Teknik Informatika, Fakultas Ilmu Komputer, Universitas Brawijaya(计算机科学系,信息学院,布拉格亚大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出了一种基于多模态深度网络的自适应融合方法,通过门控机制提升人类动作识别的准确性和鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.16848 2025-12-05 cs.CV 79%

A re-calibration method for object detection with multi-modal alignment bias in autonomous driving

面向自动驾驶的多模态对齐偏差校准方法

Zhihang Song, Dingyi Yao, Ruibo Ming, Lihui Peng, Danya Yao, Yi Zhang

机构 * Department of Automation Tsinghua University Beijing(自动化系清华大学北京)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

AI总结 本文提出一种重新校准模型,通过语义分割和定制损失函数提升自动驾驶中多模态检测的鲁棒性和性能,应对校准偏差带来的影响。

Comments Accepted for publication in IST 2025. Official IEEE Xplore entry will be available once published

Journal ref 2025 IEEE International Conference on Imaging Systems and Techniques (IST)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.00530 2025-12-04 cs.LG cs.AI cs.SI 79%

Generic Multimodal Spatially Graph Network for Spatially Embedded Network Representation Learning

通用多模态空间图网络用于空间嵌入网络表示学习

Xudong Fan, Jürgen Hackl

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

AI总结 本文提出通用多模态空间图卷积网络,通过多模态特征提升空间嵌入网络的表示准确性,实验显示在电力网络中边存在预测任务的准确率提高37.1%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03404 2025-12-04 cs.CV 79%

MOS: Mitigating Optical-SAR Modality Gap for Cross-Modal Ship Re-Identification

MOS:缓解光学-合成孔径雷达模态差距以实现跨模态舰船重识别

Yujian Zhao, Hankun Liu, Guanglin Niu

机构 * School of Artificial Intelligence, Beihang University(北京航空航天大学人工智能学院) School of Computer Science and Engineering, Beihang University(北京航空航天大学计算机科学与工程学院)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

AI总结 MOS通过模态一致表示学习和跨模态数据生成与融合,有效缓解光学与SAR图像间的模态差距,提升舰船跨模态重识别的准确率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.01513 2025-12-04 cs.CR cs.CV 79%

SafePTR: Token-Level Jailbreak Defense in Multimodal LLMs via Prune-then-Restore Mechanism

SafePTR: 通过剪枝-恢复机制实现多模态大语言模型的令牌级 Jailbreak 防御

Beitao Chen, Xinyu Lyu, Lianli Gao, Jingkuan Song, Heng Tao Shen

机构 * Shenzhen Institute for Advanced Study, University of Electronic Science and Technology of China(电子科技大学深圳研究院) Southwestern University of Finance and Economics(西南财经大学) Engineering Research Center of Intelligent Finance, Ministry of Education(教育部智能金融工程研究中心) Tongji University(同济大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 SafePTR 提出一种无需训练的多模态大语言模型防御机制,通过剪枝有害令牌并恢复良性特征,有效提升安全性并保持效率。

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01949 2025-12-02 cs.CV 79%

Script: Graph-Structured and Query-Conditioned Semantic Token Pruning for Multimodal Large Language Models

脚本:图结构和查询条件的语义令牌修剪用于多模态大语言模型

Zhongyu Yang, Dannong Xu, Wei Pang, Yingfang Yuan

机构 * BCML, Heriot-Watt University(赫瑞瓦德大学BCML中心)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 Script通过图结构和查询条件的语义令牌修剪,提升多模态大语言模型的效率和准确性,实现显著的性能提升。

Comments Published in Transactions on Machine Learning Research, Project in https://01yzzyu.github.io/script.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏