arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6903 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6903 篇

2601.00352 2026-01-05 cs.CV 83%

OmniVaT: Single Domain Generalization for Multimodal Visual-Tactile Learning

OmniVaT:单域泛化用于多模态视觉-触觉学习

Liuxiang Qiu, Hui Da, Yuzhen Niu, Tiesong Zhao, Yang Cao, Zheng-Jun Zha

机构 * Fujian Key Laboratory for Intelligent Processing and Wireless Transmission of Media Information(福建智能媒体信息处理与无线传输重点实验室) College of Physics and Information Engineering(物理与信息工程学院) Fuzhou University(福州市大学) College of Computer and Data Science(计算机与数据科学学院) MoE Key Laboratory of Brain-inspired Intelligent Perception and Cognition(MoE脑启发智能感知与认知重点实验室) University of Science and Technology of China(中国科学技术大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 OmniVaT通过多模态分数傅里叶适配器和离散树生成模块,首次实现单域泛化多模态视觉-触觉学习任务,提升跨领域适应性与泛化性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.24324 2026-01-01 cs.LG cs.AI 83%

Empower Low-Altitude Economy: A Reliability-Aware Dynamic Weighting Allocation for Multi-modal UAV Beam Prediction

赋能低空经济:一种可靠性感知的动态权重分配用于多模态无人机波束预测

Haojin Li, Anbang Zhang, Chen Sun, Chenyuan Feng, Kaiqian Qu, Tony Q. S. Quek, Haijun Zhang

机构 * University of Science and Technology Beijing(北京科技大学) Sony China Research Laboratory(索尼中国研究院) School of Control Science and Engineering, Shandong University(山东大学控制科学与工程学院) Southeast University(东南大学) College of Computer Science, University of Exeter(埃克塞特大学计算机学院) Information Systems Technology and Design Pillar, Singapore University of Technology and Design(新加坡科技设计大学信息系统技术与设计系)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.AI

AI总结 本文提出SaM2B框架,通过可靠性感知的动态权重分配和跨模态对比学习,提升多模态无人机波束预测的准确性和鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21508 2025-12-29 cs.CV 83%

Fixed-Budget Parameter-Efficient Training with Frozen Encoders Improves Multimodal Chest X-Ray Classification

固定预算参数高效训练结合冻结编码器提升多模态胸片分类

Md Ashik Khan, Md Nahid Siddique

机构 * Department of Computer Science and Engineering, Indian Institute of Technology Kharagpur, India(计算机科学与工程系,印度理工学院Kharagpur分校) Knight Foundation School of Computing and Information Sciences, Florida International University, Florida, USA(骑士基金会计算与信息科学学院,佛罗里达国际大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 本研究通过冻结编码器的参数高效训练策略,在降低计算成本的同时提升了多模态胸片分类的性能。

Comments Accepted at the 2025 28th International Conference on Computer and Information Technology (ICCIT). 6 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.20670 2025-12-25 cs.LG cs.AI 83%

Disentangling Fact from Sentiment: A Dynamic Conflict-Consensus Framework for Multimodal Fake News Detection

区分事实与情感:一种动态冲突-共识框架用于多模态虚假新闻检测

Weilin Zhou, Zonghao Ying, Junjie Mu, Shengwei Tian, Quanchen Zou, Deyue Zhang, Dongdong Yang, Xiangzheng Zhang

机构 * Xinjiang University(新疆大学) AI Security Lab(360AI安全实验室) Beihang University(北航) South China University of Technology(华南理工大学) Fudan University(复旦大学) Shenzhen University(深圳大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

AI总结 本文提出动态冲突-共识框架,通过区分事实与情感空间,利用物理启发式特征动态和冲突-共识机制,提升多模态虚假新闻检测的准确率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.18986 2025-12-23 cs.LG cs.AI 83%

R-GenIMA: Integrating Neuroimaging and Genetics with Interpretable Multimodal AI for Alzheimer's Disease Progression

R-GenIMA:整合神经影像与基因组学的可解释多模态AI用于阿尔茨海默病进展

Kun Zhao, Siyuan Dai, Yingying Zhang, Guodong Liu, Pengfei Gu, Chenghua Lin, Paul M. Thompson, Alex Leow, Heng Huang, Lifang He, Liang Zhan, Haoteng Tang

机构 * Eli and Lilly company(艾利和利公司)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

AI总结 R-GenIMA通过整合神经影像与基因组学,利用可解释的多模态AI方法,实现了对阿尔茨海默病进展的精准预测与机制揭示。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.15707 2025-12-18 cs.CV 83%

GateFusion: Hierarchical Gated Cross-Modal Fusion for Active Speaker Detection

GateFusion:用于活动说话检测的分层门控跨模态融合

Yu Wang, Juhyung Ha, Frangil M. Ramirez, Yuchen Wang, David J. Crandall

机构 * Indiana University(印第安纳大学)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV

AI总结 GateFusion通过分层门控融合解码器提升活动说话检测的跨模态融合效果,实现新的SOTA结果。

Comments accepted by WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.15250 2025-12-18 cs.LG cs.AI 83%

Leveraging Foundational Models and Simple Fusion for Multi-modal Physiological Signal Analysis

利用基础模型和简单融合进行多模态生理信号分析

Youssef Ghallab, Omar Iraqy, Mohamed Kandil, Mohamed Ashraf, Saadeldine Eletter, Morougue Ghazal, Ayman Khalafallah, Nagwa El-Makky

机构 * Computer and Communication Engineering Department, Alexandria University(亚历山大大学计算机与通信工程系) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.AI

AI总结 本文提出利用基础模型和简单融合方法,通过双掩码策略和对称编码器提升多模态生理信号分析的性能,实现情绪识别的高精度结果。

Comments Published at NeurIPS 2025 Workshop on Foundation Models for the Brain and Body

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04870 2025-12-16 eess.IV cs.CV 83%

Multi-modal Uncertainty Robust Tree Cover Segmentation For High-Resolution Remote Sensing Images

多模态不确定性鲁棒树冠覆盖分割用于高分辨率遥感图像

Yuanyuan Gui, Wei Li, Yinjian Wang, Xiang-Gen Xia, Mauro Marty, Christian Ginzler, Zuyuan Wang

机构 * School of Information and Electronics, Beijing Institute of Technology(信息与电子学院,北京理工大学) National Key Laboratory of Science and Technology on Space-Born Intelligent Information Processing(空间智能信息处理国家重点实验室) Department of Electrical and Computer Engineering, University of Delaware(电气与计算机工程系,德雷塞尔大学) Swiss Federal Institute for Forest, Snow, and Landscape Research WSL(瑞士森林、雪和景观研究联邦 institute WSL)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 MURTreeFormer通过多模态分割框架降低不确定性,提升高分辨率遥感图像中树冠分割的鲁棒性与准确性。

Journal ref IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13261 2025-12-16 cs.CV 83%

Unlocking the Potential of Difficulty Prior in RL-based Multimodal Reasoning

解锁基于强化学习的多模态推理中难度先验的潜力

Mingrui Chen, Haogeng Liu, Hao Liang, Huaibo Huang, Wentao Zhang, Ran He

机构 * School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) NLPR&MAIS, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Peking University(北京大学) Zhongguancun Academy(中关村学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV

AI总结 本文通过建模问题难度先验信息,改进基于强化学习的多模态推理性能,通过数据筛选、优势分化和难度提示提升模型推理深度和准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11901 2025-12-16 cs.CV cs.LG 83%

CLARGA: Multimodal Graph Representation Learning over Arbitrary Sets of Modalities

CLARGA:任意模态集合上的多模态图表示学习

Santosh Patapati

机构 * Santosh Patapati(独立研究者)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 CLARGA是一种通用的多模态融合架构,通过构建注意力加权图实现多模态表示学习,适用于任意模态集合,具有高效的融合能力和良好的鲁棒性。

Comments WACV; Supplementary material is available on CVF proceedings

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19312 2025-12-15 eess.SP cs.AI cs.IT cs.LG math.IT 83%

E2E Learning Massive MIMO for Multimodal Semantic Non-Orthogonal Transmission and Fusion

端到端学习大规模MIMO用于多模态语义非正交传输与融合

Minghui Wu, Zhen Gao

机构 * School of Information and Electronics, Beijing Institute of Technology (BIT)(信息与电子学院,北京理工大学) State Key Laboratory of Environment Characteristics and Effects for Near-space, Beijing(临近空间环境特征与效应国家重点实验室,北京) State Key Laboratory of CNS/ATM, Beijing(CNS/ATM国家重点实验室,北京) MIIT Key Laboratory of Complex-Field Intelligent Sensing, Beijing(工信部复杂场智能感知重点实验室,北京) BIT, Zhuhai(珠海北京理工大学) Advanced Technology Research Institute, BIT, Jinan(北京理工大学济南先进技术研究院) Yangtze Delta Region Academy, BIT, Jiaxing(长江三角洲地区学院,北京理工大学嘉兴)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

AI总结 本文提出了一种端到端学习的多模态语义非正交传输与融合框架,通过联合优化物理层和应用层任务提升大规模MIMO的频谱效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.09801 2025-12-11 cs.CV 83%

Modality-Specific Enhancement and Complementary Fusion for Semi-Supervised Multi-Modal Brain Tumor Segmentation

模态特异性增强与互补融合用于半监督多模态脑肿瘤分割

Tien-Dat Chung, Ba-Thinh Lam, Thanh-Huy Nguyen, Thien Nguyen, Nguyen Lan Vi Vu, Hoang-Loc Cao, Phat Kim Huynh, Min Xu

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出了一种半监督多模态脑肿瘤分割框架,通过模态特异性增强模块和互补信息融合模块提升分割性能。

Comments 9 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.07430 2025-12-09 cs.LG cs.AI 83%

MIDG: Mixture of Invariant Experts with knowledge injection for Domain Generalization in Multimodal Sentiment Analysis

MIDG:基于知识注入的混合不变专家用于多模态情感分析中的领域泛化

Yangle Li, Danli Luo, Haifeng Hu

机构 * School of Electronics and Information Technology(电子信息学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

AI总结 MIDG通过混合不变专家和跨模态适配器,提升多模态情感分析中领域泛化的性能与语义表达能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05515 2025-12-08 cs.CV cs.LG 83%

DashFusion: Dual-stream Alignment with Hierarchical Bottleneck Fusion for Multimodal Sentiment Analysis

DashFusion: 基于分层瓶颈融合的双流对齐多模态情感分析

Yuhua Wen, Qifei Li, Yingying Zhou, Yingming Gao, Zhengqi Wen, Jianhua Tao, Ya Li

机构 * School of Artificial Intelligence, Beijing University of Posts and Telecommunications(北京邮电大学人工智能学院) Beijing National Research Center for Information Science and Technology, Tsinghua University(清华大学信息科学与技术国家研究中心) Department of Automation, Tsinghua University(清华大学自动化系)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 DashFusion通过双流对齐与分层瓶颈融合技术,提升多模态情感分析的性能与效率。

Comments Accepted to IEEE Transactions on Neural Networks and Learning Systems (TNNLS), 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16054 2025-12-08 cs.CV 83%

Language-Instructed Reasoning for Group Activity Detection via Multimodal Large Language Model

基于多模态大语言模型的语言引导推理用于群体活动检测

Jihua Peng, Qianxiong Xu, Yichen Liu, Chenxi Liu, Cheng Long, Rui Zhao, Ziyue Li

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

AI总结 本文提出LIR-GAD框架,通过多模态大语言模型实现群体活动检测,引入活动标记和群体标记以提升语义理解和分类性能。

Comments This work is being incorporated into a larger study

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02558 2025-12-03 cs.AI 83%

Empathy Level Prediction in Multi-Modal Scenario with Supervisory Documentation Assistance

多模态场景中基于监督文档辅助的共情水平预测

Yufei Xiao, Shangfei Wang

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.AI

AI总结 本文提出一种结合视频、音频和文本信息的多模态共情预测方法,通过监督文档辅助训练提升文本特征提取,实验证明其优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00363 2025-12-02 cs.CV 83%

MM-DETR: An Efficient Multimodal Detection Transformer with Mamba-Driven Dual-Granularity Fusion and Frequency-Aware Modality Adapters

MM-DETR: 一种高效的多模态检测Transformer,采用Mamba驱动的双粒度融合和频率感知模态适配器

Jianhong Han, Yupei Wang, Yuan Zhang, Liang Chen

机构 * School of Information and Electronics, Beijing Institute of Technology(信息与电子学院,北京理工大学) Beijing Institute of Technology Chongqing Innovation Center(北京理工大学重庆创新中心) National Key Laboratory for Space-Born Intelligent Information Processing(空间智能信息处理国家级重点实验室) School of Automation, Beijing Institute of Technology(自动化学院,北京理工大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 MM-DETR通过Mamba驱动的双粒度融合和频率感知模态适配器,实现高效的多模态目标检测,提升检测精度与轻量化性能。

Comments Manuscript submitted to IEEE Transactions on Geoscience and Remote Sensing

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.23287 2025-12-01 cs.LG cs.CL 83%

Transformer-Driven Triple Fusion Framework for Enhanced Multimodal Author Intent Classification in Low-Resource Bangla

基于Transformer的三融合框架用于低资源孟加拉语多模态作者意图分类

Ariful Islam, Tanvir Mahmud, Md Rifat Hossen

机构 * Department of Computer Science(计算机科学系) Engineering Chittagong University of Engineering(工程学院恰尔达格工程大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL

AI总结 本文提出基于Transformer的三融合框架BangACMM,通过结合文本和视觉数据,在低资源孟加拉语社交媒体中实现作者意图分类,达到84.11%的宏F1得分,提升8.4个百分点。

Comments Accepted at the 28th International Conference on Computer and Information Technology (ICCIT 2025). To be published in IEEE proceedings

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22897 2025-12-01 cs.CV 83%

From Points to Clouds: Learning Robust Semantic Distributions for Multi-modal Prompts

从点到云:学习多模态提示的稳健语义分布

Weiran Li, Yeqiang Liu, Yijie Wei, Mina Han, Xin Liu, Zhenbo Li

机构 * China Agricultural University(中国农业大学)

专题命中 多模态训练与对齐 :multi-modal(title);multimodal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出P2C框架,通过动态去噪机制学习语义云分布,提升多模态提示学习的鲁棒性和泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22103 2025-12-01 cs.CV 83%

MoE3D: Mixture of Experts meets Multi-Modal 3D Understanding

MoE3D:专家混合方法与多模态3D理解

Yu Li, Yuenan Hou, Yingmei Wei, Xinge Zhu, Yuexin Ma, Wenqi Shao, Yanming Guo

机构 * National University of Defense Technology(国防科技大学) Shanghai AI Laboratory(上海人工智能实验室) The Chinese University of Hong Kong(香港中文大学) ShanghaiTech University(上海科技大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 MoE3D通过整合专家混合方法,提升多模态3D理解的性能,尤其在Multi3DRefer任务中表现更优。

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.15711 2025-11-27 cs.AI q-bio.NC 83%

Semi-supervised Multimodal Representation Learning through a Global Workspace

通过全局工作空间实现半监督多模态表示学习

Benjamin Devillers, Léopold Maytié, Rufin VanRullen

机构 * CerCo, CNRS UMR 5549, Université de Toulouse and ANITI, Artificial and Natural Intelligence Toulouse Institute(CerCo、CNRS UMR 5549、图卢兹大学和ANITI人工智能与自然智能图卢兹研究所)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

AI总结 本文提出了一种受全局工作空间概念启发的神经网络架构,通过自监督学习实现多模态表示对齐与转换,显著减少对匹配数据的需求。

Comments Under review

Journal ref IEEE Transactions on Neural Networks and Learning Systems 36 (5), 7843-7857 (2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.20167 2025-11-26 cs.MM 83%

FINE: Factorized multimodal sentiment analysis via mutual INformation Estimation

FINE: 通过互信息估计进行因子化多模态情感分析

Yadong Liu, Shangfei Wang

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.MM

AI总结 本文提出了一种基于互信息估计的因子化多模态情感分析框架,通过分解模态为共享和独特表示,抑制噪声并提升情感表示质量,从而在多个数据集上优于现有方法。

Comments 15 pages, 9 figures, conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11895 2025-11-26 cs.CV 83%

Adversarial Robustness for Unified Multi-Modal Encoders via Efficient Calibration

通过高效校准实现统一多模态编码器的对抗鲁棒性

Chih-Ting Liao, Zhangquan Chen, Chunlei Meng, Tzu-Yu Huang, Xin Cao, Xu Zheng

机构 * UNSW Sydney(新南威尔士大学悉尼分校) Tsinghua University(清华大学) Fudan University(复旦大学) UTS(澳大利亚UTS大学) HKUST(GZ)(香港理工大学(广州))

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 本研究提出高效对抗校准框架,提升统一多模态编码器的对抗鲁棒性,同时保持清洁性能,提升47.3%的鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17952 2025-11-25 cs.CV 83%

Multi-speaker Attention Alignment for Multimodal Social Interaction

多说话者注意力对齐用于多模态社交互动

Liangyang Ouyang, Yifei Huang, Mingfang Zhang, Caixin Kang, Ryosuke Furuta, Yoichi Sato

机构 * The University of Tokyo(东京大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出了一种多模态多说话者注意力对齐方法,通过动态头选择和自适应注意力偏差提升多模态社交互动理解能力,实现SOTA效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17681 2025-11-25 cs.CV 83%

Vision-Motion-Reference Alignment for Referring Multi-Object Tracking via Multi-Modal Large Language Models

基于多模态大语言模型的视觉-运动-参考对齐的指称多目标跟踪

Weiyi Lv, Ning Zhang, Hanyang Sun, Haoran Jiang, Kai Zhao, Jing Xiao, Dan Zeng

机构 * Shanghai University(上海大学) PAII Inc.(PAII公司)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 VMRMOT通过多模态大语言模型实现视觉-运动-参考对齐,提升指称多目标跟踪的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15122 2025-11-25 cs.IR cs.AI 83%

Multi-Aspect Cross-modal Quantization for Generative Recommendation

多方面跨模态量化用于生成性推荐

Fuwei Zhang, Xiaoyu Liu, Dongbo Xi, Jishen Yin, Huan Chen, Peng Yan, Fuzhen Zhuang, Zhao Zhang

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multimodal(abstract);分类 cs.AI

AI总结 MACRec通过多方面跨模态量化方法,提升生成性推荐中多模态信息利用和语义ID学习的性能。

Comments Accepted by AAAI 2026 (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14766 2025-11-20 cs.IR cs.MM 83%

OTCR: Optimal Transmission, Compression and Representation for Multimodal Information Extraction

Yang Li, Yajiao Wang, Wenhao Hu, Zhixiong Zhang, Mengting Zhang

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.MM

Comments 5 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14604 2025-11-19 cs.CV 83%

XAttn-BMD: Multimodal Deep Learning with Cross-Attention for Femoral Neck Bone Mineral Density Estimation

Yilin Zhang, Leo D. Westbury, Elaine M. Dennison, Nicholas C. Harvey, Nicholas R. Fuggle, Rahman Attar

机构 * School of Electronics and Computer Science, University of Southampton, UK(电子与计算机科学学院,索姆塞特大学,英国) MRC Lifecourse Epidemiology Centre, University of Southampton, Southampton General Hospital, UK(生命课程流行病学研究中心,索姆塞特大学,南安普顿总医院,英国) NIHR Southampton Biomedical Research Centre, University of Southampton(南安普顿生物医学研究中心,索姆塞特大学) University Hospital NHS Foundation Trust, Southampton, UK(南安普顿国家健康服务基金会信托,英国)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments 11 figures, 10 tables, 38 pages. Submitted to Artificial Intelligence in Medicine (currently with editor)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13755 2025-11-19 cs.LG cs.AI 83%

Adaptive Redundancy Regulation for Balanced Multimodal Information Refinement

Zhe Yang, Wenrui Li, Hongtao Chen, Penghong Wang, Ruiqin Xiong, Xiaopeng Fan

机构 * Department of Computer Science and Technology, Harbin Institute of Technology(计算机科学与技术系,哈尔滨工业大学) Harbin Institute of Technology Zhengzhou Research Institute(哈尔滨工业大学郑州研究所) Harbin Institute of Technology Suzhou Research Institute(哈尔滨工业大学苏州研究所) School of Mathematical Sciences, University of Electronic Science and Technology of China(数学学院,电子科学与技术大学) School of Electronic Engineering and Computer Science, Institute of Digital Media, Peking University(电子工程与计算机科学系,数字媒体研究所,北京大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12982 2025-11-18 cs.CR cs.CV 83%

SafeGRPO: Self-Rewarded Multimodal Safety Alignment via Rule-Governed Policy Optimization

Xuankun Rong, Wenke Huang, Tingfeng Wang, Daiguo Zhou, Bo Du, Mang Ye

机构 * School of Computer Science, Wuhan University(武汉大学计算机学院) MiLM Plus, Xiaomi Inc.(小米公司)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏