arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-12-09 至 2025-12-09 共收录 18 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 18 篇

2507.14997 2025-12-09 cs.CV cs.LG 85%

Language Integration in Fine-Tuning Multimodal Large Language Models for Image-Based Regression

基于图像回归的多模态大语言模型微调中的语言整合

Roy H. Jennings, Genady Paikin, Roy Shaul, Evgeny Soloveichik

机构 * Samsung Israel R&D Center(三星以色列研发中心)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出RvTC方法,通过灵活的区间法替代传统词汇受限分类,提升多模态大语言模型在图像回归任务中的性能。

Comments WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.07568 2025-12-09 cs.CV cs.AI eess.IV 84%

Dual-Stream Cross-Modal Representation Learning via Residual Semantic Decorrelation

通过残差语义去相关实现双流跨模态表示学习

Xuecheng Li, Weikuan Jia, Alisher Kurbonaliev, Qurbonaliev Alisher, Khudzhamkulov Rustam, Ismoilov Shuhratjon, Eshmatov Javhariddin, Yuanjie Zheng

机构 * School of Information Science & Engineering, Shandong Normal University(信息科学与工程学院,山东师范大学) Tajikistan State University of Law, Business Sughd(塔吉克斯坦法律、商业大学,苏赫德) Tajik State University of Law, Business and Politics Sughd(塔吉克斯坦法律、商业与政治大学,苏赫德)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV、cs.AI

AI总结 DSRSD-Net通过残差分解和语义去相关解决跨模态学习中的模态主导和冗余问题,提升预测性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.07430 2025-12-09 cs.LG cs.AI 83%

MIDG: Mixture of Invariant Experts with knowledge injection for Domain Generalization in Multimodal Sentiment Analysis

MIDG:基于知识注入的混合不变专家用于多模态情感分析中的领域泛化

Yangle Li, Danli Luo, Haifeng Hu

机构 * School of Electronics and Information Technology(电子信息学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

AI总结 MIDG通过混合不变专家和跨模态适配器,提升多模态情感分析中领域泛化的性能与语义表达能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22410 2025-12-09 stat.AP 82%

Multimodal Fusion and Interpretability in Human Activity Recognition: A Reproducible Framework for Sensor-Based Modeling

多模态融合与可解释性在人体活动识别中的应用:一种可复现的基于传感器建模框架

Yiyao Yang, Yasemin Gulbahar

专题命中 多模态训练与对齐 :multimodal(title,abstract);multi-modal(abstract)

AI总结 本文提出了一种可复现的多模态融合框架,通过统一预处理和融合策略提升人体活动识别的准确性和可解释性。

Comments 33 pages, 12 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.06281 2025-12-09 cs.CV cs.AI 81%

Unleashing the Intrinsic Visual Representation Capability of Multimodal Large Language Models

释放多模态大语言模型的内在视觉表示能力

Hengzhuang Li, Xinsong Zhang, Qiming Peng, Bin Luo, Han Hu, Dengyang Jiang, Han-Jia Ye, Teng Zhang, Hai Jin

机构 * HUST(华中科技大学) Tencent Hunyuan Research(腾讯混元研究) HKUST(香港科技大学) NJU(南京大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 本文提出LaVer框架,通过掩码图像建模提升多模态大语言模型的视觉表示能力,实验表明其在需要密集视觉能力的场景中表现优异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05996 2025-12-09 cs.CV cs.CY cs.RO eess.IV 79%

FishDetector-R1: Unified MLLM-Based Framework with Reinforcement Fine-Tuning for Weakly Supervised Fish Detection, Segmentation, and Counting

FishDetector-R1: 基于统一MLLM框架的弱监督鱼类检测、分割与计数强化微调方法

Yi Liu, Jingyu Song, Vedanth Kallakuri, Katherine A. Skinner

机构 * University of Michigan(密歇根大学)

专题命中 多模态训练与对齐 :MLLM(title,abstract);分类 cs.CV

AI总结 FishDetector-R1通过统一MLLM框架和强化学习微调,实现了弱监督下的鱼类检测、分割与计数的高效准确提升。

Comments 18 pages, under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10573 2025-12-09 cs.CV 79%

Improving Medical Visual Representation Learning with Pathological-level Cross-Modal Alignment and Correlation Exploration

通过病理级跨模态对齐和相关性探索提升医学视觉表示学习

Jun Wang, Lixing Zhu, Xiaohan Yu, Abhir Bhalerao, Yulan He

机构 * Department of Computer Science, University of Warwick(沃里克大学计算机科学系) Department of Informatics, King’s College London(伦敦国王学院信息学系) School of Computing, Macquarie University(麦考瑞大学计算机科学学院) Alan Turing Institute, UK(英国艾伦·图灵研究所)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

AI总结 本文提出PLACE框架,通过病理级跨模态对齐和相关性探索提升医学视觉表示学习,实现多下游任务的性能提升。

Comments Accepted to IEEE Journal of Biomedical and Health Informatics (JBHI).Code: https://github.com/Markin-Wang/PLACE

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.07687 2025-12-09 cs.CL cs.CV 73%

HalluShift++: Bridging Language and Vision through Internal Representation Shifts for Hierarchical Hallucinations in MLLMs

HalluShift++: 通过内部表示转移弥合语言与视觉,解决多模态大语言模型中的层级幻觉

Sujoy Nath, Arkaprabha Basu, Sharanya Dasgupta, Swagatam Das

机构 * Netaji Subhash Engineering College (NSEC)(奈尔贾伊·萨布哈工程学院) TCG Crest Electronics and Communication Sciences Unit (ECSU)(电子与通信科学单位) Indian Statistical Institute(印度统计研究所)

专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.CL

AI总结 HalluShift++通过分析MLLM内部表示转移,解决多模态大语言模型中的层级幻觉问题,提升幻觉检测的有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.07170 2025-12-09 cs.CV cs.AI 73%

Towards Unified Semantic and Controllable Image Fusion: A Diffusion Transformer Approach

迈向统一的语义和可控图像融合:一种扩散变换器方法

Jiayang Li, Chengjie Jiang, Junjun Jiang, Pengwei Liang, Jiayi Ma, Liqiang Nie

机构 * Faculty of Computing, Harbin Institute of Technology(哈尔滨工业大学计算机学院) Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院) Electronic Information School, Wuhan University(武汉大学电子信息学院)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 DiTFuse通过融合图像与自然语言指令,实现端到端、语义感知的图像融合,统一了多种融合任务并在多个基准测试中表现出色。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.06447 2025-12-09 cs.CV 70%

Towards Stable Cross-Domain Depression Recognition under Missing Modalities

迈向稳定跨领域抑郁识别下的缺失模态

Jiuyi Chen, Mingkui Tan, Haifeng Lu, Qiuna Xu, Zhihua Wang, Runhao Zeng, Xiping Hu

机构 * School of Future Technology, South China University of Technology(未来技术学院,华南理工大学) PengCheng Laboratory(鹏城实验室) School of Software Engineering, South China University of Technology(软件工程学院,华南理工大学) Artificial Intelligence Research Institute, Shenzhen MSU-BIT University(人工智能研究院,深圳MSU-BIT大学) Guangdong-Hong Kong-Macao Joint Laboratory for Emotional Intelligence and Pervasive Computing(粤港澳大湾区情感智能与泛在计算联合实验室) School of Computer Science and Technology, Guangdong University of Technology(计算机科学与技术学院,广东工业大学) Department of Computer Science, City University of Hong Kong(计算机科学系,香港城市大学)

专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract);分类 cs.CV

AI总结 本文提出SCD-MLLM框架,通过多源数据适配器和模态感知自适应融合模块,实现稳定跨领域抑郁症识别,提升多模态数据处理的鲁棒性和泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01946 2025-12-09 cs.CV 70%

3DRS: MLLMs Need 3D-Aware Representation Supervision for Scene Understanding

3DRS: MLLMs 需要 3D 意识表示监督以实现场景理解

Xiaohu Huang, Jingjing Wu, Qunyi Xie, Kai Han

机构 * Visual AI Lab, The University of Hong Kong(香港大学视觉人工智能实验室) Department of Computer Vision Technology (VIS), Baidu Inc.(百度公司计算机视觉技术部)

专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract);分类 cs.CV

AI总结 3DRS 通过引入预训练 3D 基础模型的监督,提升 MLLM 的 3D 表示能力,从而增强场景理解性能

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.06726 2025-12-09 cs.CV cs.AI cs.CL 67%

The Role of Entropy in Visual Grounding: Analysis and Optimization

熵在视觉定位中的作用:分析与优化

Shuo Li, Jiajun Sun, Zhihao Zhang, Xiaoran Fan, Senjie Jin, Hui Li, Yuming Yang, Junjie Ye, Lixing Shen, Tao Ji, Tao Gui, Qi Zhang, Xuanjing Huang

机构 * Fudan University(复旦大学) Hikvision Research Institute(海康威视研究院)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 本文提出ECVGPO算法,通过熵控制优化视觉定位任务,提升探索与利用的平衡,实现多基准的性能提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.07141 2025-12-09 cs.CV cs.CL 62%

Think-Reflect-Revise: A Policy-Guided Reflective Framework for Safety Alignment in Large Vision Language Models

思考-反思-修订:一种基于策略的反思框架,用于大型视觉语言模型的安全对齐

Fenghua Weng, Chaochao Lu, Xia Hu, Wenqi Shao, Wenjie Wang

机构 * Shanghaitech University(上海科技大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.CL

AI总结 TRR通过策略引导的反思框架提升大型视觉语言模型的安全对齐,显著提高安全响应率至87.7%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.06848 2025-12-09 cs.CL cs.CV 62%

AquaFusionNet: Lightweight VisionSensor Fusion Framework for Real-Time Pathogen Detection and Water Quality Anomaly Prediction on Edge Devices

AquaFusionNet:轻量级视觉传感器融合框架,用于边缘设备上的实时病原体检测和水质异常预测

Sepyan Purnama Kristanto, Lutfi Hakim, Hermansyah

机构 * Department of Informatics Engineering, Politeknik Negeri Banyuwangi(信息工程系,普特里克国家理工学院巴扬威angi分校) Balai Besar Teknik Kesehatan Lingkungan dan P2B Surabaya(环境与P2B技术研究所Surabaya)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV、cs.CL

AI总结 AquaFusionNet通过跨模态融合提升边缘设备上病原体检测和水质异常预测的准确率与效率。

Comments 9Pages, 3 figure, Politeknik Negeri Banyuwangi

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.07021 2025-12-09 cs.LG cs.AI 57%

Transferring Clinical Knowledge into ECGs Representation

将临床知识转移到ECG表示中

Jose Geraldo Fernandes, Luiz Facury de Souza, Pedro Robles Dutenhefner, Gisele L. Pappa, Wagner Meira

机构 * Department of Computer Science(计算机科学系) Universidade Federal de Minas Gerais(巴西矿务联邦大学) Depart. of Math. and Comp. Sc.(数学与计算机科学系) Albion College(阿尔比恩学院)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

AI总结 本文提出一种三阶段训练方法,将多模态临床数据转化为ECG表示,提升分类模型的准确性和可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.06662 2025-12-09 cs.CV 57%

Personalized Image Descriptions from Attention Sequences

基于注意力序列的个性化图像描述

Ruoyu Xue, Hieu Le, Jingyi Xu, Sounak Mondal, Abe Leite, Gregory Zelinsky, Minh Hoai, Dimitris Samaras

机构 * Stony Brook University(石溪大学) UNC-Charlotte(北卡罗来纳大学夏洛特分校) The University of Adelaide(阿德莱德大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

AI总结 DEPER通过建模个性化注意力序列,实现了基于视觉-语言模型的个性化图像描述生成,提升了描述质量与人类对齐性。

Comments 10 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05979 2025-12-09 physics.chem-ph cs.AI cs.DM cs.LG 57%

Accelerating Materials Discovery: Learning a Universal Representation of Chemical Processes for Cross-Domain Property Prediction

加速材料发现:学习化学过程的通用表示以实现跨领域性质预测

Mikhail Tsitsvero, Atsuyuki Nakao, Hisaki Ikebata

机构 * CrowdChem, Inc.(CrowdChem公司)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.AI

AI总结 本文提出了一种通用的过程图表示,通过多模态图神经网络实现跨领域性质预测,展示了大规模学习的通用表示在专门任务中的高效迁移能力。

Comments 22 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14670 2025-12-09 cs.CV 57%

Gene-DML: Dual-Pathway Multi-Level Discrimination for Gene Expression Prediction from Histopathology Images

Gene-DML:双路径多级判别用于从组织病理图像预测基因表达

Yaxuan Song, Jianan Fan, Hang Chang, Weidong Cai

机构 * The University of Sydney, Australia(悉尼大学) Lawrence Berkeley National Laboratory, USA(伯克利国家实验室)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

AI总结 Gene-DML通过双路径多级判别方法,提升组织病理图像与基因表达谱的跨模态对齐,实现高精度的基因表达预测。

Comments Accepted by The IEEE/CVF Winter Conference on Applications of Computer Vision 2026 (WACV2026). Code and data available at https://github.com/YXSong000/Gene-DML

详情

展开后加载摘要…

URL PDF HTML 收藏