arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6897 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6897 篇

2512.21476 2025-12-29 cs.CV cs.AI 62%

GPF-Net: Gated Progressive Fusion Learning for Polyp Re-Identification

GPF-Net:门控渐进融合学习用于息肉重识别

Suncheng Xiang, Xiaoyang Wang, Junjie Jiang, Hejia Wang, Dahong Qian

机构 * Shanghai Jiao Tong University(上海交通大学) Peking University(北京大学) Shanghai Fifth People's Hospital(上海第五人民医院)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 GPF-Net通过门控渐进融合学习提升息肉重识别性能,结合多模态融合策略优于现有单模态模型。

Comments Work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21452 2025-12-29 cs.CV cs.AI 62%

Intelligent recognition of GPR road hidden defect images based on feature fusion and attention mechanism

基于特征融合与注意力机制的智能GPR道路隐缺陷图像识别

Haotian Lv, Yuhui Zhang, Jiangbo Dai, Hanli Wu, Jiaji Wang, Dawei Wang

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV、cs.AI

AI总结 本研究提出基于特征融合与注意力机制的GPR道路隐缺陷图像识别框架,通过数据增强、多模态特征融合和迁移学习提升检测精度与鲁棒性。

Comments Accepted for publication in *IEEE Transactions on Geoscience and Remote Sensing*

Journal ref IEEE Transactions on Geoscience and Remote Sensing, 2025, 63, 5213217

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.18504 2025-12-23 cs.CV cs.AI 62%

GTMA: Dynamic Representation Optimization for OOD Vision-Language Models

GTMA:面向视觉-语言模型的动态表示优化

Jensen Zhang, Ningyuan Liu, Keze Wang

机构 * Sun Yat-sen University(中山大学)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 GTMA通过动态表示优化提升视觉-语言模型在分布外任务中的性能,有效解决模态不对称问题,提升零样本和少样本准确率15-20%。

Comments Under submission

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12690 2025-12-16 cs.LG cs.CL cs.CV 62%

Reassessing the Role of Supervised Fine-Tuning: An Empirical Study in VLM Reasoning

重新评估监督微调的作用:VLM推理中的实证研究

Yongcan Yu, Lingxiao He, Shuo Lu, Lijun Sheng, Yinuo Xu, Yanbo Wang, Kuangpu Guo, Jianjie Cheng, Meng Wang, Qianlong Xie, Xingxing Wang, Dapeng Hu, Jian Liang

机构 * NLPR & MAIS, Institute of Automation, Chinese Academy of Sciences(人工智能研究院 & 模式识别与人工智能研究所,中国科学院自动化研究所) School of Artificial Intelligence, University of Chinese Academy of Sciences(人工智能学院,中国科学院大学) Meituan(美团)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV、cs.CL

AI总结 本研究通过实证分析发现,监督微调在VLM推理中具有重要作用,挑战了强化学习优于监督微调的主流观点。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11849 2025-12-16 cs.RO cs.AI cs.CV cs.SY eess.IV eess.SY 62%

LocoMamba: Vision-Driven Locomotion via End-to-End Deep Reinforcement Learning with Mamba

LocoMamba:基于端到端深度强化学习的视觉驱动运动控制框架

Yinuo Wang, Gavin Tao

机构 * School of Computer Science and Statistics, Trinity College Dublin(计算机科学与统计学系,三一学院都柏林)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 LocoMamba通过端到端深度强化学习实现视觉驱动的运动控制,利用Mamba模型提升序列建模效率,有效捕捉长程依赖并增强训练效率。

Comments 14 pages. This paper has been published in Advanced Engineering Informatics. Please cite the journal version: DOI: 10.1016/j.aei.2025.104230

Journal ref Advanced Engineering Informatics, Vol. 70, Art. no. 104230 (2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17552 2025-12-12 cs.CL cs.AI 62%

Can LLMs Reason Over Non-Text Modalities in a Training-Free Manner? A Case Study with In-Context Representation Learning

LLMs能否在无训练模式下推理非文本模态?一种基于上下文表示学习的案例研究

Tianle Zhang, Wanlong Fang, Jonathan Woo, Paridhi Latawa, Deepak A. Subramanian, Alvin Chan

机构 * College of Computing and Data Science, Nanyang Technological University(南洋理工大学计算机与数据科学学院) AI-X, Interdisciplinary Graduate Programme, Nanyang Technological University(南洋理工大学人工智能交叉研究生项目) Lee Kong Chian School of Medicine, Nanyang Technological University(南洋理工大学李科钦医学院) Centre of AI in Medicine (C-AIM), Nanyang Technological University(南洋理工大学医学人工智能中心) University of Toronto(多伦多大学) Brigham and Women’s Hospital, Harvard Medical School(哈佛医学院布里洛妇女医院) Massachusetts Institute of Technology(麻省理工学院)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CL、cs.AI

AI总结 本文提出ICRL框架,使LLMs在无训练情况下利用非文本模态表示,通过少量学习实现多模态推理,为适应性泛化提供新方向。

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06249 2025-12-11 cs.CL cs.AI 62%

TRepLiNa: Layer-wise CKA+REPINA Alignment Improves Low-Resource Machine Translation in Aya-23 8B

TRepLiNa:分层CKA+REPINA对齐改进Aya-23 8B低资源机器翻译

Toshiki Nakai, Ravi Kiran Chikkala, Lena Sophie Oberkircher, Nicholas Jennings, Natalia Skachkova, Tatiana Anikina, Jesujoba Oluwadara Alabi

机构 * Saarland University(萨尔兰大学) German Research Center for Artificial Intelligence (DFKI)(德国人工智能研究中心)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CL、cs.AI

AI总结 TRepLiNa通过结合CKA和REPINA实现分层对齐,提升低资源语言Aya-23 8B的机器翻译质量,尤其在数据稀缺情况下效果显著。

Comments It is work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.07141 2025-12-09 cs.CV cs.CL 62%

Think-Reflect-Revise: A Policy-Guided Reflective Framework for Safety Alignment in Large Vision Language Models

思考-反思-修订:一种基于策略的反思框架,用于大型视觉语言模型的安全对齐

Fenghua Weng, Chaochao Lu, Xia Hu, Wenqi Shao, Wenjie Wang

机构 * Shanghaitech University(上海科技大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.CL

AI总结 TRR通过策略引导的反思框架提升大型视觉语言模型的安全对齐,显著提高安全响应率至87.7%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.06848 2025-12-09 cs.CL cs.CV 62%

AquaFusionNet: Lightweight VisionSensor Fusion Framework for Real-Time Pathogen Detection and Water Quality Anomaly Prediction on Edge Devices

AquaFusionNet:轻量级视觉传感器融合框架,用于边缘设备上的实时病原体检测和水质异常预测

Sepyan Purnama Kristanto, Lutfi Hakim, Hermansyah

机构 * Department of Informatics Engineering, Politeknik Negeri Banyuwangi(信息工程系,普特里克国家理工学院巴扬威angi分校) Balai Besar Teknik Kesehatan Lingkungan dan P2B Surabaya(环境与P2B技术研究所Surabaya)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV、cs.CL

AI总结 AquaFusionNet通过跨模态融合提升边缘设备上病原体检测和水质异常预测的准确率与效率。

Comments 9Pages, 3 figure, Politeknik Negeri Banyuwangi

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05481 2025-12-08 cs.CV cs.AI 62%

UniFS: Unified Multi-Contrast MRI Reconstruction via Frequency-Spatial Fusion

UniFS: 通过频率-空间融合实现统一的多对比MRI重建

Jialin Li, Yiwei Ren, Kai Pan, Dong Wei, Pujin Cheng, Xian Wu, Xiaoying Tang

机构 * Department of Electronic and Electrical Engineering, Southern University of Science and Technology, Shenzhen, China(电子与电气工程系,南方科技大学,深圳,中国) Jarvis Research Center, Tencent YouTu Lab, China(Jarvis研究中心,腾讯YouTu实验室,中国) Department of Electrical and Electronic Engineering, University of Hong Kong, Hong Kong, China(电子与电气工程系,香港大学,香港,中国)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 UniFS通过频率-空间融合模块实现多对比MRI重建的统一处理,提升模型泛化能力,适用于多种k空间欠采样模式。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.23075 2025-12-05 cs.CV cs.AI 62%

SpaceMind: Camera-Guided Modality Fusion for Spatial Reasoning in Vision-Language Models

SpaceMind: 基于摄像头引导的模态融合用于视觉-语言模型中的空间推理

Ruosen Zhao, Zhikang Zhang, Jialei Xu, Jiahao Chang, Dong Chen, Lingyun Li, Weijian Sun, Zizhuang Wei

机构 * Huawei(华为) The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) The University of Hong Kong(香港大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 SpaceMind通过摄像头引导的模态融合方法,提升视觉-语言模型在空间推理任务中的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21395 2025-12-01 cs.CV cs.AI 62%

Monet: Reasoning in Latent Visual Space Beyond Images and Language

Monet: 在图像和语言之外的潜在视觉空间推理

Qixun Wang, Yang Shi, Yifei Wang, Yuanxing Zhang, Pengfei Wan, Kun Gai, Xianghua Ying, Yisen Wang

机构 * Peking University(北京大学) Kling Team(Kling团队) Amazon AGI SF Lab(Amazon AGI SF实验室) MIT(麻省理工学院)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 Monet通过在潜在视觉空间中直接推理,提升多模态大语言模型的视觉推理能力,解决了潜在-视觉对齐和嵌入监督不足的问题,展示了在现实世界和抽象视觉任务中的优越表现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15691 2025-11-26 q-fin.CP cs.AI cs.CL cs.LG 62%

Exploring the Synergy of Quantitative Factors and Newsflow Representations from Large Language Models for Stock Return Prediction

探索来自大语言模型的定量因素和新闻流表示的协同效应以预测股票收益率

Tian Guo, Emmanuel Hauptmann

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CL、cs.AI

AI总结 本文提出融合学习框架和混合模型,利用大语言模型生成的定量因素和新闻流表示,提升股票收益率预测和选择的准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17547 2025-11-25 eess.SP cs.AI cs.CV cs.HC cs.LG 62%

SYNAPSE: Synergizing an Adapter and Finetuning for High-Fidelity EEG Synthesis from a CLIP-Aligned Encoder

SYNAPSE: 适配器与微调协同实现高保真EEG信号到图像的合成

Jeyoung Lee, Hochul Kang

机构 * The Catholic University of Korea(韩国天主大学)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 SYNAPSE通过结合适配器与微调技术,实现了高保真EEG信号到图像的合成,提升了EEG数据的图像生成质量和跨受体泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.06843 2025-11-25 cs.CL cs.AI 62%

Integrating Cognitive Processing Signals into Language Models: A Review of Advances, Applications and Future Directions

将认知处理信号整合进语言模型:关于进展、应用与未来方向的综述

Angela Lopez-Cardona, Sebastian Idesis, Ioannis Arapakis

机构 * Telefónica Scientific Research(Telefónica科学研究院) Universitat Politècnica de Catalunya(加泰罗尼亚理工大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CL、cs.AI

AI总结 本文综述了将认知信号,特别是眼动数据整合到语言模型中的最新进展、应用及未来方向,探讨了其在提升模型性能和解决环境成本方面的潜力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14970 2025-11-20 cs.CV cs.AI cs.RO 62%

EGSA-PT:Edge-Guided Spatial Attention with Progressive Training for Monocular Depth Estimation and Segmentation of Transparent Objects

Gbenga Omotara, Ramy Farag, Seyed Mohamad Ali Tousi, G. N. DeSouza

机构 * Vision-Guided and Intelligent Robotics Lab (ViGIR) University of Missouri(视觉引导与智能机器人实验室(ViGIR)大学)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.00275 2025-11-20 cs.CV cs.AI 62%

AdCare-VLM: Towards a Unified and Pre-aligned Latent Representation for Healthcare Video Understanding

Md Asaduzzaman Jabin, Hanqi Jiang, Yiwei Li, Patrick Kaggwa, Eugene Douglass, Juliet N. Sekandi, Tianming Liu

机构 * University of Georgia(佐治亚大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 39th Conference on Neural Information Processing Systems (NeurIPS 2025) Workshop: 7th International Workshop on Large Scale Holistic Video Understanding: Toward Video Foundation Models

Journal ref Neural Information Processing Systems (NeurIPS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14601 2025-11-19 cs.CV cs.AI 62%

MRI Embeddings Complement Clinical Predictors for Cognitive Decline Modeling in Alzheimer's Disease Cohorts

Nathaniel Putera, Daniel Vilet Rodríguez, Noah Videcrantz, Julia Machnio, Mostafa Mehdipour Ghazi

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.AI

Comments Accepted at SPIE - Medical Imaging Conference 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13575 2025-11-18 cs.CV cs.AI 62%

Hierarchical Prompt Learning for Image- and Text-Based Person Re-Identification

Linhan Zhou, Shuang Li, Neng Dong, Yonghang Tai, Yafei Zhang, Huafeng Li

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV、cs.AI

Comments 9 pages, 4 figures, accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11730 2025-11-18 cs.CV cs.AI 62%

GROVER: Graph-guided Representation of Omics and Vision with Expert Regulation for Adaptive Spatial Multi-omics Fusion

Yongjun Xiao, Dian Meng, Xinlei Huang, Yanran Liu, Shiwei Ruan, Ziyue Qiao, Xubin Zheng

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 8 pages, 3 figures, Accepted to AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11622 2025-11-18 cs.LG cs.AI cs.CL 62%

Small Vocabularies, Big Gains: Pretraining and Tokenization in Time Series Models

Alexis Roger, Gwen Legate, Kashif Rasul, Yuriy Nevmyvaka, Irina Rish

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06836 2025-11-11 cs.CV cs.AI 62%

NeuroBridge: Bio-Inspired Self-Supervised EEG-to-Image Decoding via Cognitive Priors and Bidirectional Semantic Alignment

Wenjiang Zhang, Sifeng Wang, Yuwei Su, Xinyu Li, Chen Zhang, Suyu Zhong

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV、cs.AI

Comments AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19856 2025-11-11 cs.CV cs.AI 62%

RaGS: Unleashing 3D Gaussian Splatting from 4D Radar and Monocular Cues for 3D Object Detection

Xiaokai Bai, Chenxu Zhou, Lianqing Zheng, Si-Yuan Cao, Jianan Liu, Xiaohan Zhang, Yiming Li, Zhengzhuang Zhang, Hui-liang Shen

机构 * College of Information Science and Electronic Engineering, Zhejiang University(浙江大学信息科学与电子工程学院) College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院) School of Automotive Studies, Tongji University(同济大学汽车学院) Momoni AI, Gothenburg, Sweden(Momoni AI(瑞典哥德堡)) College of Energy Engineering, Zhejiang University(浙江大学能源工程学院)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00218 2025-11-04 cs.CV cs.AI 62%

DM-QPMNET: Dual-modality fusion network for cell segmentation in quantitative phase microscopy

Rajatsubhra Chakraborty, Ana Espinosa-Momox, Riley Haskin, Depeng Xu, Rosario Porras-Aguilar

机构 * College of Computing and Informatics, University of North Carolina at Charlotte, NC, USA(计算与信息学院,北卡罗来纳大学夏洛特分校) Department of Physics and Optical Science, University of North Carolina at Charlotte, NC, USA(物理与光学科学系,北卡罗来纳大学夏洛特分校) Center for TAIMing AI, University of North Carolina at Charlotte, NC, USA(TAIMing AI中心,北卡罗来纳大学夏洛特分校)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV、cs.AI

Comments 5 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21808 2025-10-28 cs.CV cs.AI 62%

Semantic Relation-Enhanced CLIP Adapter for Domain Adaptive Zero-Shot Learning

Jiaao Yu, Mingjie Han, Jinkun Jiang, Junyu Dong, Tao Gong, Man Lan

机构 * School of Computer Science and Technology, East China Normal University, China(东华大学计算机科学与技术学院) College of Computer Science and Technology, Ocean University of China, China(中国海洋大学计算机科学与技术学院)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV、cs.AI

Comments 5 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21794 2025-10-28 cs.CV cs.AI 62%

Token-Level Inference-Time Alignment for Vision-Language Models

Kejia Chen, Jiawen Zhang, Jiacong Hu, Kewei Gao, Jian Lou, Zunlei Feng, Mingli Song

机构 * Zhejiang University(浙江大学) Sun Yat-sen University(中山大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22038 2025-10-24 cs.CV cs.AI 62%

Balanced Token Pruning: Accelerating Vision Language Models Beyond Local Optimization

Kaiyuan Li, Xiaoyue Chen, Chen Gao, Yong Li, Xinlei Chen

机构 * Tsinghua Shenzhen International Graduate School(清华大学深圳国际研究生院) BNRist, Tsinghua University(清华大学北京研究院)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted by Neurips 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01411 2025-10-17 cs.CV cs.AI 62%

ViTA-PAR: Visual and Textual Attribute Alignment with Attribute Prompting for Pedestrian Attribute Recognition

Minjeong Park, Hongbeen Park, Jinkyu Kim

机构 * Department of Computer Science and Engineering, Korea University, Seoul 02841, Korea(计算机科学与工程系,韩国大学,首尔02841,韩国)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.AI

Comments Accepted to IEEE ICIP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21976 2025-10-16 cs.CV cs.AI 62%

Geo-R1: Improving Few-Shot Geospatial Referring Expression Understanding with Reinforcement Fine-Tuning

Zilun Zhang, Zian Guan, Tiancheng Zhao, Haozhan Shen, Tianyu Li, Yuxiang Cai, Zhonggen Su, Zhaojun Liu, Jianwei Yin, Xiang Li

机构 * College of Computer Science and Technology of Zhejiang University(浙江大学计算机科学与技术学院) Polytechnic Institute of Zhejiang University(浙江大学Polytechnic学院) Om AI Research(Om AI研究机构) Binjiang Research Institute of Zhejiang University(浙江大学滨江研究机构) School of Software Engineering of Zhejiang University(浙江大学软件工程学院) School of Mathematical Sciences of Zhejiang University(浙江大学数学科学学院) China Academy of Space Technology(中国航天科技研究院) University of Bristol(布里斯托大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12482 2025-10-15 cs.CV cs.AI 62%

A Text-Image Fusion Method with Data Augmentation Capabilities for Referring Medical Image Segmentation

Shurong Chai, Rahul Kumar JAIN, Rui Xu, Shaocong Mo, Ruibo Hou, Shiyu Teng, Jiaqing Liu, Lanfen Lin, Yen-Wei Chen

机构 * College of Information Science and Engineering(信息科学与工程学院) Ritsumeikan University(立命馆大学) Tiwaki Co., Ltd.(Tiwaki公司) Dalian University of Technology(大连理工大学) School of Software(软件学院) Zhejiang University(浙江大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏