arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6903 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6903 篇

2411.08533 2026-04-29 cs.RO cs.AI 74%

ACROSS: A Deformation-Based Cross-Modal Representation for Robotic Tactile Perception

ACROSS: 一种基于变形的跨模态表示用于机器人触觉感知

Wadhah Zai El Amri, Malte Kuhlmann, Nicolás Navarro-Guerrero

机构 * L3S Research Center(L3S研究所以)

专题命中 多模态训练与对齐 :cross-modal(title);分类 cs.AI

AI总结 本文提出ACROSS框架,通过利用传感器变形信息实现触觉数据跨传感器转换,将BioTac信号转换为DIGIT传感器数据,解决现有数据集在新设备上的应用问题。

Comments Accepted to 2025 IEEE Conference on Robotics and Automation (ICRA 2025). arXiv admin note: text overlap with arXiv:2410.14310

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.21573 2026-04-24 cs.CV q-bio.QM 74%

CHRep: Cross-modal Histology Representation and Post-hoc Calibration for Spatial Gene Expression Prediction

CHRep: 跨模态组织切片表示与事后校准用于空间基因表达预测

Changfan Wang, Xinran Wang, Donghai Liu, Fei Su, Lulu Sun, Zhicheng Zhao, Zhu Meng

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Beijing Key Laboratory of Network System and Network Culture(北京网络系统与网络文化重点实验室) Peking University Third Hospital(北京大学第三医院)

专题命中 多模态训练与对齐 :cross-modal(title);分类 cs.CV

AI总结 CHRep通过两阶段框架提升组织切片到基因表达的预测鲁棒性,结合结构感知表示学习与事后校准,提升滑片级变化下的稳定性与准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.18790 2026-04-22 cs.CV 74%

EfficientPENet: Real-Time Depth Completion from Sparse LiDAR via Lightweight Multi-Modal Fusion

EfficientPENet:通过轻量多模态融合实现稀疏LiDAR的实时深度补全

Johny J. Lopez, Md Meftahul Ferdaus, Mahdi Abdelguerfi, Anton Netchaev, Steven Sloan, Ken Pathak, Kendall N. Niles

机构 * Canizaro Livingston Gulf States Center for Environmental Informatics, the University of New Orleans, New Orleans, USA(Canizaro Livingston Gulf States环境信息中心,新奥尔良大学,美国新奥尔良) US Army Corps of Engineers, Engineer Research and Development Center, Vicksburg, Mississippi, USA(美国陆军工程兵团,工程师研究与发展中心,密西西比州维克斯堡,美国)

专题命中 多模态训练与对齐 :multi-modal(title);分类 cs.CV

AI总结 本文提出EfficientPENet,通过轻量级多模态融合网络,在稀疏LiDAR和RGB图像上实现实时深度补全,具有更少的参数和更高的速度,同时保持竞争力的精度。

Comments This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.18231 2026-04-21 cs.LG cs.AI 74%

Rethinking Cross-Modal Fine-Tuning: Optimizing the Interaction Between Feature Alignment and Target Fitting

重新思考跨模态微调:优化特征对齐与目标拟合之间的交互

Trong Khiem Tran, Manh Cuong Dao, Phi Le Nguyen, Thao Nguyen Truong, Trong Nghia Hoang

机构 * Washington State University(华盛顿州立大学) National University of Singapore(新加坡国立大学) National Institute of Advanced Industrial Science and Technology(国家先进工业科学与技术研究院) Hanoi University of Science and Technology(河内科学技术大学)

专题命中 多模态训练与对齐 :cross-modal(title);分类 cs.AI

AI总结 本文提出一种原理性框架,通过特征-标签扭曲概念解释特征对齐与目标拟合的交互,建立目标误差的可证明泛化界,提升跨模态微调性能。

Comments Accepted AISTATS 20226

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.10347 2026-04-14 cs.CV 74%

Multi-modal, multi-scale representation learning for satellite imagery analysis just needs a good ALiBi

多模态、多尺度表示学习用于卫星影像分析只需一个良好的ALiBi

Patrick Kage, Pavlos Andreadis

专题命中 多模态训练与对齐 :multi-modal(title);分类 cs.CV

AI总结 本文提出Scale-ALiBi机制,通过空间编码偏置提升多尺度多模态卫星影像表示学习效果,并在GEO-Bench基准上取得改进。

Comments Originally appeared at the 4th Space Imaging Workshop at the Georgia Institute of Technology, October 7-9, 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08018 2026-04-02 cs.CV 74%

Missing No More: Dictionary-Guided Cross-Modal Image Fusion under Missing Infrared

不再缺失:字典引导的跨模态图像融合在红外缺失情况下的应用

Yafei Zhang, Meng Ma, Huafeng Li, Yu Liu

机构 * Faculty of Information Engineering and Automation, Kunming University of Science and Technology(昆明理工大学信息工程与自动化学院) Department of Biomedical Engineering, Hefei University of Technology(合肥工业大学生物医学工程系)

专题命中 多模态训练与对齐 :cross-modal(title);分类 cs.CV

AI总结 本文提出了一种基于共享卷积字典的字典引导框架,解决红外缺失时的跨模态图像融合问题,通过联合字典学习、视觉引导红外推断和自适应融合方法提升感知质量和下游检测性能。

Comments This paper has been accepted by CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16216 2026-03-31 cs.AI cs.LG 74%

FlipVQA: Scaling Multi-modal Instruction Tuning via Textbook-to-Knowledge Synthesis

FlipVQA: 通过教科书到知识合成实现多模态指令微调的扩展

Zhen Hao Wong, Jingwen Deng, Yuzhao Wang, Wenkai Yu, Jihao Huang, Runming He, Chengyu Shen, Hao Liang, Wentao Zhang

机构 * Peking University(北京大学) Zhongguancun Academy(中关村学院)

专题命中 多模态训练与对齐 :multi-modal(title);分类 cs.AI

AI总结 本文提出FlipVQA-Miner自动化流程,解决OCR文档中的长距离逻辑依赖和跨页断续问题,构建了包含83K问答对的FlipVQA-83K数据集,在保持高结构保真度的同时节省50倍成本,提升了模型推理能力和跨领域泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.11770 2026-03-27 cs.CL cs.CY cs.SI 74%

The Value of Nothing: Multimodal Extraction of Human Values Expressed by TikTok Influencers

nothing的价值:通过TikTok影响者表达的人类价值观多模态提取

Alina Starovolsky-Shitrit, Alon Neduva, Naama Appel Doron, Itamar Gafni, Ella Daniel, Oren Tsur

专题命中 多模态训练与对齐 :multimodal(title);分类 cs.CL

AI总结 本文研究通过TikTok影响者上传的视频提取隐含价值观,采用两阶段方法优于直接提取,使用少量样本的大语言模型在两个阶段表现更优,首次公开TikTok视频价值观标注数据集。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06186 2026-03-09 cs.CV 74%

SpaCRD: Multimodal Deep Fusion of Histology and Spatial Transcriptomics for Cancer Region Detection

SpaCRD:多模态深度融合组织学与空间转录组学用于癌症区域检测

Shuailin Xue, Jun Wan, Lihua Zhang, Wenwen Min

专题命中 多模态训练与对齐 :multimodal(title);分类 cs.CV

AI总结 SpaCRD通过多模态深度融合组织学与空间转录组学数据,实现跨样本、平台和批次的癌症区域检测,优于现有八种方法。

Comments Accepted by AAAI-2026-Oral

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02250 2026-03-04 cs.SD eess.AS 74%

SGPA: Spectrogram-Guided Phonetic Alignment for Feasible Shapley Value Explanations in Multimodal Large Language Models

SGPA: 基于频谱的语音对齐用于多模态大语言模型中可行的谢普利值解释

Paweł Pozorski, Jakub Muszyński, Maria Ganzha

机构 * Warsaw University of Technology(华沙技术大学)

专题命中 多模态训练与对齐 :multimodal(title);分类 eess.AS

AI总结 SGPA通过结合连接主义时间分类和频谱边界细化,实现了多模态大语言模型中可行的音频解释,显著减少了模型评估次数并保持了全局轮廓。

Comments Submitted for admission in Interspeech 2026 conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01210 2026-03-03 cs.CV cs.RO 74%

OmniVLA: Physically-Grounded Multimodal VLA with Unified Multi-Sensor Perception for Robotic Manipulation

OmniVLA:具有统一多传感器感知的物理基础多模态VLA

Heyu Guo, Shanmu Wang, Ruichun Ma, Shiqi Jiang, Yasaman Ghasempour, Omid Abari, Baining Guo, Lili Qiu

机构 * Princeton University(普林斯顿大学) University of California, Los Angeles(加州大学洛杉矶分校) Microsoft Research Asia(微软亚洲研究院)

专题命中 多模态训练与对齐 :multimodal(title);分类 cs.CV

AI总结 OmniVLA通过整合多种传感器模态,提升机器人操作的感知能力与任务成功率。

Comments Accepted by ICRA'26

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18632 2026-02-25 cs.CV 74%

Decouple, Reorganize, and Fuse: A Multimodal Framework for Cancer Survival Prediction

解耦、重组与融合:一种多模态框架用于癌症生存预测

Huayi Wang, Haochao Ying, Yuyang Xu, Qibo Qiu, Cheng Zhang, Danny Z. Chen, Ying Sun, Jian Wu

机构 * College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院) State Key Laboratory of Transvascular Implantation Devices of the Second Affiliated Hospital, Zhejiang University School of Medicine(浙江大学医学院第二附属医院血管植入设备国家重点实验室) Transvascular Implantation Devices Research Institute(血管植入设备研究院) Zhejiang Key Laboratory of Medical Imaging Artificial Intelligence(浙江省医学影像人工智能重点实验室) China Mobile (Zhejiang) Research & Innovation Institute(中国移动(浙江)研究院)

专题命中 多模态训练与对齐 :multimodal(title);分类 cs.CV

AI总结 本文提出DeReF框架,通过解耦-重组-融合策略提升多模态癌症生存预测的准确性与泛化能力。

Comments 13 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09541 2026-02-11 cs.CV 74%

Scalpel: Fine-Grained Alignment of Attention Activation Manifolds via Mixture Gaussian Bridges to Mitigate Multimodal Hallucination

Scalpel: 通过混合高斯桥梁实现细粒度注意力激活流形对齐以缓解多模态幻觉

Ziqiang Shi, Rujie Liu, Shanshan Yu, Satoshi Munakata, Koichi Shirahata

机构 * Fujitsu Research & Development Center Co.,LTD.(Fujitsu 研究与开发中心有限公司) Fujitsu Limited(Fujitsu 有限公司)

专题命中 多模态训练与对齐 :multimodal(title);分类 cs.CV

AI总结 Scalpel通过高斯混合模型和熵最优传输减少多模态幻觉,实现注意力激活流形的细粒度对齐,提升视觉-语言模型的输出一致性。

Comments WACV 2026 (It was accepted in the first round, with an acceptance rate of 6%.)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.09578 2026-01-15 cs.RO cs.CV 74%

Multimodal Signal Processing For Thermo-Visible-Lidar Fusion In Real-time 3D Semantic Mapping

多模态信号处理用于热-可见-激光雷达融合的实时3D语义制图

Jiajun Sun, Yangyi Ou, Haoyuan Zheng, Chao yang, Yue Ma

机构 * College of Mechatronics and Control Engineering, Shenzhen University(深圳大学机械与控制工程学院) School of Robotics, Xi’an-Jiaotong Liverpool University(西安交通大学利物浦大学机器人学院)

专题命中 多模态训练与对齐 :multimodal(title);分类 cs.CV

AI总结 本文提出通过多模态信号处理融合热、可见和激光雷达数据,提升实时3D语义制图的精度与语义理解能力,适用于灾害评估和工业维护等场景。

Comments 5 pages,7 figures. Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.20892 2025-12-25 cs.CV 74%

Beyond Weight Adaptation: Feature-Space Domain Injection for Cross-Modal Ship Re-Identification

超越权重适应:基于特征空间的域注入用于跨模态船舶重识别

Tingfeng Xian, Wenlve Zhou, Zhiheng Zhou, Zhelin Li

机构 * School of Electronic and Information Engineering, South China University of Technology(华南理工大学电子与信息学院) Key Laboratory of Big Data and Intelligent Robot, Ministry of Education, South China University of Technology(大数据与智能机器人教育部重点实验室)

专题命中 多模态训练与对齐 :cross-modal(title);分类 cs.CV

AI总结 本文提出基于特征空间的域注入方法,解决跨模态船舶重识别中的模态差异问题,通过轻量级模型提升性能,实现SOTA效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.20561 2025-12-24 cs.CV 74%

FlashVLM: Text-Guided Visual Token Selection for Large Multimodal Models

FlashVLM: 大规模多模态模型中的文本引导视觉令牌选择

Kaitong Cai, Jusheng Zhang, Jing Yang, Yijia Fan, Pengtao Xie, Jian Wang, Keze Wang

机构 * Sun Yat-sen University(中山大学) University of California, San Diego(加州大学圣地亚哥分校) Snap Inc.(Snap公司)

专题命中 多模态训练与对齐 :multimodal(title);分类 cs.CV

AI总结 FlashVLM通过文本引导的视觉令牌选择,实现了高效压缩和高准确率,在大规模多模态模型中表现出色。

Comments Under submission

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02972 2025-12-03 cs.CV cs.RO 74%

BEVDilation: LiDAR-Centric Multi-Modal Fusion for 3D Object Detection

BEVDilation:以LiDAR为中心的多模态融合用于3D目标检测

Guowen Zhang, Chenhang He, Liyi Chen, Lei Zhang

专题命中 多模态训练与对齐 :multi-modal(title);分类 cs.CV

AI总结 BEVDilation提出以LiDAR为中心的多模态融合方法,通过稀疏体素扩张和语义引导BEV扩张模块提升3D目标检测性能,有效缓解深度误差带来的空间错位问题。

Comments Accept by AAAI26

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11706 2025-11-21 cs.LG cs.CV 74%

Context-Aware Multimodal Representation Learning for Spatio-Temporally Explicit Environmental Modelling

面向情境的多模态表示学习用于时空明确的环境建模

Julia Peters, Karin Mora, Miguel D. Mahecha, Chaonan Ji, David Montero, Clemens Mosig, Guido Kraemer

机构 * Environmental Data Science and Remote Sensing Group(环境数据科学与遥感小组) Institute for Earth System Science and Remote Sensing(地球系统科学与遥感研究所) Leipzig University(莱比锡大学) German Centre for Integrative Biodiversity Research (iDiv)(整合生物多样性研究德国中心(iDiv))

专题命中 多模态训练与对齐 :multimodal(title);分类 cs.CV

AI总结 本文提出一种面向情境的多模态表示学习框架,整合不同地球观测模态以高时空分辨率建模环境,提升生态分析的精度与效率。

Comments 10 pages (incliding 2 pages of references), 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15139 2025-11-20 q-bio.GN cs.AI cs.LG 74%

CASPER: Cross-modal Alignment of Spatial and single-cell Profiles for Expression Recovery

Amit Kumar, Maninder Kaur, Raghvendra Mall, Sukrit Gupta

机构 * Department of Computer Science \& Engineering, Indian Institute of Technology Ropar, India Qatar Computing Research Institute, Hamad Bin Khalifa University, Doha, Qatar. Department of Biomedical Engineering, Indian Institute of Technology Ropar, India

专题命中 多模态训练与对齐 :cross-modal(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.13076 2025-10-16 cs.CV 74%

Hints of Prompt: Enhancing Visual Representation for Multimodal LLMs in Autonomous Driving

Hao Zhou, Zhanning Gao, Zhili Chen, Maosheng Ye, Qifeng Chen, Tongyi Cao, Honggang Qi

机构 * University of Chinese Academy of Sciences(中国科学院大学) The Hong Kong University of Science and Technology(香港科技大学)

专题命中 多模态训练与对齐 :multimodal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00419 2025-09-03 cs.CV 74%

LightVLM: Acceleraing Large Multimodal Models with Pyramid Token Merging and KV Cache Compression

Lianyu Hu, Fanhua Shang, Wei Feng, Liang Wan

机构 * College of Intelligence and Computing, Tianjin University(智能与计算学院,天津大学)

专题命中 多模态训练与对齐 :multimodal(title);分类 cs.CV

Comments EMNLP2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16224 2025-08-28 cs.CV 74%

LDRFusion: A LiDAR-Dominant multimodal refinement framework for 3D object detection

Jijun Wang, Yan Wu, Yujian Mo, Junqiao Zhao, Jun Yan, Yinghao Hu

机构 * School of Computer Science and Technology, Tongji University, Shanghai 201804, China(计算机科学与技术学院,同济大学,上海) School of Electronics and Information Engineering, Tongji University, Shanghai 201804, China(电子信息工程学院,同济大学,上海)

专题命中 多模态训练与对齐 :multimodal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.16873 2025-08-20 cs.CV 74%

ContrastAlign: Toward Robust BEV Feature Alignment via Contrastive Learning for Multi-Modal 3D Object Detection

Ziying Song, Hongyu Pan, Feiyang Jia, Yongchang Zhang, Lin Liu, Lei Yang, Shaoqing Xu, Peiliang Wu, Caiyan Jia, Zheng Zhang, Yadan Luo

机构 * School of Computer Science & Technology, Beijing Key Laboratory of Traffic Data Mining and Embodied Intelligence, Beijing Jiaotong University(计算机科学与技术学院,北京交通大数据挖掘与具身智能重点实验室,北京交通大学) Horizon Robotics Nanyang Technological University(南洋理工大学) University of Macau(澳门大学) School of Information Science and Engineering, Yanshan University(信息科学与工程学院,燕山大学) School of Computer Science and Technology, Harbin Institute of Technology(计算机科学与技术学院,哈尔滨理工大学) School of Information Technology and Electrical Engineering, The University of Queensland(信息技术与电气工程学院,昆士兰大学)

专题命中 多模态训练与对齐 :multi-modal(title);分类 cs.CV

Comments 12 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21121 2025-06-27 cs.CV cs.RO 74%

GoIRL: Graph-Oriented Inverse Reinforcement Learning for Multimodal Trajectory Prediction

Muleilan Pei, Shaoshuai Shi, Lu Zhang, Peiliang Li, Shaojie Shen

机构 * Department of Electronic and Computer Engineering, Hong Kong University of Science and Technology, Hong Kong, China(香港理工大学电子与计算机工程系) Voyager Research, Didi Chuxing, China(维杰研究,滴滴出行,中国) Zhuoyu Technology Co., Ltd., Shenzhen, China(珠宇科技有限公司,深圳,中国)

专题命中 多模态训练与对齐 :multimodal(title);分类 cs.CV

Comments Accepted by ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.03494 2025-05-07 cs.CV 74%

UPMAD-Net: A Brain Tumor Segmentation Network with Uncertainty Guidance and Adaptive Multimodal Feature Fusion

Zhanyuan Jia, Ni Yao, Danyang Sun, Chuang Han, Yanting Li, Jiaofen Nan, Fubao Zhu, Chen Zhao, Weihua Zhou

专题命中 多模态训练与对齐 :multimodal(title);分类 cs.CV

Comments 21 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.03284 2025-05-07 cs.CV cs.RO 74%

OccCylindrical: Multi-Modal Fusion with Cylindrical Representation for 3D Semantic Occupancy Prediction

Zhenxing Ming, Julie Stephany Berrio, Mao Shan, Yaoqi Huang, Hongyu Lyu, Nguyen Hoang Khoi Tran, Tzu-Yun Tseng, Stewart Worrall

机构 * Australian Centre for Robotics (ACFR) at the University of Sydney(澳大利亚机器人中心(ACFR)于悉尼大学)

专题命中 多模态训练与对齐 :multi-modal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.19271 2025-04-29 cs.CV 74%

Leveraging Multi-Modal Saliency and Fusion for Gaze Target Detection

Athul M. Mathew, Arshad Ali Khan, Thariq Khalid, Faroq AL-Tam, Riad Souissi

机构 * Elm Company(埃尔姆公司)

专题命中 多模态训练与对齐 :multi-modal(title);分类 cs.CV

Comments accepted at NeurIPS 2023 Gaze Meets ML Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.22285 2025-03-31 cs.CV 74%

RUNA: Object-level Out-of-Distribution Detection via Regional Uncertainty Alignment of Multimodal Representations

Bin Zhang, Jinggang Chen, Xiaoyang Qu, Guokuan Li, Kai Lu, Jiguang Wan, Jing Xiao, Jianzong Wang

专题命中 多模态训练与对齐 :multimodal(title);分类 cs.CV

Comments 9 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.11409 2025-03-17 cs.CV cs.RO 74%

LuSeg: Efficient Negative and Positive Obstacles Segmentation via Contrast-Driven Multi-Modal Feature Fusion on the Lunar

Shuaifeng Jiao, Zhiwen Zeng, Zhuoqun Su, Xieyuanli Chen, Zongtan Zhou, Huimin Lu

专题命中 多模态训练与对齐 :multi-modal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.11848 2025-03-14 cs.CV 74%

GraphBEV: Towards Robust BEV Feature Alignment for Multi-Modal 3D Object Detection

Ziying Song, Lei Yang, Shaoqing Xu, Lin Liu, Dongyang Xu, Caiyan Jia, Feiyang Jia, Li Wang

专题命中 多模态训练与对齐 :multi-modal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏