arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-02-05 至 2026-02-05 共收录 11 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 11 篇

2602.04231 2026-02-05 cs.RO 82%

GeoLanG: Geometry-Aware Language-Guided Grasping with Unified RGB-D Multimodal Learning

GeoLanG: 基于几何的语言引导抓取与统一RGB-D多模态学习

Rui Tang, Guankun Wang, Long Bai, Huxin Gao, Jiewen Lai, Chi Kit Ng, Jiazheng Wang, Fan Zhang, Hongliang Ren

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract)

AI总结 GeoLanG通过统一RGB-D多模态学习,结合深度引导几何模块和自适应通道集成,实现鲁棒的语言引导抓取,提升复杂环境中的抓取精度和泛化能力。

Comments IEEE ICRA 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.04016 2026-02-05 eess.SP cs.LG 82%

A Multi-Modal Foundational Model for Wireless Communication and Sensing

一种用于无线通信和传感的多模态基础模型

Vahid Yazdnian, Yasaman Ghasempour

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract)

AI总结 本文提出了一种多模态基础模型,通过物理指导的自监督预训练策略,实现无线通信和传感任务的稳健泛化与高效适应。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21436 2026-02-05 cs.LG cs.AI cs.CL cs.CV 82%

From Consistency to Complementarity: Aligned and Disentangled Multi-modal Learning for Time Series Understanding and Reasoning

从一致性到互补性:面向时间序列理解和推理的对齐与解缠多模态学习

Hang Ni, Weijia Zhang, Fei Wang, Zezhi Shao, Hao Liu

机构 * The Hong Kong University of Science(香港科学与技术大学) Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所) University of Chinese Academy of Sciences(中国科学院大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 MADI通过细粒度对齐和解缠交互提升多模态时间序列理解和推理能力,实现更精确的数值-视觉模态整合。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.04405 2026-02-05 cs.CV cs.MM 81%

Interactive Spatial-Frequency Fusion Mamba for Multi-Modal Image Fusion

交互式空间-频率融合Mamba用于多模态图像融合

Yixin Zhu, Long Lv, Pingping Zhang, Xuehu Liu, Tongdan Tang, Feng Tian, Weibing Sun, Huchuan Lu

机构 * School of Future Technology, Dalian University of Technology and the Key Laboratory of Data Science and Smart Education (Hainan Normal University), Ministry of Education(未来技术学院,大连理工大学和数据科学与智能教育关键实验室(海南师范大学),教育部) Affiliated Zhongshan Hospital of Dalian University(大连大学附属中山医院) School of Computer Science and Artificial Intelligence, Wuhan University of Technology(计算机科学与人工智能学院,武汉理工大学) Central Hospital of Dalian University of Technology(大连理工大学中心医院) School of Information and Communication Engineering, Dalian University of Technology(信息与通信工程学院,大连理工大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV、cs.MM

AI总结 本文提出交互式空间-频率融合Mamba框架,通过多尺度频率融合和交互式融合提升多模态图像融合性能。

Comments This work is accepted by IEEE Transactions on Image Processing. More modifications may be performed

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.16419 2026-02-05 cs.CL cs.CV 81%

Learning Domain Knowledge in Multimodal Large Language Models through Reinforcement Fine-Tuning

通过强化微调学习多模态大语言模型中的领域知识

Qinglong Cao, Yuntian Chen, Chao Ma, Xiaokang Yang

机构 * MoE Key Lab of Artificial Intelligence, AI Institute, Shanghai Jiao Tong University, Shanghai, China(人工智能大规模并行计算实验室,人工智能研究院,上海交通大学,上海)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL

AI总结 本文提出通过强化微调框架在优化层面整合领域知识,提升多模态大语言模型在专门领域任务中的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.04512 2026-02-05 q-bio.NC cs.AI 79%

BrainVista: Modeling Naturalistic Brain Dynamics as Multimodal Next-Token Prediction

BrainVista: 以多模态下一项令牌预测建模自然主义脑动力学

Xuanhua Yin, Runkai Zhao, Lina Yao, Weidong Cai

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

AI总结 BrainVista通过多模态自回归框架建模自然主义脑动力学,采用网络级令牌化器和空间混合头以解构系统特定动态并捕捉跨网络信息流,通过S2B掩码机制实现严格因果条件,提升fMRI编码性能。

Comments 17 pages, 7 figures, 11 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04670 2026-02-05 cs.AI 79%

Improving Multimodal Brain Encoding Model with Dynamic Subject-awareness Routing

改进多模态脑编码模型的动态主体感知路由

Xuanhua Yin, Runkai Zhao, Weidong Cai

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

AI总结 AFIRE和MIND通过动态主体感知路由提升多模态脑编码模型的性能,增强跨受试者泛化能力并实现可解释的专家模式。

Comments 7 pages, 4 figures, accepted by ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20977 2026-02-05 cs.CL 79%

Evaluating and Steering Modality Preferences in Multimodal Large Language Model

评估和引导多模态大语言模型中的模态偏好

Yu Zhang, Jinlong Ma, Yongshuai Hou, Xuefeng Bai, Kehai Chen, Yang Xiang, Jun Yu, Min Zhang

机构 * Harbin Institute of Technology, Shenzhen, China(哈尔滨工业大学(深圳)) Peng Cheng Laboratory, Shenzhen, China(鹏城实验室)

专题命中 多模态训练与对齐 :multimodal(title);multi-modal(abstract);分类 cs.CL

AI总结 本文提出MC²基准测试,通过受控证据冲突场景评估和引导多模态大语言模型的模态偏好,揭示了其可通过指令引导和潜在表示控制,并展示了通过表示工程方法提升多模态任务性能的潜力。

Comments Modality Preference

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.04565 2026-02-05 cs.CV 57%

Understanding Degradation with Vision Language Model

理解视觉退化

Guanzhou Lan, Chenyi Liao, Yuqi Yang, Qianli Ma, Zhigang Wang, Dong Wang, Bin Zhao, Xuelong Li

机构 * School of Artificial Intelligence, OPtics and ElectroNics (iOPEN), Northwestern Polytechnical University(人工智能学院、光学与电子学(iOPEN)、西北工业大学) Honors College, Northwestern Polytechnical University(荣誉学院、西北工业大学) Shanghai AI Laboratory(上海人工智能实验室) School of Artificial Intelligence, Shanghai Jiaotong University(人工智能学院、上海交通大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

AI总结 本文提出DU-VLM模型,通过结构化奖励强化学习解决视觉退化理解问题,引入大规模数据集并实现高保真图像恢复。

Comments 17 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.08420 2026-02-05 cs.RO 50%

LiDAR, GNSS and IMU Sensor Fine Alignment through Dynamic Time Warping to Construct 3D City Maps

通过动态时间规整实现LiDAR、GNSS和IMU传感器精细对齐以构建3D城市地图

Haitian Wang, Hezam Albaqami, Xinyu Wang, Muhammad Ibrahim, Zainy M. Malakan, Abdullah M. Algamdi, Mohammed H. Alghamdi, Ajmal Mian

机构 * University of Western Australia(西澳大学) University of Jeddah(朱尔法大学) Umm Al-Qura University(乌姆·阿勒·盖拉大学) King Khalid University(国王·卡利德大学)

专题命中 多模态训练与对齐 :multimodal(abstract)

AI总结 本文提出通过动态时间规整实现LiDAR、GNSS和IMU传感器精细对齐,以提高3D城市地图构建的精度和一致性。

Comments This paper has been submitted to IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing (JSTARS) and is currently under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.02542 2026-02-05 cs.LG cs.CR 50%

OverThink: Slowdown Attacks on Reasoning LLMs

OverThink: 对推理大语言模型的减速攻击

Abhinav Kumar, Jaechul Roh, Ali Naseh, Marzena Karpinska, Mohit Iyyer, Amir Houmansadr, Eugene Bagdasarian

机构 * University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校) Simon Fraser University(西蒙弗雷泽大学) University of Maryland, College Park(马里兰大学学院公园分校)

专题命中 多模态训练与对齐 :multi-modal(abstract)

AI总结 OverThink攻击通过注入伪装推理问题,迫使推理大语言模型消耗更多token,从而增加延迟和成本,同时探讨了其防御和影响。

详情

展开后加载摘要…

URL PDF HTML 收藏