arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-01-01 至 2026-01-01 共收录 11 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 11 篇

2506.06970 2026-01-01 cs.CV 89%

Guiding Cross-Modal Representations with MLLM Priors via Preference Alignment

通过偏好对齐引导跨模态表示

Pengfei Zhao, Rongbo Luan, Wei Zhang, Peng Wu, Sifeng He

机构 * Apple(苹果公司)

专题命中 多模态训练与对齐 :MLLM(title,abstract);cross-modal(title,abstract);multimodal(abstract);分类 cs.CV

AI总结 MAPLE通过利用多模态大语言模型的细粒度对齐先验,引导跨模态表示学习,提升细粒度检索性能。

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.23936 2026-01-01 cs.CV 88%

MGML: A Plug-and-Play Meta-Guided Multi-Modal Learning Framework for Incomplete Multimodal Brain Tumor Segmentation

MGML: 一种插件式元引导多模态学习框架用于不完整多模态脑肿瘤分割

Yulong Zou, Bo Liu, Cun-Jing Zheng, Yuan-ming Geng, Siyue Li, Qiankun Zuo, Shuihua Wang, Yudong Zhang, Jin Hong

专题命中 多模态训练与对齐 :multimodal(title,abstract);multi-modal(title,abstract);分类 cs.CV

AI总结 MGML框架通过元引导和一致性正则化模块提升不完整多模态脑肿瘤分割性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07307 2026-01-01 cs.CV cs.AI 84%

MCITlib: Multimodal Continual Instruction Tuning Library and Benchmark

MCITlib: 多模态持续指令微调库与基准

Haiyang Guo, Fei Zhu, Hongbo Zhao, Fanhu Zeng, Wenzhuo Liu, Shijie Ma, Da-Han Wang, Xu-Yao Zhang

机构 * School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences(中国科学院大学先进交叉学科学院) State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所多模态人工智能系统国家重点实验室) Centre for Artificial Intelligence and Robotics, Hong Kong Institute of Science and Innovation, Chinese Academy of Sciences(中国科学院香港创新科学研究院人工智能与机器人中心) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Fujian Key Laboratory of Pattern Recognition and Image Understanding, School of Computer and Information Engineering, Xiamen University of Technology(福建 pattern recognition and image understanding 工程学院,厦门大学科技学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 MCITlib提供多模态持续学习的库和基准,支持8种算法并评估3个基准,旨在解决灾难性遗忘和跨模态协调问题。

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.24324 2026-01-01 cs.LG cs.AI 83%

Empower Low-Altitude Economy: A Reliability-Aware Dynamic Weighting Allocation for Multi-modal UAV Beam Prediction

赋能低空经济:一种可靠性感知的动态权重分配用于多模态无人机波束预测

Haojin Li, Anbang Zhang, Chen Sun, Chenyuan Feng, Kaiqian Qu, Tony Q. S. Quek, Haijun Zhang

机构 * University of Science and Technology Beijing(北京科技大学) Sony China Research Laboratory(索尼中国研究院) School of Control Science and Engineering, Shandong University(山东大学控制科学与工程学院) Southeast University(东南大学) College of Computer Science, University of Exeter(埃克塞特大学计算机学院) Information Systems Technology and Design Pillar, Singapore University of Technology and Design(新加坡科技设计大学信息系统技术与设计系)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.AI

AI总结 本文提出SaM2B框架,通过可靠性感知的动态权重分配和跨模态对比学习,提升多模态无人机波束预测的准确性和鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.00742 2026-01-01 cs.CV cs.AI eess.IV 81%

Zoomer: Adaptive Image Focus Optimization for Black-box MLLM

Zoomer: 为黑盒大语言模型实现自适应图像聚焦优化

Jiaxu Qian, Chendong Wang, Yifan Yang, Chaoyun Zhang, Huiqiang Jiang, Xufang Luo, Yu Kang, Qingwei Lin, Anlan Zhang, Shiqi Jiang, Ting Cao, Tianjun Mao, Suman Banerjee, Guyue Liu, Saravan Rajmohan, Dongmei Zhang, Yuqing Yang, Qi Zhang, Lili Qiu

机构 * Microsoft(微软公司) Peking University(北京大学) University of Wisconsin Madison(威斯康星大学麦迪逊分校) University of Southern California(南加州大学)

专题命中 多模态训练与对齐 :MLLM(title);multimodal(abstract);分类 cs.CV、cs.AI

AI总结 Zoomer通过自适应图像聚焦优化提升黑盒MLLM的多模态理解能力,显著提升准确性并减少token使用

Comments TMLR accepted

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22748 2026-01-01 cs.CV 79%

TrimTokenator-LC: Towards Adaptive Visual Token Pruning for Large Multimodal Models with Long Contexts

TrimTokenator-LC: 向大型多模态模型长上下文的自适应视觉标记修剪迈进

Hao Zhang, Mengsi Lyu, Bo Huang, Yulong Ao, Yonghua Lin

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 TrimTokenator-LC通过自适应视觉标记修剪方法,在长上下文和多图像场景中有效减少视觉标记数量,同时保持性能。

Comments 17 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.24793 2026-01-01 cs.LG cs.NE 78%

Self-Supervised Neural Architecture Search for Multimodal Deep Neural Networks

多模态深度神经网络的自监督神经架构搜索

Shota Suzuki, Satoshi Ono

专题命中 多模态训练与对齐 :multimodal(title,abstract)

AI总结 本文提出了一种自监督学习方法,用于多模态深度神经网络的架构搜索,能够在未标记数据上有效设计DNN架构。

Journal ref IEICE Transactions on Information and Systems, Vol.E108.D, No. 6, pp. 640-643, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.24826 2026-01-01 cs.CV cs.AI 62%

Video and Language Alignment in 2D Systems for 3D Multi-object Scenes with Multi-Information Derivative-Free Control

用于多物体3D场景的2D系统中视频与语言对齐的多信息无导数控制

Jason Armitage, Rico Sennnrich

机构 * University of Zurich(苏黎世大学)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出了一种无导数优化方法,用于在多物体3D场景中实现视频与语言的对齐,通过在线适应物体遮挡和区分特征来提升跨模态任务性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.17221 2026-01-01 cs.CV 57%

DAVE: A VLM Vision Encoder for Document Understanding and Web Agents

DAVE: 一种用于文档理解与网络代理的视觉编码器

Brandon Huang, Hang Hua, Zhuoran Yu, Trevor Darrell, Rogerio Feris, Roei Herzig

机构 * MIT-IBM Watson AI Lab(MIT-IBM Watson AI实验室) UC Berkeley(加州大学伯克利分校) University of Wisconsin–Madison(威斯康星大学麦迪逊分校)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

AI总结 DAVE是一种专为文档理解和网络代理设计的视觉编码器,通过自监督和监督预训练结合模型融合策略,提升对文档和网络任务的适应性与性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.24284 2026-01-01 cs.RO cs.AI 57%

DRL-TH: Jointly Utilizing Temporal Graph Attention and Hierarchical Fusion for UGV Navigation in Crowded Environments

DRL-TH:联合利用时序图注意力和层次融合用于拥挤环境下的无人地面车辆导航

Ruitong Li, Lin Zhang, Yuenan Zhao, Chengxin Liu, Ran Song, Wei Zhang

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.AI

AI总结 DRL-TH通过结合时序图注意力和层次融合,提升无人地面车辆在拥挤环境中的导航性能和动态适应性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.23335 2026-01-01 cs.CV cs.LG 57%

Visual Language Hypothesis

视觉语言假说

Xiu Li

机构 * Independent Researcher(独立研究者)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

AI总结 该研究提出视觉理解依赖于语义语言的假说,并通过拓扑结构解释语义抽象和不变性的实现机制。

详情

展开后加载摘要…

URL PDF HTML 收藏