arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-01-01 至 2026-01-01 共收录 57 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 15 篇

2512.24013 2026-01-01 cs.CV 57%

Bridging the Perception-Cognition Gap:Re-engineering SAM2 with Hilbert-Mamba for Robust VLM-based Medical Diagnosis

弥合感知-认知鸿沟:通过希尔伯特-马amba重新工程SAM2以实现鲁棒的基于VLM的医学诊断

Hao Wu, Hui Li, Yiyun Su

机构 * Southern University of Science and Technology(南方科技大学) School of Informatics, Xiamen University(厦门大学信息学院) Rutgers University(罗格斯大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

AI总结 通过重新设计SAM2架构,引入希尔伯特空间填充曲线和希尔伯特-马amba交叉注意力机制,提升基于VLM的医学图像分割与诊断准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21684 2026-01-01 cs.CV 57%

SlideChain: Semantic Provenance for Lecture Understanding via Blockchain Registration

SlideChain: 通过区块链注册实现讲座理解的语义溯源

Md Motaleb Hossen Manik, Md Zabirul Islam, Ge Wang

机构 * Department of Computer Science Rensselaer Polytechnic Institute(计算机科学系罗切斯特理工学院) Department of Biomedical Engineering Rensselaer Polytechnic Institute(生物医学工程系罗切斯特理工学院)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

AI总结 SlideChain通过区块链注册实现讲座理解的语义溯源,提供可验证的完整性与持久基准,解决多模态教育内容的可复现性和审计性问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23105 2026-01-01 cs.CV 57%

Towards Comprehensive Interactive Change Understanding in Remote Sensing: A Large-scale Dataset and Dual-granularity Enhanced VLM

迈向遥感中全面交互变化理解:一个大规模数据集和双粒度增强VLM

Junxiao Xue, Quan Deng, Xuecheng Wu, Kelu Yao, Xinyi Yin, Fei Yu, Wei Zhou, Yanfei Zhong, Yang Liu, Dingkang Yang

机构 * Research Center for Space Computing System, Zhejiang Lab(浙江省实验室空间计算系统研究中心) Hangzhou Institute for Advanced Study, University of Chinese Academy of Sciences(中国科学院大学杭州高等研究院) School of Computer Science and Technology, Xi’an Jiaotong University(西安交通大学计算机科学与技术学院) School of Cyber Science and Engineering, Zhengzhou University(郑州大学网络科学与工程学院) Liaoning University of Technology(辽宁科技学院) School of Computer Science and Informatics, Cardiff University(卡迪夫大学计算机科学与信息学院) State Key Laboratory of Information Engineering in Surveying, Mapping and Remote Sensing (LIESMARS), Wuhan University(武汉测绘遥感与信息工程国家重点实验室(LIESMARS)) College of Electronic and Information Engineering, Tongji University(同济大学电子与信息工程学院) College of Intelligent Robotics and Advanced Manufacturing, Fudan University(复旦大学智能机器人与先进制造学院)

专题命中 多模态评测 :cross-modal(abstract);分类 cs.CV

AI总结 本文提出ChangeIMTI数据集和ChangeVG模型,通过双粒度增强方法提升遥感图像变化理解的准确性和交互性。

Comments Junxiao Xue, Quan Deng, and Xuecheng Wu deserve equal contributions

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17597 2026-01-01 cs.CV 57%

BCWildfire: A Long-term Multi-factor Dataset and Deep Learning Benchmark for Boreal Wildfire Risk Prediction

BCWildfire: 一种长期多因素数据集和深度学习基准,用于北极火灾风险预测

Zhengsen Xu, Sibo Cheng, Lanying Wang, Hongjie He, Wentao Sun, Jonathan Li, Lincoln Linlin Xu

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

AI总结 BCWildfire数据集和基准为长期多因素火灾风险预测提供了深度学习评估框架,包含38个协变量和多种模型评估。

Comments This paper has been accepted by AAAI-26

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 多模态Agent 6 篇

2512.23906 2026-01-01 eess.SP cs.AI cs.CV cs.LG 81%

A multimodal Transformer for InSAR-based ground deformation forecasting with cross-site generalization across Europe

基于InSAR的地面形变预测的多模态Transformer及欧洲跨站点泛化

Wendong Yao, Binhua Huang, Soumyabrata Dev

机构 * The ADAPT SFI Research Centre, Dublin, Ireland(ADAPT SFI研究机构,都柏林,爱尔兰) School of Computer Science, University College Dublin, Belfield, Dublin, Ireland(计算机科学学院,都柏林大学学院,贝尔菲德,都柏林,爱尔兰)

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 本文提出多模态Transformer模型,用于基于InSAR的地面形变预测,实现欧洲跨站点泛化,提升预测精度和稳定性。

Comments submitted to ISPRS Journal of Photogrammetry and Remote Sensing for review

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22998 2026-01-01 cs.AI 79%

TIM-PRM: Verifying multimodal reasoning with Tool-Integrated PRM

TIM-PRM:利用工具集成PRM验证多模态推理

Peng Kuang, Xiangxiang Wang, Wentao Liu, Jian Dong, Kaidi Xu

机构 * Zhejiang University(浙江大学) iFLYTEK AI Research Institute(iFLYTEK人工智能研究院) City University of Hong Kong(香港城市大学)

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.AI

AI总结 TIM-PRM通过工具集成PRM验证多模态推理,有效解决视觉幻觉和逻辑不一致问题,实验表明其在性能和可解释性上优于现有模型。

Comments 12 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.24023 2026-01-01 cs.CV cs.AI 79%

RSAgent: Learning to Reason and Act for Text-Guided Segmentation via Multi-Turn Tool Invocations

RSAgent: 通过多轮工具调用学习推理与行动进行文本引导的分割

Xingqi He, Yujie Zhang, Shuyong Gao, Wenjie Li, Lingyi Hong, Mingxi Chen, Kaixun Jiang, Jiyuan Fu, Wenqiang Zhang

机构 * Shanghai Key Lab of Intelligent Information Processing, College of Computer Science and Artificial Intelligence, Fudan University(上海智能信息处理关键实验室,计算机科学与人工智能学院,复旦大学) Artificial Intelligence, Fudan University(人工智能,复旦大学)

专题命中 多模态Agent :multimodal(abstract);MLLM(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 RSAgent通过多轮工具调用实现文本引导分割的推理与行动,采用两阶段框架提升分割性能,达到领域内和领域外基准的最先进水平。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.24404 2026-01-01 cs.LG cs.CV 57%

Lifting Vision: Ground to Aerial Localization with Reasoning Guided Planning

提升视觉:基于推理引导的地面到空中定位

Soham Pahari, M. Srinivas

机构 * School of Computer Science, UPES(UPES计算机科学学院) Department of CS&E, NIT Warangal(NIT Warangal计算机科学与工程系)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV

AI总结 本文提出ViReLoc框架,通过视觉推理实现地面到空中定位,无需依赖GPS数据,提升空间推理和跨视角检索性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.17821 2026-01-01 cs.RO cs.AI cs.LG 57%

CAML: Collaborative Auxiliary Modality Learning for Multi-Agent Systems

CAML: 多智能体系统中的协同辅助模态学习

Rui Liu, Yu Shen, Peng Gao, Pratap Tokekar, Ming Lin

机构 * University of Maryland, College Park(马里兰大学 College Park 分校) Adobe Research(Adobe 研究院) North Carolina State University(北卡罗来纳州立大学)

专题命中 多模态Agent :multi-modal(abstract);分类 cs.AI

AI总结 CAML通过多智能体协作和共享多模态数据,提升多模态学习在资源受限环境下的性能,实现事故检测和语义分割的显著改进。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.24645 2026-01-01 cs.SD 50%

AudioFab: Building A General and Intelligent Audio Factory through Tool Learning

通过工具学习构建通用且智能的音频工厂

Cheng Zhu, Jing Han, Qianshuai Xue, Kehan Wang, Huan Zhao, Zixing Zhang

机构 * Hunan University(湖南大学) University of Cambridge(剑桥大学)

专题命中 多模态Agent :multimodal(abstract)

AI总结 AudioFab通过工具学习构建通用智能音频处理框架,提供模块化设计和自然语言接口,提升音频任务的效率和准确性。

Journal ref ACM Multimedia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 多模态训练与对齐 11 篇

2506.06970 2026-01-01 cs.CV 89%

Guiding Cross-Modal Representations with MLLM Priors via Preference Alignment

通过偏好对齐引导跨模态表示

Pengfei Zhao, Rongbo Luan, Wei Zhang, Peng Wu, Sifeng He

机构 * Apple(苹果公司)

专题命中 多模态训练与对齐 :MLLM(title,abstract);cross-modal(title,abstract);multimodal(abstract);分类 cs.CV

AI总结 MAPLE通过利用多模态大语言模型的细粒度对齐先验,引导跨模态表示学习,提升细粒度检索性能。

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.23936 2026-01-01 cs.CV 88%

MGML: A Plug-and-Play Meta-Guided Multi-Modal Learning Framework for Incomplete Multimodal Brain Tumor Segmentation

MGML: 一种插件式元引导多模态学习框架用于不完整多模态脑肿瘤分割

Yulong Zou, Bo Liu, Cun-Jing Zheng, Yuan-ming Geng, Siyue Li, Qiankun Zuo, Shuihua Wang, Yudong Zhang, Jin Hong

专题命中 多模态训练与对齐 :multimodal(title,abstract);multi-modal(title,abstract);分类 cs.CV

AI总结 MGML框架通过元引导和一致性正则化模块提升不完整多模态脑肿瘤分割性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07307 2026-01-01 cs.CV cs.AI 84%

MCITlib: Multimodal Continual Instruction Tuning Library and Benchmark

MCITlib: 多模态持续指令微调库与基准

Haiyang Guo, Fei Zhu, Hongbo Zhao, Fanhu Zeng, Wenzhuo Liu, Shijie Ma, Da-Han Wang, Xu-Yao Zhang

机构 * School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences(中国科学院大学先进交叉学科学院) State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所多模态人工智能系统国家重点实验室) Centre for Artificial Intelligence and Robotics, Hong Kong Institute of Science and Innovation, Chinese Academy of Sciences(中国科学院香港创新科学研究院人工智能与机器人中心) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Fujian Key Laboratory of Pattern Recognition and Image Understanding, School of Computer and Information Engineering, Xiamen University of Technology(福建 pattern recognition and image understanding 工程学院,厦门大学科技学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 MCITlib提供多模态持续学习的库和基准,支持8种算法并评估3个基准,旨在解决灾难性遗忘和跨模态协调问题。

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.24324 2026-01-01 cs.LG cs.AI 83%

Empower Low-Altitude Economy: A Reliability-Aware Dynamic Weighting Allocation for Multi-modal UAV Beam Prediction

赋能低空经济:一种可靠性感知的动态权重分配用于多模态无人机波束预测

Haojin Li, Anbang Zhang, Chen Sun, Chenyuan Feng, Kaiqian Qu, Tony Q. S. Quek, Haijun Zhang

机构 * University of Science and Technology Beijing(北京科技大学) Sony China Research Laboratory(索尼中国研究院) School of Control Science and Engineering, Shandong University(山东大学控制科学与工程学院) Southeast University(东南大学) College of Computer Science, University of Exeter(埃克塞特大学计算机学院) Information Systems Technology and Design Pillar, Singapore University of Technology and Design(新加坡科技设计大学信息系统技术与设计系)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.AI

AI总结 本文提出SaM2B框架,通过可靠性感知的动态权重分配和跨模态对比学习,提升多模态无人机波束预测的准确性和鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.00742 2026-01-01 cs.CV cs.AI eess.IV 81%

Zoomer: Adaptive Image Focus Optimization for Black-box MLLM

Zoomer: 为黑盒大语言模型实现自适应图像聚焦优化

Jiaxu Qian, Chendong Wang, Yifan Yang, Chaoyun Zhang, Huiqiang Jiang, Xufang Luo, Yu Kang, Qingwei Lin, Anlan Zhang, Shiqi Jiang, Ting Cao, Tianjun Mao, Suman Banerjee, Guyue Liu, Saravan Rajmohan, Dongmei Zhang, Yuqing Yang, Qi Zhang, Lili Qiu

机构 * Microsoft(微软公司) Peking University(北京大学) University of Wisconsin Madison(威斯康星大学麦迪逊分校) University of Southern California(南加州大学)

专题命中 多模态训练与对齐 :MLLM(title);multimodal(abstract);分类 cs.CV、cs.AI

AI总结 Zoomer通过自适应图像聚焦优化提升黑盒MLLM的多模态理解能力,显著提升准确性并减少token使用

Comments TMLR accepted

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22748 2026-01-01 cs.CV 79%

TrimTokenator-LC: Towards Adaptive Visual Token Pruning for Large Multimodal Models with Long Contexts

TrimTokenator-LC: 向大型多模态模型长上下文的自适应视觉标记修剪迈进

Hao Zhang, Mengsi Lyu, Bo Huang, Yulong Ao, Yonghua Lin

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 TrimTokenator-LC通过自适应视觉标记修剪方法,在长上下文和多图像场景中有效减少视觉标记数量,同时保持性能。

Comments 17 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.24793 2026-01-01 cs.LG cs.NE 78%

Self-Supervised Neural Architecture Search for Multimodal Deep Neural Networks

多模态深度神经网络的自监督神经架构搜索

Shota Suzuki, Satoshi Ono

专题命中 多模态训练与对齐 :multimodal(title,abstract)

AI总结 本文提出了一种自监督学习方法,用于多模态深度神经网络的架构搜索,能够在未标记数据上有效设计DNN架构。

Journal ref IEICE Transactions on Information and Systems, Vol.E108.D, No. 6, pp. 640-643, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.24826 2026-01-01 cs.CV cs.AI 62%

Video and Language Alignment in 2D Systems for 3D Multi-object Scenes with Multi-Information Derivative-Free Control

用于多物体3D场景的2D系统中视频与语言对齐的多信息无导数控制

Jason Armitage, Rico Sennnrich

机构 * University of Zurich(苏黎世大学)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出了一种无导数优化方法,用于在多物体3D场景中实现视频与语言的对齐,通过在线适应物体遮挡和区分特征来提升跨模态任务性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.17221 2026-01-01 cs.CV 57%

DAVE: A VLM Vision Encoder for Document Understanding and Web Agents

DAVE: 一种用于文档理解与网络代理的视觉编码器

Brandon Huang, Hang Hua, Zhuoran Yu, Trevor Darrell, Rogerio Feris, Roei Herzig

机构 * MIT-IBM Watson AI Lab(MIT-IBM Watson AI实验室) UC Berkeley(加州大学伯克利分校) University of Wisconsin–Madison(威斯康星大学麦迪逊分校)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

AI总结 DAVE是一种专为文档理解和网络代理设计的视觉编码器,通过自监督和监督预训练结合模型融合策略,提升对文档和网络任务的适应性与性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.24284 2026-01-01 cs.RO cs.AI 57%

DRL-TH: Jointly Utilizing Temporal Graph Attention and Hierarchical Fusion for UGV Navigation in Crowded Environments

DRL-TH:联合利用时序图注意力和层次融合用于拥挤环境下的无人地面车辆导航

Ruitong Li, Lin Zhang, Yuenan Zhao, Chengxin Liu, Ran Song, Wei Zhang

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.AI

AI总结 DRL-TH通过结合时序图注意力和层次融合,提升无人地面车辆在拥挤环境中的导航性能和动态适应性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.23335 2026-01-01 cs.CV cs.LG 57%

Visual Language Hypothesis

视觉语言假说

Xiu Li

机构 * Independent Researcher(独立研究者)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

AI总结 该研究提出视觉理解依赖于语义语言的假说,并通过拓扑结构解释语义抽象和不变性的实现机制。

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 其他多模态 6 篇

2512.24786 2026-01-01 physics.med-ph 78%

A Dual-Tuned Concentric Multimodal RF Coil for 7T 1H/31P MRSI: Concurrently Enhancing B1 Efficiency Over Single-Tuned References

双调谐同心多模RF线圈用于7T 1H/31P MRSI:同时提升31P的B1效率与单调谐参考的1H性能

Yunkun Zhao, Xiaoliang Zhang

专题命中 其他多模态 :multimodal(title,abstract)

AI总结 本研究提出了一种双调谐同心多模RF线圈,用于7T 1H/31P MRSI,显著提升31P B1效率并改善1H性能。

Comments 21 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.07666 2026-01-01 cs.LG cs.AI cs.CL cs.CV 67%

Model Merging in LLMs, MLLMs, and Beyond: Methods, Theories, Applications and Opportunities

在大语言模型、多模态大语言模型及更广泛的领域中进行模型融合:方法、理论、应用与机遇

Enneng Yang, Li Shen, Guibing Guo, Xingwei Wang, Xiaochun Cao, Jie Zhang, Dacheng Tao

机构 * Shenzhen Campus of Sun Yat-sen University, China(中山大学深圳校区) Northeastern University China(东北大学) Shenzhen Campus of Sun Yat-sen University China(中山大学深圳校区) Nanyang Technological University Singapore(南洋理工大学) Northeastern University(东北大学) Shenzhen Campus of Sun Yat-sen University(中山大学深圳校区) Nanyang Technological University(南洋理工大学) Institute for Clarity in Documentation Dublin Ohio USA(文档清晰研究所) Inria Paris-Rocquencourt Rocquencourt France(巴黎-罗quentourt研究所) Rajiv Gandhi University Doimukh Arunachal Pradesh India(拉贾·甘地大学) Tsinghua University Haidian Qu Beijing Shi China(清华大学) Palmer Research Laboratories San Antonio Texas USA(帕勒研究中心) Institute for Clarity in Documentation(文档清晰研究所) Inria Paris-Rocquencourt(巴黎-罗quentourt研究所) Rajiv Gandhi University(拉贾·甘地大学) Tsinghua University(清华大学) Palmer Research Laboratories(帕勒研究中心)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 本文综述了模型融合的方法、理论、应用及未来方向,提出新的分类方法并探讨其在多个机器学习领域的应用及挑战。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.24007 2026-01-01 cs.NE cs.AI 57%

TESO Tabu Enhanced Simulation Optimization for Noisy Black Box Problems

噪声黑盒问题的TESO禁忌增强仿真优化

Bulent Soykan, Sean Mondesire, Ghaith Rabadi

机构 * Institute for Simulation and Training(模拟与培训研究所) University of Central Florida(中央佛罗里达大学) School of Modeling, Simulation, and Training(建模、模拟与培训学院)

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

AI总结 TESO通过结合禁忌列表和精英记忆,提升噪声黑盒问题的仿真优化效率与稳定性。

Comments 11 pages, 2 figures, Presented at the Winter Simulation Conference 2025, Seattle, Washington (December 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.24220 2026-01-01 cond-mat.mes-hall 50%

Non-Euclidean interfaces decode the continuous landscape of graphene-induced surface reconstructions

非欧几里得界面解码石墨烯诱导表面重构的连续景观

Li-Qun Shen, Hao-Jin Wang, Mengzhao Sun, Yang Xiang, Xin-Ning Tian, Yue Chai, Yue Yang, Feng Ding, Xiao Kong, Marc-Georg Willinger, Zhu-Jun Wang

专题命中 其他多模态 :multimodal(abstract)

AI总结 非欧几里得界面通过石墨烯-铜表面模型系统,解码了石墨烯诱导表面重构的连续景观,并揭示了其热力学机制。

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.16152 2026-01-01 math.LO 50%

A vector logic for extensional formal semantics

一个扩展式形式语义的向量逻辑

Daniel Quigley

专题命中 其他多模态 :multimodal(abstract)

AI总结 本文提出一种向量逻辑,证明扩展式形式语义与分布向量空间语义的同构关系,展示二者在结构上的兼容性。

Comments 37 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.14737 2026-01-01 cond-mat.mtrl-sci nlin.AO 50%

Unveiling dynamic bifurcation of Resch-patterned origami for self-adaptive impact mitigation structure

揭示Resch图案折纸的动态分叉行为用于自适应冲击缓解结构

Yasuhiro Miyazawa, Dahun Lee, Seonghyun Kim, Chia-Yung Chang, Qixun Li, Ryan Tenu Ahn, Minho Cha, Koshiro Yamaguchi, Yuyang Song, Shinnosuke Shimokawa, Umesh Gandhi, Jinkyu Yang

专题命中 其他多模态 :multi-modal(abstract)

AI总结 基于Resch图案折纸设计自适应冲击缓解结构,通过动态分叉行为实现多模式变形,提升能量耗散能力。

详情

展开后加载摘要…

URL PDF HTML 收藏