arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-12-24 至 2025-12-24 共收录 41 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态Agent 7 篇

2503.11006 2025-12-24 cs.CV cs.AI 62%

Fine-Grained Instruction-Guided Graph Reasoning for Vision-and-Language Navigation

细粒度指令引导的图推理用于视觉-语言导航

Yaohua Liu, Xinyuan Song, Yunfu Deng, Yifan Xie, Binkai Ou, Yan Zhong

机构 * Guangdong Institute of Intelligence Science and Technology(广东智能科学与技术研究院) Department of Computer Science, Emory University(埃默里大学计算机科学系) Department of Computer Science, University of Wisconsin-Madison(威斯康星大学麦迪逊分校计算机科学系) Tsinghua University(清华大学) Innovation and Research and Development Department, BoardWare Information System Company(BoardWare信息系统公司创新与研发部) School of Mathematics, Peking University(北京大学数学学院)

专题命中 多模态Agent :cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出细粒度指令引导的图推理框架OIKG,通过解耦角度与视觉提示并增强空间表示,提升视觉-语言导航中指令理解和跨模态对齐能力。

Comments 10 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.20206 2025-12-24 cs.AI 57%

TongSIM: A General Platform for Simulating Intelligent Machines

TongSIM: 一个用于模拟智能机器的通用平台

Zhe Sun, Kunlun Wu, Chuanjian Fu, Zeming Song, Langyong Shi, Zihe Xue, Bohan Jing, Ying Yang, Xiaomeng Gao, Aijia Li, Tianyu Guo, Huiying Li, Xueyuan Yang, Rongkai Liu, Xinyi He, Yuxi Wang, Yue Li, Mingyuan Liu, Yujie Lu, Hongzhao Xie, Shiyun Zhao, Bo Dai, Wei Wang, Tao Yuan, Song-Chun Zhu, Yujia Peng, Zhenliang Zhang

机构 * State Key Laboratory of General Artificial Intelligence(通用人工智能国家重点实验室) BIGAI Beijing, China(中国北京)

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

AI总结 TongSIM是一个通用平台,用于训练和评估具身智能代理,提供多样化的场景和交互丰富的环境,支持从低级导航到高级复合活动的广泛研究需求。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.08674 2025-12-24 cs.AI cs.MA 57%

Multi-Agent Intelligence for Multidisciplinary Decision-Making in Gastrointestinal Oncology

多智能体智能在胃肠肿瘤多学科决策中的应用

Rongzhao Zhang, Junqiao Wang, Shuyun Yang, Mouxiao Bian, Chihao Zhang, Dongyang Wang, Qiujuan Yan, Yun Zhong, Yuwei Bai, Guanxu Zhu, Kangkun Mao, Miao Wang, Chao Ding, Renjie Lu, Lei Wang, Lei Zheng, Tao Zheng, Xi Wang, Zhuo Fan, Bing Han, Meiling Liu, Luyi Jiang, Dongming Shan, Wenzhong Jin, Jiwei Yu, Zheng Wang, Jie Xu, Meng Luo

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Shanghai Ninth People’s Hospital, Shanghai Jiao Tong University School of Medicine(上海第九人民医院,上海交通大学医学院) Renji Hospital, Shanghai Jiao Tong University School of Medicine(仁济医院,上海交通大学医学院) Shanghai Institute of Infectious Disease and Biosecurity, Fudan University(上海市传染病防治研究所暨生物安全研究所,复旦大学) Shanghai Health Development Research Center (Shanghai Medical Information Center)(上海市卫生健康发展研究中心(上海市医疗信息中心)) Shanghai Kupas Technology Co., Ltd.(上海库帕斯科技有限公司)

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

AI总结 本文提出基于多智能体的框架,用于提升胃肠肿瘤多学科决策的自动化支持,通过模拟多学科团队协作,提高推理逻辑和医疗准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 多模态训练与对齐 6 篇

2512.19934 2025-12-24 cs.CV cs.AI cs.LG 86%

Vehicle-centric Perception via Multimodal Structured Pre-training

基于多模态结构预训练的车辆感知

Wentao Wu, Xiao Wang, Chenglong Li, Jin Tang, Bin Luo

机构 * Information Materials and Intelligent Sensing Laboratory of Anhui Province(安徽省信息材料与智能感知实验室) Anhui Provincial Key Laboratory of Multimodal Cognitive Computation(安徽省多模态认知计算重点实验室) the School of Artificial Intelligence, Anhui University(安徽大学人工智能学院) School of Computer Science and Technology, Anhui University(安徽大学计算机科学与技术学院) Institute of Artificial Intelligence, Hefei Comprehensive National Science Center(合肥综合性国家科学中心人工智能研究院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);image-text(abstract);分类 cs.CV、cs.AI

AI总结 本文提出VehicleMAE-V2,通过多模态结构先验知识提升车辆感知的预训练能力,采用SMM、CRM和SRM模块增强模型对车辆结构和语义的理解。

Comments Journal extension of VehicleMAE (AAAI 2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.20084 2025-12-24 cs.LG cs.AI 79%

QE-Catalytic: A Graph-Language Multimodal Base Model for Relaxed-Energy Prediction in Catalytic Adsorption

QE-Catalytic: 一种图-语言多模态基础模型,用于催化吸附中放松能量的预测

Yanjie Li, Jian Xu, Xueqing Chen, Lina Yu, Shiming Xiang, Weijun Li, Cheng-lin Liu

机构 * AnnLab(安实验室) Institute of Semiconductors, Chinese Academy of Sciences(半导体研究所,中国科学院) Zhongguancun Academy(中关村学院) State Key Laboratory of Multimodal Artificial Intelligence Systems(多模态人工智能系统国家重点实验室) Institute of Automation, Chinese Academy of Sciences(自动化研究所,中国科学院) University of Chinese Academy of Sciences(中国科学院大学) Computer Network Information Center(计算机网络信息中心)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

AI总结 QE-Catalytic结合语言模型与图Transformer,实现高精度催化吸附能量预测及逆向设计

Comments 25 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.20026 2025-12-24 cs.CV 79%

MAPI-GNN: Multi-Activation Plane Interaction Graph Neural Network for Multimodal Medical Diagnosis

MAPI-GNN:多激活平面交互图神经网络用于多模态医学诊断

Ziwei Qin, Xuhui Song, Deqing Huang, Na Qin, Jun Li

机构 * Ziwei Qin(独立研究者) Xuhui Song(独立研究者) Deqing Huang(独立研究者) Na Qin(独立研究者) Jun Li(独立研究者)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 MAPI-GNN通过多激活平面交互机制,有效建模患者特异性病理关系,提升多模态医学诊断的准确性。

Comments Accepted by Proceedings of the AAAI Conference on Artificial Intelligence 40 (AAAI-26)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.20561 2025-12-24 cs.CV 74%

FlashVLM: Text-Guided Visual Token Selection for Large Multimodal Models

FlashVLM: 大规模多模态模型中的文本引导视觉令牌选择

Kaitong Cai, Jusheng Zhang, Jing Yang, Yijia Fan, Pengtao Xie, Jian Wang, Keze Wang

机构 * Sun Yat-sen University(中山大学) University of California, San Diego(加州大学圣地亚哥分校) Snap Inc.(Snap公司)

专题命中 多模态训练与对齐 :multimodal(title);分类 cs.CV

AI总结 FlashVLM通过文本引导的视觉令牌选择,实现了高效压缩和高准确率,在大规模多模态模型中表现出色。

Comments Under submission

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.20556 2025-12-24 cs.CV 57%

Multi-Grained Text-Guided Image Fusion for Multi-Exposure and Multi-Focus Scenarios

多粒度文本引导图像融合用于多曝光和多聚焦场景

Mingwei Tang, Jiahao Nie, Guang Yang, Ziqing Cui, Jie Li

机构 * Xidian University(西安电子科技大学) Nanyang Technological University(新加坡国立大学) Xi’an University of Technology(西安理工大学)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

AI总结 本文提出MTIF方法,通过多粒度文本描述和跨模态调节模块提升多曝光和多聚焦图像融合性能。

Comments Accepted to WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.20034 2025-12-24 cs.IR 50%

VSA:Visual-Structural Alignment for UI-to-Code

VSA:面向UI到代码的视觉-结构对齐

Xian Wu, Ming Zhang, Zhiyu Fang, Fei Li, Bin Wang, Yong Jiang, Hao Zhou

专题命中 多模态训练与对齐 :multimodal(abstract)

AI总结 VSA通过视觉-结构对齐提升UI到代码的模块化和一致性,生成类型安全的组件以提高软件工程效率。

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 其他多模态 2 篇

2512.18279 2025-12-24 cs.CV 79%

UniMPR: A Unified Framework for Multimodal Place Recognition with Heterogeneous Sensor Configurations

UniMPR: 一种用于异构传感器配置的多模态地点识别统一框架

Zhangshuo Qi, Jingyi Xu, Luqi Cheng, Shichen Wen, Yiming Ma, Guangming Xiong

机构 * Beijing Institute of Technology(北京理工大学) Shanghai Jiao Tong University(上海交通大学) The University of New South Wales(新南威尔士大学)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 UniMPR提出了一种统一框架,能够适应多种异构传感器配置,通过极坐标BEV特征空间和多分支网络实现多模态地点识别的高效与鲁棒性。

Comments 14 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.11181 2025-12-24 cs.LG cs.AI 57%

Mixture of Experts in Large Language Models

大语言模型中的专家混合架构

Danyang Zhang, Junhao Song, Ziqian Bi, Xinyuan Song, Yingfang Yuan, Tianyang Wang, Joe Yeong, Junfeng Hao

机构 * Department of Research(研究部) ByteDance Inc(字节跳动公司) Department of CS(计算机科学系) Imperial College London(伦敦帝国理工学院) Purdue University(普渡大学) Emory University(埃默里大学) Department of Computer Science(计算机科学系) Heriot-Watt University(赫罗特-瓦特大学) AI Agent Lab(AI代理实验室) Vokram Group(Vokram集团) Department of Anatomical Pathology(解剖病理学系) Singapore General Hospital(新加坡中央医院)

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

AI总结 本文研究了大语言模型中专家混合架构的性能提升与应用挑战,分析了其核心机制与优化方法,探讨了MoE在任务特定性能和模型扩展方面的优势。

详情

展开后加载摘要…

URL PDF HTML 收藏