arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-11-14 至 2025-11-14 共收录 8 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 8 篇

2511.10035 2025-11-14 cs.CV 79%

DGFusion: Dual-guided Fusion for Robust Multi-Modal 3D Object Detection

Feiyang Jia, Caiyan Jia, Ailin Liu, Shaoqing Xu, Qiming Xia, Lin Liu, Lei Yang, Yan Gong, Ziying Song

机构 * School of Computer Science and Technology, Beijing Key Laboratory of Traffic Data Mining and Embodied Intelligence, Beijing Jiaotong University(计算机科学与技术学院、交通数据挖掘与具身智能北京市重点实验室、北京交通大学) State Key Laboratory of Internet of Things for Smart City and Department of Electrome chanical Engineering, University of Macau(智能城市物联网国家重点实验室、澳门大学机电工程系) Fujian Key Laboratory of Sensing and Computing for Smart Cities, Xiamen University(智能城市感知与计算福建省重点实验室、厦门大学) School of Mechanical and Aerospace Engineering, Nanyang Technological University(机械与航空航天工程学院、南洋理工大学) State Key Laboratory of Robotics and System, Harbin Institute of Technology(机器人系统国家重点实验室、哈尔滨工业大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12174 2025-11-14 cs.CV cs.RO 79%

UniGS: Unified Geometry-Aware Gaussian Splatting for Multimodal Rendering

Yusen Xie, Zhenmin Huang, Jianhao Jiao, Dimitrios Kanoulas, Jun Ma

机构 * HKUST (GZ)(香港科技大学(广州)) HKUST(香港科技大学) UCL(伦敦大学学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10081 2025-11-14 cs.CV 70%

GridPrune: From "Where to Look" to "What to Select" in Visual Token Pruning for MLLMs

Yuxiang Duan, Ao Li, Yingqin Li, Luyu Li, Pengwei Wang

机构 * Shandong University(山东大学)

专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10211 2025-11-14 cs.CV 57%

HeatV2X: Scalable Heterogeneous Collaborative Perception via Efficient Alignment and Interaction

Yueran Zhao, Zhang Zhang, Chao Sun, Tianze Wang, Chao Yue, Nuoran Li

机构 * Beijing Institute of Technology(北京理工大学)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

Comments 10 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10087 2025-11-14 cs.RO cs.AI cs.LG 57%

Opinion: Towards Unified Expressive Policy Optimization for Robust Robot Learning

Haidong Huang, Haiyue Zhu. Jiayu Song, Xixin Zhao, Yaohua Zhou, Jiayi Zhang, Yuze Zhai, Xiaocong Li

机构 * Eastern Institute of Technology(东部技术研究所) University of Nottingham(诺丁汉大学) SIMTech, Agency for Science, Technology and Research (A*STAR)(SIMTech,科技研究局(A*STAR)) Southern University of Science and Technology(南方科技大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

Comments Accepted by NeurIPS 2025 Workshop on Embodied World Models for Decision Making

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08903 2025-11-14 cs.CV 57%

LLM-Guided Probabilistic Fusion for Label-Efficient Document Layout Analysis

Ibne Farabi Shihab, Sanjeda Akter, Anuj Sharma

机构 * Department of Computer Science, Iowa State University(计算机科学系,爱荷华州立大学) Department of Civil, Construction & Environmental Engineering, Iowa State University(土木、建设与环境工程系,爱荷华州立大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09725 2025-11-14 physics.plasm-ph cs.LG physics.data-an 50%

The Data Fusion Labeler (dFL): Challenges and Solutions to Data Harmonization, Labeling, and Provenance in Fusion Energy

Craig Michoski, Matthew Waller, Brian Sammuli, Zeyu Li, Tapan Ganatma Nakkina, Raffi Nazikian, Sterling Smith, David Orozco, Dongyang Kuang, Martin Foltin, Erik Olofsson, Mike Fredrickson, Jerry Louis-Jeune, David R. Hatch, Todd A. Oliver, Mitchell Clark, Steph-Yves Louis

机构 * University of Texas, Austin, TX, USA(德克萨斯大学)

专题命中 多模态训练与对齐 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.02909 2025-11-14 cs.LG 50%

Fine-grained Token Allocation Via Operation Pruning for Efficient MLLMs

Aoming Liu, Reuben Tan, Boqing Gong, Bryan A. Plummer

机构 * Boston University(波士顿大学) Microsoft Research(微软研究院)

专题命中 多模态训练与对齐 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏