arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-11-13 至 2025-11-13 共收录 6 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6 篇

2508.02557 2025-11-13 eess.IV cs.CV 83%

RL-U$^2$Net: A Dual-Branch UNet with Reinforcement Learning-Assisted Multimodal Feature Fusion for Accurate 3D Whole-Heart Segmentation

Jierui Qu, Jianchun Zhao

专题命中 多模态训练与对齐 :multimodal(title);multi-modal(abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.07722 2025-11-13 cs.AI 83%

Is Cognition Consistent with Perception? Assessing and Mitigating Multimodal Knowledge Conflicts in Document Understanding

Zirui Shao, Feiyu Gao, Zhaoqing Zhu, Chuwei Luo, Hangdi Xing, Zhi Yu, Qi Zheng, Ming Yan, Jiajun Bu

机构 * Zhejiang Key Laboratory of Accessible Perception and Intelligent Systems, Zhejiang University(浙江可感知智能系统重点实验室,浙江大学) Alibaba Group(阿里巴巴集团) Hangzhou High-Tech Zone (Binjiang) Institute of Blockchain and DataSecurity(杭州高新技术区(滨江)区块链与数据安全研究院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.AI

Comments Accepted at EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09286 2025-11-13 cs.CV 79%

Enriching Knowledge Distillation with Cross-Modal Teacher Fusion

Amir M. Mansourian, Amir Mohammad Babaei, Shohreh Kasaei

机构 * Image Processing Lab, Sharif University of Technology(沙斐大学技术实验室)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

Comments 11 pages, 5 figures, 8 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00452 2025-11-13 cs.IR cs.AI 79%

M^2VAE: Multi-Modal Multi-View Variational Autoencoder for Cold-start Item Recommendation

Chuan He, Yongchao Liu, Qiang Li, Wenliang Zhong, Chuntao Hong, Xinwei Yao

机构 * Ant Group(蚂蚁集团)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08970 2025-11-13 astro-ph.SR 78%

JW-Flare: Accurate Solar Flare Forecasting Method Based on Multimodal Large Language Models

Mingfu Shao, Hui Wang, Yuyang Li, Jiaben Lin, Jifeng Liu, Baolin Tan, Juan Guo, Yin Zhang, Jing Huang, Jiangtao Su, Yingzi Sun, Haiqing Xu, Jie Chen, Suo Liu, Yuanyong Deng, Liyue Tong, Yang Bai, Cunshi Wang, Kaifan Ji, Yuqing Zhou

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Comments 12 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15062 2025-11-13 cs.RO cs.AI cs.LG 70%

Touch in the Wild: Learning Fine-Grained Manipulation with a Portable Visuo-Tactile Gripper

Xinyue Zhu, Binghao Huang, Yunzhu Li

机构 * Columbia University(哥伦比亚大学)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.AI

Comments More videos can be found on our website:https://binghao-huang.github.io/touch_in_the_wild/

详情

展开后加载摘要…

URL PDF HTML 收藏