arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-09-19 至 2025-09-19 共收录 6 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6 篇

2509.14735 2025-09-19 cs.CL 88%

Decoupled Proxy Alignment: Mitigating Language Prior Conflict for Multimodal Alignment in MLLM

Chenkun Tan, Pengyu Wang, Shaojun Zhou, Botian Jiang, Zhaowei Li, Dong Zhang, Xinghao Wang, Yaqian Zhou, Xipeng Qiu

机构 * Fudan University(复旦大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(title,abstract);分类 cs.CL

Comments Accepted by Findings of EMNLP2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14383 2025-09-19 cs.RO cs.CV 83%

RLBind: Adversarial-Invariant Cross-Modal Alignment for Unified Robust Embeddings

Yuhong Lu

机构 * Samueli School of Engineering, Electrical and Computer Engineering, UCLA(UCLA电气与计算机工程学院萨姆利学校)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV

Comments This paper is submitted to IEEE International Conference on Robotics and Automation (ICRA) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.18042 2025-09-19 cs.CV cs.AI 81%

VLM-E2E: Enhancing End-to-End Autonomous Driving with Multimodal Driver Attention Fusion

Pei Liu, Haipeng Liu, Haichao Liu, Xin Liu, Jinxin Ni, Jun Ma

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Li Auto Inc.(Li汽车公司) the School of Aeronautics and Astronautics, Xiamen University(厦门大学航空航天学院) The Hong Kong University of Science and Technology(香港科技大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15096 2025-09-19 cs.CV 79%

OmniSegmentor: A Flexible Multi-Modal Learning Framework for Semantic Segmentation

Bo-Wen Yin, Jiao-Long Cao, Xuying Zhang, Yuming Chen, Ming-Ming Cheng, Qibin Hou

机构 * VCIP, CS, Nankai University(视觉计算研究所,计算机科学,南开大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14739 2025-09-19 cs.CV 57%

FMGS-Avatar: Mesh-Guided 2D Gaussian Splatting with Foundation Model Priors for 3D Monocular Avatar Reconstruction

Jinlong Fan, Bingyu Hu, Xingguang Li, Yuxiang Yang, Jing Zhang

机构 * HangZhou Dianzi University(杭州电子大学) Shenzhen Polytechnic University(深圳职业技术大学) WuHan University(武汉大学)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14084 2025-09-19 cs.CV 57%

AD-DINOv3: Enhancing DINOv3 for Zero-Shot Anomaly Detection with Anomaly-Aware Calibration

Jingyi Yuan, Jianxiong Ye, Wenkang Chen, Chenqiang Gao

机构 * School of Intelligent Systems Engineering, Sun Yat-Sen University(智能系统工程学院,中山大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏