arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-09-24 至 2025-09-24 共收录 14 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 14 篇

2509.19212 2025-09-24 cs.CL cs.AI 84%

Steering Multimodal Large Language Models Decoding for Context-Aware Safety

Zheyuan Liu, Zhangchen Xu, Guangyao Dou, Xiangchi Yuan, Zhaoxuan Tan, Radha Poovendran, Meng Jiang

机构 * University of Notre Dame(notre dame 大学) University of Washington(华盛顿大学) Johns Hopkins University(约翰霍普金斯大学) Georgia Institute of Technology(佐治亚理工学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CL、cs.AI

Comments A lightweight and model-agnostic decoding framework that dynamically adjusts token generation based on multimodal context

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.12623 2025-09-24 cs.SD cs.AI cs.CL cs.MM eess.AS 83%

DeepResonance: Enhancing Multimodal Music Understanding via Music-centric Multi-way Instruction Tuning

Zhuoyuan Mao, Mengjie Zhao, Qiyu Wu, Hiromi Wakaki, Yuki Mitsufuji

机构 * Sony Group Corporation(索尼集团公司) Sony AI(索尼人工智能)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL、cs.AI、cs.MM

Comments Accepted to EMNLP 2025 main conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18221 2025-09-24 cs.AI cs.LG 83%

Multimodal Health Risk Prediction System for Chronic Diseases via Vision-Language Fusion and Large Language Models

Dingxin Lu, Shurui Wu, Xinyi Huang

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19018 2025-09-24 cs.LG 82%

OmniBridge: Unified Multimodal Understanding, Generation, and Retrieval via Latent Space Alignment

Teng Xiao, Zuchao Li, Lefei Zhang

机构 * School of Computer Science, Wuhan University(武汉大学计算机学院) School of Artificial Intelligence, Wuhan University(武汉大学人工智能学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17492 2025-09-24 cs.CV cs.AI 81%

Multimodal Medical Image Classification via Synergistic Learning Pre-training

Qinghua Lin, Guang-Hai Liu, Zuoyong Li, Yang Li, Yuting Jiang, Xiang Wu

机构 * College of Biomedical Engineering, Fudan University(复旦大学生物医学工程学院) College of Computer Science and Engineering, Guangxi Normal University(广西师范大学计算机科学与工程学院) Fujian Provincial Key Laboratory of Information Processing and Intelligent Control, School of Computer and Big Data, Minjiang University(福建省信息处理与智能控制重点实验室,闽江学院计算机与大数据学院) Department of Automation Science and Electrical Engineering, Beihang University(北京航空航天大学自动化科学与电气工程学院) Department of Digestive Endoscopy, Fuzhou University Affiliated Provincial Hospital, Provincial Clinical Medical College of Fujian Medical University(福州市大学附属省医院消化内镜科,福建医科大学省临床医学学院) Department of Urology, Fuzhou University Affiliated Provincial Hospital, Provincial Clinical Medical College of Fujian Medical University(福州市大学附属省医院泌尿科,福建医科大学省临床医学学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19208 2025-09-24 cs.CV 79%

Enabling Plant Phenotyping in Weedy Environments using Multi-Modal Imagery via Synthetic and Generated Training Data

Earl Ranario, Ismael Mayanja, Heesup Yun, Brian N. Bailey, J. Mason Earles

专题命中 多模态训练与对齐 :multi-modal(title);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19047 2025-09-24 cs.RO 67%

ManipForce: Force-Guided Policy Learning with Frequency-Aware Representation for Contact-Rich Manipulation

Geonhyup Lee, Yeongjin Lee, Kangmin Kim, Seongju Lee, Sangjun Noh, Seunghyeok Back, Kyoobin Lee

机构 * Department of AI Convergence, Gwangju Institute of Science and Technology (GIST)(人工智能融合系,全州科学技术院) Department of AI Machinery, Korea Institute of Machinery & Materials (KIMM)(人工智能机械系,韩国机械材料研究院)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract)

Comments 9 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.04183 2025-09-24 cs.CL cs.AI 62%

GALLa: Graph Aligned Large Language Models for Improved Source Code Understanding

Ziyin Zhang, Hang Yu, Shijie Li, Peng Di, Jianguo Li, Rui Wang

机构 * Ant Group(蚂蚁集团) Shanghai Jiao Tong University(上海交通大学)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CL、cs.AI

Comments ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18743 2025-09-24 cs.CV 57%

TriFusion-AE: Language-Guided Depth and LiDAR Fusion for Robust Point Cloud Processing

Susmit Neogi

机构 * Department of Mechanical Engineering(机械工程系) Indian Institute of Technology Bombay(印度理工学院班加罗尔) Mumbai, India(孟买,印度)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments 39th Conference on Neural Information Processing Systems (NeurIPS 2025) Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18738 2025-09-24 cs.CV 57%

HyPSAM: Hybrid Prompt-driven Segment Anything Model for RGB-Thermal Salient Object Detection

Ruichao Hou, Xingyuan Li, Tongwei Ren, Dongming Zhou, Gangshan Wu, Jinde Cao

机构 * State Key Laboratory for Novel Software Technology, Nanjing University(南京大学新型软件技术国家重点实验室) School of Information Science and Engineering, Yunnan University(云南大学信息科学与工程学院) School of Mathematics, Southeast University(东南大学数学学院)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18733 2025-09-24 cs.CV 57%

Knowledge Transfer from Interaction Learning

Yilin Gao, Kangyi Chen, Zhongxing Peng, Hengjie Lu, Shugong Xu

机构 * Shanghai University(上海大学) Xi’an Jiaotong-Liverpool University(西安交通大学利物浦大学)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

Comments Accepted by ICCV2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18613 2025-09-24 cs.CV 57%

MLF-4DRCNet: Multi-Level Fusion with 4D Radar and Camera for 3D Object Detection in Autonomous Driving

Yuzhi Wu, Li Xiao, Jun Liu, Guangfeng Jiang, XiangGen Xia

机构 * MoE Key Laboratory of Brain-Inspired Intelligence Perception and Cognition, University of Science and Technology of China(脑启发智能感知与认知教育部重点实验室,中国科学技术大学) Institute of Artificial Intelligence, Hefei Comprehensive National Science Center(人工智能研究院,合肥国家科学中心) Department of Electronic Engineering and Information Science, University of Science and Technology of China(电子工程与信息科学系,中国科学技术大学) Department of Electrical and Computer Engineering, University of Delaware(电气与计算机工程系,德克萨斯大学)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18152 2025-09-24 cs.LG cs.AI 57%

WLFM: A Well-Logs Foundation Model for Multi-Task and Cross-Well Geological Interpretation

Zhenyu Qi, Qing Yu, Jichen Wang, Yun-Bo Zhao, Zerui Li, Wenjun Lv

机构 * Institute of Advanced Technology, University of Science and Technology of China(科学技术大学先进技术研究所) Department of Automation, University of Science and Technology of China(科学技术大学自动化系) Institute of Artificial Intelligence, Hefei Comprehensive National Science Center(合肥综合国家科学中心人工智能研究所)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.19972 2025-09-24 cs.CV 57%

DAOcc: 3D Object Detection Assisted Multi-Sensor Fusion for 3D Occupancy Prediction

Zhen Yang, Yanpeng Dong, Jiayu Wang, Heng Wang, Lichao Ma, Zijian Cui, Qi Liu, Haoran Pei, Kexin Zhang, Chao Zhang

机构 * Beijing Mechanical Equipment Institute(北京机械设备研究所)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

Comments TCSVT Accepted version (not the final published version)

Journal ref IEEE Transactions on Circuits and Systems for Video Technology, 2025, Print ISSN: 1051-8215, Online ISSN: 1558-2205

详情

展开后加载摘要…

URL PDF HTML 收藏