arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-10-03 至 2025-10-03 共收录 49 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 15 篇

2510.01954 2025-10-03 cs.CV 83%

Patch-as-Decodable-Token: Towards Unified Multi-Modal Vision Tasks in MLLMs

Yongyi Su, Haojie Zhang, Shijie Li, Nanqing Liu, Jingyi Liao, Junyi Pan, Yuan Liu, Xiaofen Xing, Chong Sun, Chen Li, Nancy F. Chen, Shuicheng Yan, Xulei Yang, Xun Xu

机构 * South China University of Technology(华南理工大学) Institute for Infocomm Research (I 2 R), A*STAR(信息通信研究所(I 2 R),A*STAR) WeChat Vision, Tencent Inc.(微信视觉,腾讯公司) Foshan University(佛山大学) Nanyang Technological University(南洋理工大学) National University of Singapore(新加坡国立大学)

专题命中 多模态训练与对齐 :multi-modal(title);multimodal(abstract);MLLM(abstract);分类 cs.CV

Comments 24 pages, 12 figures and 9 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01677 2025-10-03 cs.LG cs.CV 83%

Beyond Simple Fusion: Adaptive Gated Fusion for Robust Multimodal Sentiment Analysis

Han Wu, Yanming Sun, Yunhe Yang, Derek F. Wong

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01428 2025-10-03 q-bio.QM cs.AI 83%

BioVERSE: Representation Alignment of Biomedical Modalities to LLMs for Multi-Modal Reasoning

Ching-Huei Tsou, Michal Ozery-Flato, Ella Barkan, Diwakar Mahajan, Ben Shapira

机构 * IBM T.J. Watson Research Center(IBM T.J. Watson研究所以) IBM Research(IBM研究所以)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00320 2025-10-03 cs.CV 83%

TrimTokenator: Towards Adaptive Visual Token Pruning for Large Multimodal Models

Hao Zhang, Mengsi Lyu, Chenrui He, Yulong Ao, Yonghua Lin

机构 * Beijing Academy of Artificial Intelligence (BAAI)(北京人工智能研究院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments 15 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01780 2025-10-03 cs.CR cs.AI cs.CY cs.LG 79%

Secure Multi-Modal Data Fusion in Federated Digital Health Systems via MCP

Aueaphum Aueawatthanaphisut

机构 * School of Information, Computer, and Communication Technology(信息、计算机与通信技术学院) Sirindhorn International Institute of Technology, Thammasat University(泰国朱拉安吞国际技术学院,泰国 Thammasat 大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.AI

Comments 6 pages, 8 figures, 7 equations, 1 algorithm

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01606 2025-10-03 cs.IR cs.AI cs.CL 76%

Bridging Collaborative Filtering and Large Language Models with Dynamic Alignment, Multimodal Fusion and Evidence-grounded Explanations

Bo Ma, LuYao Liu, Simon Lau, Chandler Yuan, and XueY Cui, Rosie Zhang

机构 * Department of Software \& Microelectronics, Peking University, Beijing, China Economic Law School, China University of Political Science Financial Media, Peking University, ChangSha, China

专题命中 多模态训练与对齐 :multimodal(title);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14497 2025-10-03 cs.CV cs.CL 73%

Efficient Whole Slide Pathology VQA via Token Compression

Weimin Lyu, Qingqiao Hu, Kehan Qi, Zhan Shi, Wentao Huang, Saumya Gupta, Chao Chen

机构 * Stony Brook University(石溪大学)

专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01546 2025-10-03 cs.CV cs.LG 70%

Growing Visual Generative Capacity for Pre-Trained MLLMs

Hanyu Wang, Jiaming Han, Ziyan Yang, Qi Zhao, Shanchuan Lin, Xiangyu Yue, Abhinav Shrivastava, Zhenheng Yang, Hao Chen

机构 * University of Maryland, College Park(马里兰大学学院公园分校) CUHK MMLab(香港大学MMLab) ByteDance(字节跳动)

专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract);分类 cs.CV

Comments Project page: https://hywang66.github.io/bridge/

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19118 2025-10-03 cs.CV 70%

Cross Spatial Temporal Fusion Attention for Remote Sensing Object Detection via Image Feature Matching

Abu Sadat Mohammad Salehin Amit, Xiaoli Zhang, Md Masum Billa Shagar, Zhaojun Liu, Xiongfei Li, Fanlong Meng

机构 * College of Computer Science and Technology(计算机科学与技术学院) Jilin University(吉林大学) College of Data Science(数据科学学院) Taiyuan University of Technology(太原理工大学) Department of Computer Science(计算机科学系)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01454 2025-10-03 cs.CV cs.LG 57%

Data Selection for Fine-tuning Vision Language Models via Cross Modal Alignment Trajectories

Nilay Naharas, Dang Nguyen, Nesihan Bulut, Mohammadhossein Bateni, Vahab Mirrokni, Baharan Mirzasoleiman

机构 * Department of Computer Science, University of California Los Angeles(加州大学洛杉矶分校计算机科学系) Google Research(谷歌研究)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

Comments 30 pages, 10 figures, 5 tables, link: https://bigml-cs-ucla.github.io/XMAS-project-page/

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.11997 2025-10-03 cs.LG cs.AI 57%

Can LLMs Find Fraudsters? Multi-level LLM Enhanced Graph Fraud Detection

Tairan Huang, Yili Wang, Qiutong Li, Changlong He, Jianliang Gao

机构 * Central South University(中南大学) Hongkong University of Science and Technology (Guangzhou)(香港科技大学(广州))

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

Comments Accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.19186 2025-10-03 cs.CV cs.RO 57%

LRFusionPR: A Polar BEV-Based LiDAR-Radar Fusion Network for Place Recognition

Zhangshuo Qi, Luqi Cheng, Zijie Zhou, Guangming Xiong

机构 * Beijing Institute of Technology(北京理工大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments Accepted by IEEE Robotics and Automation Letters (RAL) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01970 2025-10-03 cs.LG 50%

Moon: A Modality Conversion-based Efficient Multivariate Time Series Anomaly Detection

Yuanyuan Yao, Yuhan Shi, Lu Chen, Ziquan Fang, Yunjun Gao, Leong Hou U, Yushuai Li, Tianyi Li

机构 * College of Computer Science, Zhejiang University(浙江大学计算机科学学院) Department of Computer and Information Science, University of Macau(澳门大学计算机与信息科学系) Department of Computer Science, Aalborg University(奥胡斯大学计算机科学系)

专题命中 多模态训练与对齐 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01634 2025-10-03 cs.LG 50%

CAT: Curvature-Adaptive Transformers for Geometry-Aware Learning

Ryan Y. Lin, Siddhartha Ojha, Nicholas Bai

机构 * Division of Engineering and Applied Science(工程与应用科学系) California Institute of Technology(加州理工学院)

专题命中 多模态训练与对齐 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09404 2025-10-03 cs.LG 50%

Scaling Laws for Optimal Data Mixtures

Mustafa Shukor, Louis Bethune, Dan Busbridge, David Grangier, Enrico Fini, Alaaeldin El-Nouby, Pierre Ablin

机构 * Sorbonne University(索邦大学) Apple(苹果公司)

专题命中 多模态训练与对齐 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 其他多模态 4 篇

2510.01845 2025-10-03 cs.CL cs.CV 81%

Model Merging to Maintain Language-Only Performance in Developmentally Plausible Multimodal Models

Ece Takmaz, Lisa Bylinina, Jakub Dotlacil

机构 * Utrecht University(乌特勒支大学)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments Accepted to the EMNLP 2025 workshop BabyLM: Accelerating language modeling research with cognitively plausible datasets

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.17674 2025-10-03 cs.CL 74%

Push the Limit of Multi-modal Emotion Recognition by Prompting LLMs with Receptive-Field-Aware Attention Weighting

Han Zhang, Yu Lu, Liyun Zhang, Dian Ding, Dinghua Zhao, Yi-Chao Chen, Ye Wu, Guangtao Xue

机构 * Shanghai Jiao Tong University, Shanghai, China(上海交通大学) Xi'an Jiaotong-Liverpool University, Suzhou, China(西安交通大学-利物浦大学)

专题命中 其他多模态 :multi-modal(title);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.01730 2025-10-03 cs.NI 71%

A Deep Incremental Framework for Multi-Service Multi-Modal Devices in NextG AI-RAN Systems

Mrityunjoy Gain, Kitae Kim, Avi Deb Raha, Apurba Adhikary, Walid Saad, Zhu Han, Choong Seon Hong

专题命中 其他多模态 :multi-modal(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01506 2025-10-03 cond-mat.soft physics.flu-dyn 50%

Creases as elastocapillary gates for autonomous droplet control

Zixuan Wu, Gavin Linton, Stefan Karpitschka, Anupam Pandey

专题命中 其他多模态 :multimodal(abstract)

Comments 11 Pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏