arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-09-30 至 2025-09-30 共收录 24 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 24 篇

2509.24896 2025-09-30 cs.CV 86%

DAM: Dual Active Learning with Multimodal Foundation Model for Source-Free Domain Adaptation

Xi Chen, Hongxun Yao, Zhaopan Xu, Kui Jiang

机构 * Harbin Institute of Technology(哈尔滨工业大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);multimodal foundation model(title);分类 cs.CV

Comments 5 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24693 2025-09-30 q-bio.NC 86%

Brain Harmony: A Multimodal Foundation Model Unifying Morphology and Function into 1D Tokens

Zijian Dong, Ruilin Li, Joanna Su Xian Chong, Niousha Dehestani, Yinghui Teng, Yi Lin, Zhizhou Li, Yichi Zhang, Yapei Xie, Leon Qi Rong Ooi, B. T. Thomas Yeo, Juan Helen Zhou

专题命中 多模态训练与对齐 :multimodal(title,abstract);multimodal foundation model(title)

Comments NeurIPS 2025. The first two authors contributed equally

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24298 2025-09-30 cs.HC cs.AI cs.CL cs.CY cs.MM 85%

Bridging the behavior-neural gap: A multimodal AI reveals the brain's geometry of emotion more accurately than human self-reports

Changde Du, Yizhuo Lu, Zhongyu Huang, Yi Sun, Zisen Zhou, Shaozheng Qin, Huiguang He

机构 * State Key Laboratory of Brain Cognition and Brain-inspired Intelligence Technology, Institute of Automation, Chinese Academy of Sciences(脑认知与脑启发智能技术重点实验室,自动化研究所,中国科学院) School of Artificial Intelligence, University of Chinese Academy of Sciences(人工智能学院,中国科学院大学) School of Future Technology, University of Chinese Academy of Sciences(未来技术学院,中国科学院大学) State Key Laboratory of Cognitive Neuroscience and Learning, Beijing Normal University(认知神经科学与学习国家重点实验室,北京师范大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CL、cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22697 2025-09-30 cs.CV cs.AI cs.LG 84%

Learning Hyperspectral Images with Curated Text Prompts for Efficient Multimodal Alignment

Abhiroop Chatterjee, Susmita Ghosh

机构 * Jadavpur University(贾瓦帕尔大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted at the IEEE/CVF International Conference on Computer Vision (ICCV 2025), Workshop on Curated Data for Efficient Learning

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24505 2025-09-30 cs.CV 83%

Robust Multimodal Semantic Segmentation with Balanced Modality Contributions

Jiaqi Tan, Xu Zheng, Fangyu Li, Yang Liu

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) HKUST(GZ)(香港科技大学(广州)) INSAIT

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24776 2025-09-30 cs.CV cs.AI 81%

VTPerception-R1: Enhancing Multimodal Reasoning via Explicit Visual and Textual Perceptual Grounding

Yizhuo Ding, Mingkang Chen, Zhibang Feng, Tong Xiao, Wanying Qu, Wenqi Shao, Yanwei Fu

机构 * Fudan University(复旦大学) Shanghai AI Laboratory(上海人工智能实验室) The University of Hong Kong(香港大学) Shenzhen University(深圳大学) University of Science and Technology of China(中国科学技术大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24734 2025-09-30 cs.LG cs.AI cs.CV 81%

A TRIANGLE Enables Multimodal Alignment Beyond Cosine Similarity

Giordano Cicchetti, Eleonora Grassucci, Danilo Comminiello

机构 * Department of Information Engineering, Electronics, and Telecommunications(信息工程、电子与电信系)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23109 2025-09-30 cs.AI cs.CV 81%

AttAnchor: Guiding Cross-Modal Token Alignment in VLMs with Attention Anchors

Junyang Zhang, Tianyi Zhu, Thierry Tambe

机构 * California Institute of Technology(加州理工学院) Stanford University(斯坦福大学)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV、cs.AI

Comments 31 pages, 17 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22729 2025-09-30 cs.CL cs.AI 81%

Multi-Modal Sentiment Analysis with Dynamic Attention Fusion

Sadia Abdulhalim, Muaz Albaghdadi, Moshiur Farazi

机构 * University of Doha for Science and Technology(多哈科学技术大学)

专题命中 多模态训练与对齐 :multi-modal(title);multimodal(abstract);分类 cs.CL、cs.AI

Comments Paper accepted for presentation at the ACS/IEEE 22nd International Conference on Computer Systems and Applications (AICCSA 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25037 2025-09-30 cs.CL 79%

GateMABSA: Aspect-Image Gated Fusion for Multimodal Aspect-based Sentiment Analysis

Adamu Lawan, Haruna Yunusa

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

Comments 6 pages, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23677 2025-09-30 cs.CV 79%

MSD-KMamba: Bidirectional Spatial-Aware Multi-Modal 3D Brain Segmentation via Multi-scale Self-Distilled Fusion Strategy

Dayu Tan, Ziwei Zhang, Yansan Su, Xin Peng, Yike Dai, Chunhou Zheng, Weimin Zhong

机构 * Key Laboratory of Intelligent Computing and Signal Processing, Ministry of Education, Anhui University(智能计算与信号处理重点实验室,教育部,安徽大学) Key Laboratory of Smart Manufacturing in Energy Chemical Process, Ministry of Education, East China University of Science and Technology(能源化工过程智能制造重点实验室,教育部,东华大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.19904 2025-09-30 cs.RO cs.MM eess.SP 79%

WildFusion: Multimodal Implicit 3D Reconstructions in the Wild

Yanbaihui Liu, Boyuan Chen

机构 * Duke University(杜克大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.MM

Comments Our project website is at: http://generalroboticslab.com/WildFusion

Journal ref 2025 IEEE International Conference on Robotics and Automation (ICRA)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24431 2025-09-30 cs.LG 78%

Semantic Compression via Multimodal Representation Learning

Eleonora Grassucci, Giordano Cicchetti, Aurelio Uncini, Danilo Comminiello

机构 * Dept. of Information Engineering, Electronics, and Telecomm.(信息工程、电子与电信系)

专题命中 多模态训练与对齐 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24165 2025-09-30 cs.CV cs.AI 76%

LatXGen: Towards Radiation-Free and Accurate Quantitative Analysis of Sagittal Spinal Alignment Via Cross-Modal Radiographic View Synthesis

Moxin Zhao, Nan Meng, Jason Pui Yin Cheung, Chris Yuk Kwan Tang, Chenxi Yu, Wenting Zhong, Pengyu Lu, Chang Shi, Yipeng Zhuang, Teng Zhang

机构 * Department of Orthopaedics and Traumatology, The University of Hong Kong(香港大学骨科与创伤学系) Department of Joint Surgery, Shandong Provincial Hospital Affiliated to Shandong First Medical University(山东省第一医科大学附属山东省人民医院骨科)

专题命中 多模态训练与对齐 :cross-modal(title);分类 cs.CV、cs.AI

Comments 8 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23442 2025-09-30 eess.IV cs.AI cs.CV cs.LG eess.SP 76%

S$^3$F-Net: A Multi-Modal Approach to Medical Image Classification via Spatial-Spectral Summarizer Fusion Network

Md. Saiful Bari Siddiqui, Mohammed Imamul Hassan Bhuiyan

机构 * Department of Computer Science and Engineering, BRAC University(计算机科学与工程系,布拉克大学) Department of Electrical and Electronic Engineering, Bangladesh University of Engineering and Technology(电气与电子工程系,孟加拉国工程与技术大学)

专题命中 多模态训练与对齐 :multi-modal(title);分类 cs.CV、cs.AI

Comments Submitted to IEEE Journal of Biomedical and Health Informatics (JBHI). This preprint includes few additional details not present in the journal submission

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22853 2025-09-30 q-bio.QM cs.AI cs.CL cs.LG 73%

Patient-specific Biomolecular Instruction Tuning

Irsyad Adam, Zekai Chen, David Laub, Shaun Porwal, Arda Pekis, Kevin Brown

机构 * Standard Model Biomedicine(标准模型生物医学) University of California, San Diego(加州大学圣地亚哥分校)

专题命中 多模态训练与对齐 :multimodal(abstract);multi-modal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11425 2025-09-30 cs.SD cs.AI cs.CL eess.AS 67%

FuseCodec: Semantic-Contextual Fusion and Supervision for Neural Codecs

Md Mubtasim Ahasan, Rafat Hasan Khan, Tasnim Mohiuddin, Aman Chadha, Tariq Iqbal, M Ashraful Amin, Amin Ahsan Ali, Md Mofijul Islam, A K M Mahbubur Rahman

机构 * Center for Computational & Data Sciences, Independent University, Bangladesh(计算与数据科学中心,独立大学,孟加拉国) Amazon GenAI(亚马逊生成人工智能) Qatar Computing Research Institute(卡塔尔计算研究所) University of Virginia(弗吉尼亚大学)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CL、cs.AI、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22913 2025-09-30 cs.LG stat.ML 67%

Guided Manifold Alignment with Geometry-Regularized Twin Autoencoders

Jake S. Rhodes, Adam G. Rustad, Marshall S. Nielsen, Morgan Chase McClellan, Dallan Gardner, Dawson Hedges

机构 * Department of Statistics(统计学系) Department of Computer Science(计算机科学系) Neuroscience Center(神经科学中心) Neuroscience Center and Department of Psychology(神经科学中心和心理学系) Department of Psychology(心理学系)

专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract)

Comments 10 pages, 4 figures, 7 tables. Accepted at the MMAI workshop at ICDM, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24878 2025-09-30 cs.CV cs.RO 57%

ThermalGen: Style-Disentangled Flow-Based Generative Models for RGB-to-Thermal Image Translation

Jiuhong Xiao, Roshan Nayak, Ning Zhang, Daniel Tortei, Giuseppe Loianno

机构 * New York University(纽约大学) Technology Innovation Institute(技术创新研究所) University of California, Berkeley(加州大学伯克利分校)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

Comments 23 pages including the checklist and appendix. Accepted at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18640 2025-09-30 cs.LG cs.AI 57%

ThanoRA: Task Heterogeneity-Aware Multi-Task Low-Rank Adaptation

Jian Liang, Wenke Huang, Xianda Guo, Guancheng Wan, Bo Du, Mang Ye

机构 * Wuhan University(武汉大学) ByteDance(字节跳动) National University of Defense Technology(国防科技大学) Nanyang Technological University(南洋理工大学) The AGH University of Krakow(克拉科夫AGH大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23772 2025-09-30 cs.CV stat.AP 57%

A Modality-Tailored Graph Modeling Framework for Urban Region Representation via Contrastive Learning

Yaya Zhao, Kaiqi Zhao, Zixuan Tang, Zhiyuan Liu, Xiaoling Lu, Yalei Du

机构 * Center for Applied Statistics, School of Statistics, Innovation Platform, Renmin University of China(应用统计中心、统计学院、创新平台、中国人民大学) Harbin Institute of Technology, Shenzhen(哈尔滨工业大学深圳学院)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23641 2025-09-30 cs.CV cs.RO 57%

From Static to Dynamic: a Survey of Topology-Aware Perception in Autonomous Driving

Yixiao Chen, Ruining Yang, Xin Chen, Jia He, Dongliang Xu, Yue Yao

机构 * Sems Shandong University(山东大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments 13 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23273 2025-09-30 cs.CV 57%

SynDoc: A Hybrid Discriminative-Generative Framework for Enhancing Synthetic Domain-Adaptive Document Key Information Extraction

Yihao Ding, Soyeon Caren Han, Yanbei Jiang, Yan Li, Zechuan Li, Yifan Peng

机构 * The University of Western Australia(西澳大学) The University of Melbourne(墨尔本大学) The University of Sydney(悉尼大学) Weill Cornell Medicine, Cornell University(韦尔·科恩医学中心,康奈尔大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments Work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24381 2025-09-30 cs.DC 50%

RServe: Overlapping Encoding and Prefill for Efficient LMM Inference

Tianyu Guo, Tianming Xu, Xianjie Chen, Junru Chen, Nong Xiao, Xianwei Zhang

专题命中 多模态训练与对齐 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏