arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-09-30 至 2025-09-30 共收录 142 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 24 篇

2509.23109 2025-09-30 cs.AI cs.CV 81%

AttAnchor: Guiding Cross-Modal Token Alignment in VLMs with Attention Anchors

Junyang Zhang, Tianyi Zhu, Thierry Tambe

机构 * California Institute of Technology(加州理工学院) Stanford University(斯坦福大学)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV、cs.AI

Comments 31 pages, 17 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22729 2025-09-30 cs.CL cs.AI 81%

Multi-Modal Sentiment Analysis with Dynamic Attention Fusion

Sadia Abdulhalim, Muaz Albaghdadi, Moshiur Farazi

机构 * University of Doha for Science and Technology(多哈科学技术大学)

专题命中 多模态训练与对齐 :multi-modal(title);multimodal(abstract);分类 cs.CL、cs.AI

Comments Paper accepted for presentation at the ACS/IEEE 22nd International Conference on Computer Systems and Applications (AICCSA 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25037 2025-09-30 cs.CL 79%

GateMABSA: Aspect-Image Gated Fusion for Multimodal Aspect-based Sentiment Analysis

Adamu Lawan, Haruna Yunusa

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

Comments 6 pages, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23677 2025-09-30 cs.CV 79%

MSD-KMamba: Bidirectional Spatial-Aware Multi-Modal 3D Brain Segmentation via Multi-scale Self-Distilled Fusion Strategy

Dayu Tan, Ziwei Zhang, Yansan Su, Xin Peng, Yike Dai, Chunhou Zheng, Weimin Zhong

机构 * Key Laboratory of Intelligent Computing and Signal Processing, Ministry of Education, Anhui University(智能计算与信号处理重点实验室,教育部,安徽大学) Key Laboratory of Smart Manufacturing in Energy Chemical Process, Ministry of Education, East China University of Science and Technology(能源化工过程智能制造重点实验室,教育部,东华大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.19904 2025-09-30 cs.RO cs.MM eess.SP 79%

WildFusion: Multimodal Implicit 3D Reconstructions in the Wild

Yanbaihui Liu, Boyuan Chen

机构 * Duke University(杜克大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.MM

Comments Our project website is at: http://generalroboticslab.com/WildFusion

Journal ref 2025 IEEE International Conference on Robotics and Automation (ICRA)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24431 2025-09-30 cs.LG 78%

Semantic Compression via Multimodal Representation Learning

Eleonora Grassucci, Giordano Cicchetti, Aurelio Uncini, Danilo Comminiello

机构 * Dept. of Information Engineering, Electronics, and Telecomm.(信息工程、电子与电信系)

专题命中 多模态训练与对齐 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24165 2025-09-30 cs.CV cs.AI 76%

LatXGen: Towards Radiation-Free and Accurate Quantitative Analysis of Sagittal Spinal Alignment Via Cross-Modal Radiographic View Synthesis

Moxin Zhao, Nan Meng, Jason Pui Yin Cheung, Chris Yuk Kwan Tang, Chenxi Yu, Wenting Zhong, Pengyu Lu, Chang Shi, Yipeng Zhuang, Teng Zhang

机构 * Department of Orthopaedics and Traumatology, The University of Hong Kong(香港大学骨科与创伤学系) Department of Joint Surgery, Shandong Provincial Hospital Affiliated to Shandong First Medical University(山东省第一医科大学附属山东省人民医院骨科)

专题命中 多模态训练与对齐 :cross-modal(title);分类 cs.CV、cs.AI

Comments 8 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23442 2025-09-30 eess.IV cs.AI cs.CV cs.LG eess.SP 76%

S$^3$F-Net: A Multi-Modal Approach to Medical Image Classification via Spatial-Spectral Summarizer Fusion Network

Md. Saiful Bari Siddiqui, Mohammed Imamul Hassan Bhuiyan

机构 * Department of Computer Science and Engineering, BRAC University(计算机科学与工程系,布拉克大学) Department of Electrical and Electronic Engineering, Bangladesh University of Engineering and Technology(电气与电子工程系,孟加拉国工程与技术大学)

专题命中 多模态训练与对齐 :multi-modal(title);分类 cs.CV、cs.AI

Comments Submitted to IEEE Journal of Biomedical and Health Informatics (JBHI). This preprint includes few additional details not present in the journal submission

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22853 2025-09-30 q-bio.QM cs.AI cs.CL cs.LG 73%

Patient-specific Biomolecular Instruction Tuning

Irsyad Adam, Zekai Chen, David Laub, Shaun Porwal, Arda Pekis, Kevin Brown

机构 * Standard Model Biomedicine(标准模型生物医学) University of California, San Diego(加州大学圣地亚哥分校)

专题命中 多模态训练与对齐 :multimodal(abstract);multi-modal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11425 2025-09-30 cs.SD cs.AI cs.CL eess.AS 67%

FuseCodec: Semantic-Contextual Fusion and Supervision for Neural Codecs

Md Mubtasim Ahasan, Rafat Hasan Khan, Tasnim Mohiuddin, Aman Chadha, Tariq Iqbal, M Ashraful Amin, Amin Ahsan Ali, Md Mofijul Islam, A K M Mahbubur Rahman

机构 * Center for Computational & Data Sciences, Independent University, Bangladesh(计算与数据科学中心,独立大学,孟加拉国) Amazon GenAI(亚马逊生成人工智能) Qatar Computing Research Institute(卡塔尔计算研究所) University of Virginia(弗吉尼亚大学)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CL、cs.AI、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22913 2025-09-30 cs.LG stat.ML 67%

Guided Manifold Alignment with Geometry-Regularized Twin Autoencoders

Jake S. Rhodes, Adam G. Rustad, Marshall S. Nielsen, Morgan Chase McClellan, Dallan Gardner, Dawson Hedges

机构 * Department of Statistics(统计学系) Department of Computer Science(计算机科学系) Neuroscience Center(神经科学中心) Neuroscience Center and Department of Psychology(神经科学中心和心理学系) Department of Psychology(心理学系)

专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract)

Comments 10 pages, 4 figures, 7 tables. Accepted at the MMAI workshop at ICDM, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24878 2025-09-30 cs.CV cs.RO 57%

ThermalGen: Style-Disentangled Flow-Based Generative Models for RGB-to-Thermal Image Translation

Jiuhong Xiao, Roshan Nayak, Ning Zhang, Daniel Tortei, Giuseppe Loianno

机构 * New York University(纽约大学) Technology Innovation Institute(技术创新研究所) University of California, Berkeley(加州大学伯克利分校)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

Comments 23 pages including the checklist and appendix. Accepted at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18640 2025-09-30 cs.LG cs.AI 57%

ThanoRA: Task Heterogeneity-Aware Multi-Task Low-Rank Adaptation

Jian Liang, Wenke Huang, Xianda Guo, Guancheng Wan, Bo Du, Mang Ye

机构 * Wuhan University(武汉大学) ByteDance(字节跳动) National University of Defense Technology(国防科技大学) Nanyang Technological University(南洋理工大学) The AGH University of Krakow(克拉科夫AGH大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23772 2025-09-30 cs.CV stat.AP 57%

A Modality-Tailored Graph Modeling Framework for Urban Region Representation via Contrastive Learning

Yaya Zhao, Kaiqi Zhao, Zixuan Tang, Zhiyuan Liu, Xiaoling Lu, Yalei Du

机构 * Center for Applied Statistics, School of Statistics, Innovation Platform, Renmin University of China(应用统计中心、统计学院、创新平台、中国人民大学) Harbin Institute of Technology, Shenzhen(哈尔滨工业大学深圳学院)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23641 2025-09-30 cs.CV cs.RO 57%

From Static to Dynamic: a Survey of Topology-Aware Perception in Autonomous Driving

Yixiao Chen, Ruining Yang, Xin Chen, Jia He, Dongliang Xu, Yue Yao

机构 * Sems Shandong University(山东大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments 13 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23273 2025-09-30 cs.CV 57%

SynDoc: A Hybrid Discriminative-Generative Framework for Enhancing Synthetic Domain-Adaptive Document Key Information Extraction

Yihao Ding, Soyeon Caren Han, Yanbei Jiang, Yan Li, Zechuan Li, Yifan Peng

机构 * The University of Western Australia(西澳大学) The University of Melbourne(墨尔本大学) The University of Sydney(悉尼大学) Weill Cornell Medicine, Cornell University(韦尔·科恩医学中心,康奈尔大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments Work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24381 2025-09-30 cs.DC 50%

RServe: Overlapping Encoding and Prefill for Efficient LMM Inference

Tianyu Guo, Tianming Xu, Xianjie Chen, Junru Chen, Nong Xiao, Xianwei Zhang

专题命中 多模态训练与对齐 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 其他多模态 5 篇

2509.12227 2025-09-30 cs.LG cs.AI 79%

Learning to Route: Per-Sample Adaptive Routing for Multimodal Multitask Prediction

Marzieh Ajirak, Oded Bein, Ellen Rose Bowen, Dora Kanellopoulos, Avital Falk, Faith M. Gunning, Nili Solomonov, Logan Grosenick

机构 * Department of Psychiatry, Weill Cornell Medicine(威立·科恩医学部) Feil Family Brain & Mind Research Institute, Weill Cornell Medicine(费尔家族脑与心灵研究研究所)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23475 2025-09-30 cs.CV 79%

Robust Multi-Modal Face Anti-Spoofing with Domain Adaptation: Tackling Missing Modalities, Noisy Pseudo-Labels, and Model Degradation

Ming-Tsung Hsu, Fang-Yu Hsu, Yi-Ting Lin, Kai-Heng Chien, Jun-Ren Chen, Cheng-Hsiang Su, Yi-Chen Ou, Chiou-Ting Hsu, Pei-Kai Huang

机构 * College of Computer and Cyber Security, Fujian Normal University(计算机与网络安全部,福建师范大学) Department of Computer Science, National Tsing Hua University(计算机科学系,国立清华大学)

专题命中 其他多模态 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.03522 2025-09-30 stat.ML cond-mat.dis-nn cs.LG 78%

Optimal thresholds and algorithms for a model of multi-modal learning in high dimensions

Christian Keup, Lenka Zdeborová

机构 * Statistical Physics of Computation Laboratory(计算物理实验室)

专题命中 其他多模态 :multi-modal(title,abstract)

Journal ref J. Stat. Mech. (2025) 093302

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24873 2025-09-30 cs.LG cs.AI 57%

Uncertainty-Guided Expert-AI Collaboration for Efficient Soil Horizon Annotation

Teodor Chiaburu, Vipin Singh, Frank Haußer, Felix Bießmann

机构 * Einstein Center Digital Future Berlin Germany Einstein Center Digital Future

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

Comments 11 pages, 7 figures, presented at ECAI 2025, CLEAR-AI Workshop, Bologna

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24124 2025-09-30 cs.RO cs.AI cs.LG 57%

Ancestry Tree Clustering for Particle Filter Diversity Maintenance

Ilari Vallivaara, Bingnan Duan, Yinhuan Dong, Tughrul Arslan

机构 * Visiting Research Fellow University of Edinburgh Edinburgh, UK(访问研究员 爱丁堡大学 英国爱丁堡) School of Engineering University of Edinburgh Edinburgh, UK(工程学院 爱丁堡大学 英国爱丁堡)

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

Comments 15th International Conference on Indoor Positioning and Indoor Navigation, 15-18 September 2025, Tampere, Finland Originally 8 pages. The online version with appendices is 14 pages

详情

展开后加载摘要…

URL PDF HTML 收藏