arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-08-19 至 2025-08-19 共收录 10 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 10 篇

2508.13072 2025-08-19 cs.AI 83%

A Language-Signal-Vision Multimodal Framework for Multitask Cardiac Analysis

Yuting Zhang, Tiantian Geng, Luoying Hao, Xinxing Cheng, Alexander Thorley, Xiaoxia Wang, Wenqi Lu, Sandeep S Hothi, Lei Wei, Zhaowen Qiu, Dipak Kotecha, Jinming Duan

机构 * School of Computer Science, University of Birmingham, Birmingham, UK Department of Cardiovascular Sciences, University of Birmingham, Birmingham, UK NIHR Birmingham Biomedical Research Centre West Midlands NHS Secure Data Environment, University Hospitals Birmingham NHS Foundation Trust, Birmingham, UK Department of Computing Mathematics, Manchester Metropolitan University, Manchester, UK Department of Cardiology, Heart Lung Centre, Royal Wolverhampton NHS Trust, Wolverhampton, UK Department of Cardiovascular Surgery, The First Affiliated Hospital with Nanjing Medical University , Nanjing,China College of Computer Control Engineering, Northeast Forestry University, Harbin, China Julius Center, University Medical Center Utrecht, the Netherlands Data Sciences, University of Manchester, Manchester, UK

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12917 2025-08-19 cs.CV 83%

CMF-IoU: Multi-Stage Cross-Modal Fusion 3D Object Detection with IoU Joint Prediction

Zhiwei Ning, Zhaojiang Liu, Xuanang Gao, Yifan Zuo, Jie Yang, Yuming Fang, Wei Liu

机构 * School of Automation and Intelligent Sensing & Institute of Image Processing and Pattern Recognition, Shanghai Jiao Tong University(自动化与智能感知学院及图像处理与模式识别研究所,上海交通大学) School of Computing and Artificial Intelligence, Jiangxi University of Finance and Economics(计算机与人工智能学院,江西财经大学) School of Automation and Intelligent Sensing & Institute of Image Processing and Pattern Recognition & Institute of Medical Robotics, Shanghai Jiao Tong University(自动化与智能感知学院及图像处理与模式识别研究所及医学机器人研究所,上海交通大学)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV

Comments The Paper is Accepted by TCSVT

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02133 2025-08-19 cs.HC 82%

Hierarchical MoE: Continuous Multimodal Emotion Recognition with Incomplete and Asynchronous Inputs

Yitong Zhu, Lei Han, Guanxuan Jiang, PengYuan Zhou, Yuyang Wang

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11886 2025-08-19 cs.CV cs.AI cs.CL cs.LG eess.IV 82%

EVTP-IVS: Effective Visual Token Pruning For Unifying Instruction Visual Segmentation In Multi-Modal Large Language Models

Wenhui Zhu, Xiwen Chen, Zhipeng Wang, Shao Tang, Sayan Ghosh, Xuanzhao Dong, Rajat Koner, Yalin Wang

机构 * Arizona State University(亚利桑那州立大学) Clemson University(克莱姆森大学) LinkedIn Corporation(领英公司) Ludwig Maximilian University of Munich(慕尼黑路德维希-马克西米利安大学)

专题命中 多模态训练与对齐 :multi-modal(title);multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.11319 2025-08-19 cs.CV cs.AI 81%

GeoSAM: Fine-tuning SAM with Multi-Modal Prompts for Mobility Infrastructure Segmentation

Rafi Ibn Sultan, Chengyin Li, Hui Zhu, Prashant Khanduri, Marco Brocanelli, Dongxiao Zhu

机构 * Department of Computer Science, Wayne State University, Detroit, MI, USA 48202(计算机科学系,韦恩州立大学) Department of Radiation Oncology, Henry Ford Health, Detroit, MI, USA 48202(放射肿瘤科,亨利福特健康) Department of Electrical and Computer Engineering, The Ohio State University, Columbus, Ohio, USA 43210(电气与计算机工程系,俄亥俄州立大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted by European Conference on Artificial Intelligence (ECAI 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11737 2025-08-19 cs.CV cs.AI cs.CL cs.LG 67%

Ovis2.5 Technical Report

Shiyin Lu, Yang Li, Yu Xia, Yuwei Hu, Shanshan Zhao, Yanqing Ma, Zhichao Wei, Yinglun Li, Lunhao Duan, Jianshan Zhao, Yuxuan Han, Haijun Li, Wanying Chen, Junke Tang, Chengkun Hou, Zhixing Du, Tianli Zhou, Wenjie Zhang, Huping Ding, Jiahe Li, Wen Li, Gui Hu, Yiliang Gu, Siran Yang, Jiamang Wang, Hailong Sun, Yibo Wang, Hui Sun, Jinlong Huang, Yuping He, Shengze Shi, Weihong Zhang, Guodong Zheng, Junpeng Jiang, Sensen Gao, Yi-Feng Wu, Sijia Chen, Yuhui Chen, Qing-Guo Chen, Zhao Xu, Weihua Luo, Kaifu Zhang

机构 * Alibaba Group(阿里巴巴集团)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12036 2025-08-19 cs.CV cs.AI 62%

Q-FSRU: Quantum-Augmented Frequency-Spectral Fusion for Medical Visual Question Answering

Rakesh Thakur, Yusra Tariq

机构 * Amity Centre for Artificial Intelligence, Amity University, Noida(阿米蒂人工智能中心,阿米蒂大学,诺伊达)

专题命中 多模态训练与对齐 :image-text(abstract);分类 cs.CV、cs.AI

Comments 8 pages, 4 figures Submitted to AAAI 26

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12609 2025-08-19 cs.CV 57%

Not All Tokens and Heads Are Equally Important: Dual-Level Attention Intervention for Hallucination Mitigation

Lexiang Tang, Xianwei Zhuang, Bang Yang, Zhiyuan Hu, Hongxiang Li, Lu Ma, Jinghan Ru, Yuexian Zou

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.01212 2025-08-19 cs.CV cs.HC 57%

Towards Consumer-Grade Cybersickness Prediction: Multi-Model Alignment for Real-Time Vision-Only Inference

Yitong Zhu, Zhuowen Liang, Yiming Wu, Tangyao Li, Yuyang Wang

机构 * The Hong Kong University of Science and Technology(Guangzhou)(香港科学与技术大学(广州)) Nanyang Technological University(南洋理工大学) Nanyang Technological University Singapore(南洋理工大学新加坡)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12418 2025-08-19 cs.LG 50%

Bi-Axial Transformers: Addressing the Increasing Complexity of EHR Classification

Rachael DeVries, Casper Christensen, Marie Lisandra Zepeda Mendoza, Ole Winther

机构 * University of Copenhagen(哥本哈根大学) Novo Nordisk A/S(诺华北欧制药有限公司) Novo Nordisk Research Center Oxford Ltd.(诺华奥克斯研究中心) Technical University of Denmark(丹麦技术大学)

专题命中 多模态训练与对齐 :multi-modal(abstract)

Comments 18 pages, 7 figures. Submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏