arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-09-05 至 2025-09-05 共收录 33 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 4 篇

2509.04324 2025-09-05 cs.RO cs.CV 79%

OVGrasp: Open-Vocabulary Grasping Assistance via Multimodal Intent Detection

Chen Hu, Shan Luo, Letizia Gionfrida

机构 * Department of Informatics, King's College London(伦敦国王学院信息学院) Department of Engineering, King's College London(伦敦国王学院工程学院) John A. Paulson School of Engineering and Applied Sciences, Harvard University(哈佛大学约翰·A·保罗森工程与应用科学学院)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14904 2025-09-05 cs.CV cs.AI 73%

TriCLIP-3D: A Unified Parameter-Efficient Framework for Tri-Modal 3D Visual Grounding based on CLIP

Fan Li, Zanyi Wang, Zeyi Huang, Guang Dai, Jingdong Wang, Mengmeng Wang

机构 * Xi’an Jiaotong University(西安交通大学) SGIT AI Lab(SGIT人工智能实验室) Zhejiang University of Technology(浙江工业大学) Huawei(华为)

专题命中 图文多模态 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03800 2025-09-05 cs.CV 57%

MedVista3D: Vision-Language Modeling for Reducing Diagnostic Errors in 3D CT Disease Detection, Understanding and Reporting

Yuheng Li, Yenho Chen, Yuxiang Lai, Jike Zhong, Vanessa Wildman, Xiaofeng Yang

机构 * Department of Biomedical Engineering(生物医学工程系) Georgia Institute of Technology(佐治亚理工学院) Department of Machine Learning(机器学习系) Department of Radiation Oncology(放射肿瘤科) Emory University School of Medicine(埃默里大学医学院) University of Southern California(南加州大学)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04162 2025-09-05 cs.AR 50%

Real Time FPGA Based Transformers & VLMs for Vision Tasks: SOTA Designs and Optimizations

Safa Mohammed Sali, Mahmoud Meribout, Ashiyana Abdul Majeed

专题命中 图文多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 音频语音多模态 3 篇

2509.04215 2025-09-05 cs.SD cs.IR cs.MM 79%

PianoBind: A Multimodal Joint Embedding Model for Pop-piano Music

Hayeon Bang, Eunjin Choi, Seungheon Doh, Juhan Nam

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.MM

Comments Accepted for publication at the 26th International Society for Music Information Retrieval Conference (ISMIR 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04392 2025-09-05 cs.SD 50%

Denoising GER: A Noise-Robust Generative Error Correction with LLM for Speech Recognition

Yanyan Liu, Minqiang Xu, Yihao Chen, Liang He, Lei Fang, Sian Fang, Lin Liu

机构 * School of Computer Science and Technology(计算机科学与技术学院) Xinjiang University(新疆大学) Hefei iFly Digital Technology Co. Ltd.(合肥iFly数字技术有限公司) University of Science and Technology of China(中国科学技术大学) Tsinghua University(清华大学)

专题命中 音频语音多模态 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04356 2025-09-05 cs.HC cs.RO 50%

SRWToolkit: An Open Source Wizard of Oz Toolkit to Create Social Robotic Avatars

Atikkhan Faridkhan Nilgar, Kristof Van Laerhoven, Ayub Kinoti

机构 * University of Siegen(施皮格恩大学) Honda Research Institute Europe GmbH(本田欧洲研究院) Dedan Kimathi University of Technology(德丹·基马蒂技术大学)

专题命中 音频语音多模态 :multimodal(abstract)

Journal ref 2025 International Conference on Social Robotics (ICSR)

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频多模态 5 篇

2509.04254 2025-09-05 cs.HC 82%

MuMTAffect: A Multimodal Multitask Affective Framework for Personality and Emotion Recognition from Physiological Signals

Meisam Jamshidi Seikavandi, Fabricio Batista Narcizo, Ted Vucurevich, Andrew Burke Dittberner, Paolo Burelli

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04330 2025-09-05 cs.IR 78%

Temporal Interest-Driven Multimodal Personalized Content Generation

Tian Miao

专题命中 视频多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04210 2025-09-05 cs.CE cs.LG 78%

COBRA: Multimodal Sensing Deep Learning Framework for Remote Chronic Obesity Management via Wrist-Worn Activity Monitoring

Zhengyang Shen, Bo Gao, Mayue Shi

机构 * Department of Electrical and Electronic Engineering, Imperial College London, London SW7 2AZ, UK(帝国理工学院电子与电气工程系) Institute of Biomedical Engineering, Department of Engineering Science, University of Oxford, Oxford OX3 7DQ, UK(牛津大学生物医学工程研究所)

专题命中 视频多模态 :multimodal(title,abstract)

Comments 19 pages, 4 figures. *Correspondence: m.shi16@imperial.ac.uk. Accepted by the IUPESM World Congress on Medical Physics and Biomedical Engineering 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19346 2025-09-05 cs.LG 78%

Short-Form Video Recommendations with Multimodal Embeddings: Addressing Cold-Start and Bias Challenges

Andrii Dzhoha, Katya Mirylenka, Egor Malykh, Marco-Andrea Buchmann, Francesca Catino

机构 * Zalando SE Berlin Germany(泽尔安多德国分公司) Zalando Switzerland AG Zürich Switzerland(泽尔安多瑞士分公司)

专题命中 视频多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04117 2025-09-05 cs.CV 57%

DVS-PedX: Synthetic-and-Real Event-Based Pedestrian Dataset

Mustafa Sakhai, Kaung Sithu, Min Khant Soe Oke, Maciej Wielgosz

机构 * Faculty of Computer Science, Electronics and Telecommunications(计算机科学与电子技术学院) AGH University of Science and Technology(AGH科技大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments 12 pages, 8 figures, 3 tables; dataset descriptor paper introducing DVS-PedX (synthetic-and-real event-based pedestrian dataset with baselines) External URL: https://doi.org/10.5281/zenodo.17030898

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 多模态生成 4 篇

2509.03535 2025-09-05 cs.CL cs.AI 81%

QuesGenie: Intelligent Multimodal Question Generation

Ahmed Mubarak, Amna Ahmed, Amira Nasser, Aya Mohamed, Fares El-Sadek, Mohammed Ahmed, Ahmed Salah, Youssef Sobhy

专题命中 多模态生成 :multimodal(title);multi-modal(abstract);分类 cs.CL、cs.AI

Comments 7 pages, 8 figures, 12 tables. Supervised by Dr. Ahmed Salah and TA Youssef Sobhy

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04269 2025-09-05 cs.CV 57%

TauGenNet: Plasma-Driven Tau PET Image Synthesis via Text-Guided 3D Diffusion Models

Yuxin Gong, Se-in Jang, Wei Shao, Yi Su, Kuang Gong

机构 * J. Crayton Pruitt Family Department of Biomedical Engineering, University of Florida(J. Crayton Pruitt家族生物医学工程系,佛罗里达大学) Department of Radiology & Biomedical Imaging, Yale University(放射学与生物医学成像系,耶鲁大学) Department of Medicine, University of Florida(医学系,佛罗里达大学) Banner Alzheimer’s Institute(Banner阿尔茨海默症研究所)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments 9 pages, 4 figures, submitted to IEEE Transactions on Radiation and Plasma Medical Sciences

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03550 2025-09-05 cs.AI 57%

Diffusion-RL Based Air Traffic Conflict Detection and Resolution Method

Tonghe Li, Jixin Liu, Weili Zeng, Hao Jiang

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI

Comments 59 pages,13 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.16507 2025-09-05 cs.CV cs.LG 57%

Straighter Flow Matching via a Diffusion-Based Coupling Prior

Siyu Xing, Jie Cao, Huaibo Huang, Haichao Shi, Xiao-Yu Zhang

机构 * Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络空间安全学院) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 多模态评测 8 篇

2509.03986 2025-09-05 cs.CV cs.AI cs.CL cs.LG 82%

Promptception: How Sensitive Are Large Multimodal Models to Prompts?

Mohamed Insaf Ismithdeen, Muhammad Uzair Khattak, Salman Khan

机构 * Mohamed Bin Zayed University of Artificial Intelligence(莫罕默德·本·扎耶德人工智能大学) Swiss Federal Institute of Technology Lausanne (EPFL)(洛桑联邦理工学院) Australian National University(澳大利亚国立大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted to EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03529 2025-09-05 cs.CL cs.AI eess.AS 82%

Multimodal Proposal for an AI-Based Tool to Increase Cross-Assessment of Messages

Alejandro Álvarez Castro, Joaquín Ordieres-Meré

机构 * AI master(人工智能硕士) Universidad Politécnica de Madrid(马德里理工大学)

专题命中 多模态评测 :multimodal(title);multi-modal(abstract);分类 cs.CL、cs.AI、eess.AS

Comments Presented at NLMLT2025 (https://airccse.org/csit/V15N16.html), 15 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02175 2025-09-05 cs.CV cs.AI cs.CL cs.LG 67%

Understanding Space Is Rocket Science -- Only Top Reasoning Models Can Solve Spatial Understanding Tasks

Nils Hoehing, Mayug Maniparambil, Ellen Rushe, Noel E. O'Connor, Anthony Ventresque

机构 * School of Computer Science(计算机科学学院) University College Dublin(都柏林大学) School of Computing(计算机科学学院) Dublin City University(都柏林城市大学) School of Electronic Engineering(电子工程学院) Trinity College Dublin(都柏林三一学院)

专题命中 多模态评测 :image-text(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03791 2025-09-05 cs.CL cs.AI 62%

SiLVERScore: Semantically-Aware Embeddings for Sign Language Generation Evaluation

Saki Imai, Mert İnan, Anthony Sicilia, Malihe Alikhani

机构 * Northeastern University(东北大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04326 2025-09-05 cs.CV 57%

Efficient Odd-One-Out Anomaly Detection

Silvio Chito, Paolo Rabino, Tatiana Tommasi

机构 * Politecnico di Torino(托斯尼亚理工学院)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

Comments Accepted at ICIAP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03741 2025-09-05 cs.HC cs.AI 57%

Designing Gaze Analytics for ELA Instruction: A User-Centered Dashboard with Conversational AI Support

Eduardo Davalos, Yike Zhang, Shruti Jain, Namrata Srivastava, Trieu Truong, Nafees-ul Haque, Tristan Van, Jorge Salas, Sara McFadden, Sun-Joo Cho, Gautam Biswas, Amanda Goodwin

机构 * Trinity University(特里尼蒂大学) St. Mary's University(圣玛丽大学) Vanderbilt University(范德比大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.AI

Comments 22 pages, 9 figures, 3 tables, submitted to IUI2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02349 2025-09-05 cs.SD cs.AI cs.LG 57%

AudioCodecBench: A Comprehensive Benchmark for Audio Codec Evaluation

Lu Wang, Hao Chen, Siyu Wu, Zhiyue Wu, Hao Zhou, Chengfeng Zhang, Ting Wang, Haodi Zhang

机构 * Lu Wang(卢王) Hao Chen(何晨) Siyu Wu(武士) Zhiyue Wu(吴致岳) Hao Zhou(周浩) Chengfeng Zhang(张成峰) Ting Wang(王婷) Haodi Zhang(张浩迪)

专题命中 多模态评测 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04322 2025-09-05 cs.LG 50%

Characteristic Energy Behavior Profiling of Non-Residential Buildings

Haley Dozier, Althea Henslee

机构 * Information Technology Laboratory U.S. Army Engineer Research Development Center Vicksburg, M.S. U.S.A

专题命中 多模态评测 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 多模态Agent 1 篇

2509.03536 2025-09-05 cs.AI cs.HC 57%

PG-Agent: An Agent Powered by Page Graph

Weizhi Chen, Ziwei Wang, Leyang Yang, Sheng Zhou, Xiaoxuan Tang, Jiajun Bu, Yong Li, Wei Jiang

机构 * Zhejiang Key Lab of Accessible Perception \& Intelligent Systems, Zhejiang University Hangzhou China Zhejiang University

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

Comments Paper accepted to ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

7. 多模态训练与对齐 6 篇

2506.09556 2025-09-05 cs.CL 83%

MEDUSA: A Multimodal Deep Fusion Multi-Stage Training Framework for Speech Emotion Recognition in Naturalistic Conditions

Georgios Chatzichristodoulou, Despoina Kosmopoulou, Antonios Kritikos, Anastasia Poulopoulou, Efthymios Georgiou, Athanasios Katsamanis, Vassilis Katsouros, Alexandros Potamianos

机构 * National Technical University of Athens(希腊国家技术大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL

Comments Interspeech 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03961 2025-09-05 cs.CV cs.AI 81%

Multimodal Feature Fusion Network with Text Difference Enhancement for Remote Sensing Change Detection

Yijun Zhou, Yikui Zhai, Zilu Ying, Tingfeng Xian, Wenlve Zhou, Zhiheng Zhou, Xiaolin Tian, Xudong Jia, Hongsheng Zhang, C. L. Philip Chen

机构 * College of Electronics and Information Engineering, Wuyi University(威怡大学电子与信息工程学院) School of Electronic and Information Engineering and the Key Laboratory of Big Data and Intelligent Robot, Ministry of Education, South China University of Technology(电子与信息工程学院和大数据与智能机器人重点实验室,华南理工大学) State Key Laboratory of Lunar and Planetary Sciences, Macau University of Science and Technology(澳门大学地球和行星科学国家重点实验室) College of Engineering and Computer Science, California State University, Northridge(工程与计算机科学学院,加州大学北岭分校) Department of Geography, The University of Hong Kong(地理系,香港大学) Faculty of Computer Science and Engineering, S(计算机科学与工程学院,S)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03999 2025-09-05 cs.CV 79%

SliceSemOcc: Vertical Slice Based Multimodal 3D Semantic Occupancy Representation

Han Huang, Han Sun, Ningzhong Liu, Huiyu Zhou, Jiaquan Shen

机构 * Nanjing University of Aeronautics and Astronautics(南京航空航天大学) University of Leicester(莱斯特大学) Luoyang Normal University(洛阳师范学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments 14 pages, accepted by PRCV2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03837 2025-09-05 cs.LG cs.IT math.IT 78%

Vehicle-to-Infrastructure Collaborative Spatial Perception via Multimodal Large Language Models

Kimia Ehsani, Walid Saad

机构 * Bradley Department of Electrical and Computer Engineering(电气与计算机工程系)

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Comments Accepted at IEEE GLOBECOM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03872 2025-09-05 cs.CV 57%

Focus Through Motion: RGB-Event Collaborative Token Sparsification for Efficient Object Detection

Nan Yang, Yang Wang, Zhanwen Liu, Yuchao Dai, Yang Liu, Xiangmo Zhao

机构 * School of Information Engineering, Chang’an University(信息工程学院,长安大学) School of Electronics and Information, Northwestern Polytechnical University(电子信息学院,西北工业大学) School of Vehicle and Mobility, Tsinghua University(车辆与移动学院,清华大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏