arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-09-12 至 2025-09-12 共收录 40 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 2 篇

2509.09397 2025-09-12 cs.CV 70%

Decoupling Clinical and Class-Agnostic Features for Reliable Few-Shot Adaptation under Shift

Umaima Rahman, Raza Imam, Mohammad Yaqub, Dwarikanath Mahapatra

机构 * Mohamed Bin Zayed University of Artificial Intelligence(莫扎德人工智能大学) Khalifa University(卡利法大学)

专题命中 图文多模态 :cross-modal(abstract);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09311 2025-09-12 cs.CV 57%

Image Recognition with Vision and Language Embeddings of VLMs

Illia Volkov, Nikita Kisel, Klara Janouskova, Jiri Matas

机构 * Visual Recognition Group, Faculty of Electrical Engineering, Czech Technical University in Prague(视觉识别组,电气工程学院,布拉格捷克技术大学)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 音频语音多模态 2 篇

2404.02359 2025-09-12 cs.LG 78%

Attribution Regularization for Multimodal Paradigms

Sahiti Yerramilli, Jayant Sravan Tamarapalli, Jonathan Francis, Eric Nyberg

专题命中 音频语音多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.11538 2025-09-12 cs.CL cs.AI eess.AS 67%

MERaLiON-SpeechEncoder: Towards a Speech Foundation Model for Singapore and Beyond

Muhammad Huzaifah, Geyu Lin, Tianchi Liu, Hardik B. Sailor, Kye Min Tan, Tarun K. Vangani, Qiongqiong Wang, Jeremy H. M. Wong, Jinyang Wu, Nancy F. Chen, Ai Ti Aw

机构 * MERaLiON Team(MERaLiON团队) Institute for Infocomm Research (I 2 R), A*STAR, Singapore(信息通信研究所(I2R),A*STAR,新加坡)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、cs.AI、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频多模态 2 篇

2509.09584 2025-09-12 cs.CV cs.RO 57%

Visual Grounding from Event Cameras

Lingdong Kong, Dongyue Lu, Ao Liang, Rong Li, Yuhao Dong, Tianshuai Hu, Lai Xing Ng, Wei Tsang Ooi, Benoit R. Cottereau

机构 * NUS(新加坡国立大学) HKUST(GZ)(香港科技大学(广州)) NTU(南洋理工大学) HKUST(香港科技大学) I 2 R, A*STAR(新加坡科技研究局) IPAL, CNRS(法国国家科学研究中心IPAL) CerCo, CNRS(法国国家科学研究中心CerCo)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments Abstract Paper (Non-Archival) @ ICCV 2025 NeVi Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09263 2025-09-12 cs.CV 57%

DATE: Dynamic Absolute Time Enhancement for Long Video Understanding

Chao Yuan, Yang Yang, Yehui Yang, Zach Cheng

机构 * Beihang University(北京航空航天大学) Dcar, ByteDance(字节跳动Dcar部门) Qfin Holdings,Inc(Qfin控股公司) MAIS, Institute of Automation, Chinese Academy of Sciences(自动化研究所,中国科学院)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 跨模态检索 2 篇

2509.09118 2025-09-12 cs.CV 70%

Gradient-Attention Guided Dual-Masking Synergetic Framework for Robust Text-based Person Retrieval

Tianlu Zheng, Yifan Zhang, Xiang An, Ziyong Feng, Kaicheng Yang, Qichuan Ding

机构 * Northeastern University(东北大学) South China University of Technology(南方科技大学) DeepGlint

专题命中 跨模态检索 :cross-modal(abstract);image-text(abstract);分类 cs.CV

Comments Accepted by EMNLP2025 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09306 2025-09-12 eess.AS eess.IV 57%

Listening for "You": Enhancing Speech Image Retrieval via Target Speaker Extraction

Wenhao Yang, Jianguo Wei, Wenhuan Lu, Xinyue Song, Xianghu Yue

专题命中 跨模态检索 :multimodal(abstract);分类 eess.AS

Comments 5 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 多模态生成 4 篇

2509.09456 2025-09-12 cs.CV 83%

FlexiD-Fuse: Flexible number of inputs multi-modal medical image fusion based on diffusion model

Yushen Xu, Xiaosong Li, Yuchun Wang, Xiaoqi Cheng, Huafeng Li, Haishu Tan

机构 * School of Physics(物理学院) Optoelectronic Engineering, Foshan University(光电工程学院,佛山大学) Guangdong-HongKong-Macao Joint Laboratory for Intelligent Micro-Nano Optoelectronic Technology(粤港澳联合智能微纳光电技术实验室) Guangdong Provincial Key Laboratory of Industrial Intelligent Inspection Technology(广东省工业智能检测技术重点实验室) School of Information Engineering(信息工程学院) Automation, Kunming University of Science(自动化系,昆明理工大学)

专题命中 多模态生成 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

Journal ref Expert Systems with Applications, 2025: 128895

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08847 2025-09-12 cs.AI cs.CL cs.LG cs.SE 81%

Automated Unity Game Template Generation from GDDs via NLP and Multi-Modal LLMs

Amna Hassan

机构 * UET Taxila(塔希尔大学工程学院)

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20877 2025-09-12 cs.CV 74%

Deep Learning Framework for Early Detection of Pancreatic Cancer Using Multi-Modal Medical Imaging Analysis

Dennis Slobodzian, Amir Kordijazi

专题命中 多模态生成 :multi-modal(title);分类 cs.CV

Comments 21 pages, 17 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08775 2025-09-12 cs.RO 50%

Joint Model-based Model-free Diffusion for Planning with Constraints

Wonsuhk Jung, Utkarsh A. Mishra, Nadun Ranawaka Arachchige, Yongxin Chen, Danfei Xu, Shreyas Kousik

机构 * Georgia Institute of Technology(佐治亚理工学院)

专题命中 多模态生成 :multi-modal(abstract)

Comments The first two authors contributed equally. Last three authors advised equally. Accepted to CoRL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 多模态评测 13 篇

2509.09160 2025-09-12 cs.CL cs.AI 84%

Target-oriented Multimodal Sentiment Classification with Counterfactual-enhanced Debiasing

Zhiyue Liu, Fanrong Ma, Xin Ling

机构 * School of Computer, Electronics and Information(计算机、电子与信息学院) Guangxi University(广西大学) Guangxi Key Laboratory of Multimedia Communications(广西多媒体通信与网络技术重点实验室) School of Sociology and Anthropology(社会学与人类学学院) Sun Yat-sen University(中山大学)

专题命中 多模态评测 :multimodal(title,abstract);image-text(abstract);分类 cs.CL、cs.AI

Comments Accepted by the IEEE International Conference on Multimedia and Expo (ICME 2025). © 2025 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09307 2025-09-12 cs.CV cs.AI cs.CL cs.MM 83%

Can Multimodal LLMs See Materials Clearly? A Multimodal Benchmark on Materials Characterization

Zhengzhao Lai, Youbin Zheng, Zhenyang Cai, Haonan Lyu, Jinpu Yang, Hongqing Liang, Yan Hu, Benyou Wang

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.07084 2025-09-12 cs.RO 82%

DriveSOTIF: Advancing Perception SOTIF Through Multimodal Large Language Models

Shucheng Huang, Freda Shi, Chen Sun, Jiaming Zhong, Minghao Ning, Yufeng Yang, Yukun Lu, Hong Wang, Amir Khajepour

机构 * MVSLab, Department of Mechanical and Mechatronics Engineering, University of Waterloo(滑铁卢大学机械与机电工程系MVSLab) CompLING Lab, David R. Cheriton School of Computer Science, University of Waterloo(滑铁卢大学大卫·R·切里顿计算机科学学院CompLING Lab) Department of Data and Systems Engineering, University of Hong Kong(香港大学数据与系统工程系) Department of Mechanical Engineering, University of New Brunswick(新不伦瑞克大学机械工程系) School of Vehicle and Mobility, Tsinghua University(清华大学车辆与移动性学院)

专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract)

Comments This work has been accepted to IEEE Transactions on Vehicular Technology. Please refer to the copyright notice for additional information

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09254 2025-09-12 cs.CV cs.MM 81%

Towards Better Dental AI: A Multimodal Benchmark and Instruction Dataset for Panoramic X-ray Analysis

Jing Hao, Yuxuan Fan, Yanpeng Sun, Kaixin Guo, Lizhuo Lin, Jinrong Yang, Qi Yong H. Ai, Lun M. Wong, Hao Tang, Kuo Feng Hung

机构 * Faculty of Dentistry, The University of Hong Kong(香港大学牙科学院) The Hong Kong University of Science and Technology (GZ)(香港科学与技术大学) National University of Singapore(新加坡国立大学) CVTE Sun Yat-sen University(孙中山大学) Department of Diagnostic Radiology, The University of Hong Kong(香港大学放射科) Imaging and Interventional Radiology, Faculty of Medicine, The Chinese University of Hong Kong(香港中文大学医学院影像与介入放射科) School of Computer Science, Peking University(北京大学计算机科学系)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.MM

Comments 40 pages, 26 figures, 9 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09014 2025-09-12 cs.CV cs.CL 81%

COCO-Urdu: A Large-Scale Urdu Image-Caption Dataset with Multimodal Quality Estimation

Umair Hassan

机构 * Independent Researcher(独立研究者)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments 17 pages, 3 figures, 3 tables. Dataset available at https://huggingface.co/datasets/umairhassan02/urdu-translated-coco-captions-subset. Scripts and notebooks to reproduce results available at https://github.com/umair-hassan2/COCO-Urdu

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09190 2025-09-12 cs.CV 79%

VQualA 2025 Challenge on Visual Quality Comparison for Large Multimodal Models: Methods and Results

Hanwei Zhu, Haoning Wu, Zicheng Zhang, Lingyu Zhu, Yixuan Li, Peilin Chen, Shiqi Wang, Chris Wei Zhou, Linhan Cao, Wei Sun, Xiangyang Zhu, Weixia Zhang, Yucheng Zhu, Jing Liu, Dandan Zhu, Guangtao Zhai, Xiongkuo Min, Zhichao Zhang, Xinyue Li, Shubo Xu, Anh Dao, Yifan Li, Hongyuan Yu, Jiaojiao Yi, Yiding Tian, Yupeng Wu, Feiran Sun, Lijuan Liao, Song Jiang

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Comments ICCV VQualA Workshop 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.19662 2025-09-12 physics.ed-ph 78%

Multimodal large language models and physics visual tasks: comparative analysis of performance and costs

Giulia Polverini, Bor Gregorcic

专题命中 多模态评测 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.04077 2025-09-12 cs.CL cs.SD eess.AS 73%

A Novel Data Augmentation Approach for Automatic Speaking Assessment on Opinion Expressions

Chung-Chun Wang, Jhen-Ke Lin, Hao-Chien Lu, Hong-Yun Lin, Berlin Chen

机构 * National Taiwan Normal University(台湾国立台湾师范大学)

专题命中 多模态评测 :multimodal(abstract);cross-modal(abstract);分类 cs.CL、eess.AS

Comments submitted to the ISCA SLaTE-2025 Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09473 2025-09-12 cs.CL 57%

Mitigating Language Barriers in Education: Developing Multilingual Digital Learning Materials with Machine Translation

Lucie Poláková, Martin Popel, Věra Kloudová, Michal Novák, Mariia Anisimova, Jiří Balhar

机构 * Charles University, Faculty of Mathematics and Physics(查理大学数学与物理系)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CL

Comments 8 pages, 2 figures

Journal ref L. Poláková, M. Popel, V. Kloudová, M. Novák, M. Anisimova, J. Balhar (2025). Mitigating Language Barriers in Education: Developing Multilingual Digital Learning Materials with Machine Translation, EDULEARN25, pp. 8754-8760

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09324 2025-09-12 cs.CV 57%

Fine-Grained Customized Fashion Design with Image-into-Prompt benchmark and dataset from LMM

Hui Li, Yi You, Qiqi Chen, Bingfeng Zhang, George Q. Huang

机构 * The Hong Kong Polytechnic University, China(香港理工大学) China University of Petroleum (East China), China(中国石油大学(华东))

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09227 2025-09-12 eess.IV cs.CV 57%

Dynamic Structural Recovery Parameters Enhance Prediction of Visual Outcomes After Macular Hole Surgery

Yinzheng Zhao, Zhihao Zhao, Rundong Jiang, Louisa Sackewitz, Quanmin Liang, Mathias Maier, Daniel Zapp, Peter Charbel Issa, Mohammad Ali Nasseri

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

Comments TVST

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.10546 2025-09-12 cs.CV cs.RO 57%

The Oxford Spires Dataset: Benchmarking Large-Scale LiDAR-Visual Localisation, Reconstruction and Radiance Field Methods

Yifu Tao, Miguel Ángel Muñoz-Bañón, Lintong Zhang, Jiahao Wang, Lanke Frank Tarimo Fu, Maurice Fallon

机构 * Oxford Robotics Inst., Dept. of Eng. Science, Univ. of Oxford, UK(牛津大学机器人研究所、工程科学系) Group of Automation, Robotics and Computer Vision (AUROVA), University of Alicante, Spain(自动化、机器人与计算机视觉小组(AUROVA)、阿尔基兰特大学)

专题命中 多模态评测 :multi-modal(abstract);分类 cs.CV

Comments Accepted by IJRR. Website: https://dynamic.robots.ox.ac.uk/datasets/oxford-spires/

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05019 2025-09-12 cs.CE 50%

FinMultiTime: A Four-Modal Bilingual Dataset for Financial Time-Series Analysis

Wenyan Xu, Dawei Xiang, Yue Liu, Xiyu Wang, Yanxiang Ma, Liang Zhang, Shu Hu, Chang Xu, Jiaheng Zhang

专题命中 多模态评测 :multimodal(abstract)

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏

7. 多模态Agent 3 篇

2504.15917 2025-09-12 cs.SE 78%

Towards Test Generation from Task Description for Mobile Testing with Multi-modal Reasoning

Hieu Huynh, Hai Phung, Hao Pham, Tien N. Nguyen, Vu Nguyen

专题命中 多模态Agent :multi-modal(title,abstract)

Comments Change the method and experimentation

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09154 2025-09-12 cs.AI cs.CV 62%

Mind Meets Space: Rethinking Agentic Spatial Intelligence from a Neuroscience-inspired Perspective

Bui Duc Manh, Soumyaratna Debnath, Zetong Zhang, Shriram Damodaran, Arvind Kumar, Yueyi Zhang, Lu Mi, Erik Cambria, Lin Wang

机构 * School of EEE, Nanyang Technological University(南洋理工大学电子工程系) Department of Civil Engineering, Tsinghua University(清华大学土木工程系) Dr. B. R. Ambedkar National Institute of Technology, Jalandhar(B.R.阿姆贝卡尔国家理工学院,贾兰德赫) KTH Royal Institute of Technology(皇家理工学院) College of AI, Tsinghua University(清华大学人工智能学院) CCDS, Nanyang Technological University(南洋理工大学CCDS)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV、cs.AI

Comments 54 pages, journal

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03700 2025-09-12 cs.HC cs.AI 57%

MagicGUI: A Foundational Mobile GUI Agent with Scalable Data Pipeline and Reinforcement Fine-tuning

Liujian Tang, Shaokang Dong, Yijia Huang, Minqi Xiang, Hongtao Ruan, Bin Wang, Shuo Li, Zhiheng Xi, Zhihui Cao, Hailiang Pang, Heng Kong, He Yang, Mingxu Chai, Zhilin Gao, Xingyu Liu, Yingnan Fu, Jiaming Liu, Xuanjing Huang, Yu-Gang Jiang, Tao Gui, Qi Zhang, Kang Wang, Yunke Zhang, Yuran Wang

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

8. 多模态训练与对齐 9 篇

2505.19455 2025-09-12 cs.CV cs.AI cs.LG 84%

MM-Prompt: Cross-Modal Prompt Tuning for Continual Visual Question Answering

Xu Li, Fan Lyu

机构 * Khoury College of Computer Sciences, Northeastern University(东北大学克劳尔计算机科学学院) New Laboratory of Pattern Recognition, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所模式识别新实验室)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09427 2025-09-12 cs.CV 83%

FS-Diff: Semantic guidance and clarity-aware simultaneous multimodal image fusion and super-resolution

Yuchan Jie, Yushen Xu, Xiaosong Li, Fuqiang Zhou, Jianming Lv, Huafeng Li

机构 * School of Computer Science and Engineering, South China University of Technology(华南理工大学计算机科学与工程学院) School of Physics and Optoelectronic Engineering, Foshan University(佛山大学物理与光电工程学院) School of Instrumentation Science and Optelectronics Engineering, Beihang University(北航仪器科学与光电工程学院) School of Information Engineering and Automation, Kunming University of Science and Technology(昆明理工大学信息工程与自动化学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Journal ref Information Fusion, 2025, 121: 103146

详情

展开后加载摘要…

URL PDF HTML 收藏