arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-09-25 至 2025-09-25 共收录 59 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 14 篇

2410.17241 2025-09-25 eess.IV cs.CV 61%

Frontiers in Intelligent Colonoscopy

Ge-Peng Ji, Jingyi Liu, Peng Xu, Nick Barnes, Fahad Shahbaz Khan, Salman Khan, Deng-Ping Fan

机构 * Nankai Institute of Advanced Research (SHENZHEN FUTIAN)(南开先进研究院(深圳福田)) College of Computer Science & VCIP, Nankai University(计算机科学与VCIP学院,南开大学) School of Computing, Australian National University(计算学院,澳大利亚国立大学) Graduate School of Science and Technology, Keio University(科学与技术研究生院,庆应大学) Department of Electronic Engineering, Tsinghua University(电子工程系,清华大学) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)

专题命中 多模态评测 :multimodal(abstract,comments);分类 cs.CV

Comments [Work in progress] A comprehensive survey of intelligent colonoscopy in the multimodal era. [Updated Version V2] New training strategy for colonoscopy-specific multimodal language model

Journal ref Machine Intelligence Research 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.18583 2025-09-25 cs.CL 57%

What are Foundation Models Cooking in the Post-Soviet World?

Anton Lavrouk, Tarek Naous, Alan Ritter, Wei Xu

机构 * Georgia Institute of Technology(佐治亚理工学院)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CL

Comments Accepted to EMNLP 2025 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19926 2025-09-25 cs.LG 50%

MMSE-Calibrated Few-Shot Prompting for Alzheimer's Detection

Jana Sweidan, Mounim A. El-Yacoubi, Nasredine Semmar

专题命中 多模态评测 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19856 2025-09-25 cs.LG 50%

Oversampling and Downsampling with Core-Boundary Awareness: A Data Quality-Driven Approach

Samir Brahim Belhaouari, Yunis Carreon Kahalan, Humaira Shaffique, Ismael Belhaouari, Ashhadul Islam

机构 * Hamad Bin Khalifa University(哈马德·本·卡尔法大学) Maastricht University(马斯特里赫特大学) KTH Royal Institute of Technology(皇家理工学院)

专题命中 多模态评测 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19342 2025-09-25 eess.SP cs.IT cs.LG math.IT 50%

A Measurement Report Data-Driven Framework for Localized Statistical Channel Modeling

Xinyu Qin, Ye Xue, Qi Yan, Shutao Zhang, Bingsheng Peng, Tsung-Hui Chang

机构 * Shenzhen Research Institute of Big Data, School of Science and Engineering, The Chinese University of Hong Kong, Shenzhen, Guangdong(深圳大数据研究院,科学与工程学院,香港中文大学(深圳)) Shenzhen Research Institute of Big Data, School of Data Science, The Chinese University of Hong Kong, Shenzhen, Guangdong(深圳大数据研究院,数据科学学院,香港中文大学(深圳)) Networking and User Experience Lab, Huawei Technologies(网络与用户体验实验室,华为技术)

专题命中 多模态评测 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 多模态Agent 4 篇

2509.20021 2025-09-25 cs.AI cs.CL cs.RO 73%

Embodied AI: From LLMs to World Models

Tongtong Feng, Xin Wang, Yu-Gang Jiang, Wenwu Zhu

机构 * Department of Computer Science and Technology, Tsinghua University(清华大学计算机科学与技术系) Beijing National Research Center for Information Science and Technology(北京信息科学与技术国家研究中心) Institute of Trustworthy Embodied AI, Fudan University(复旦大学可信具身人工智能研究院)

专题命中 多模态Agent :multimodal(abstract);MLLM(abstract);分类 cs.CL、cs.AI

Comments Accepted by IEEE CASM

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.15275 2025-09-25 cs.CV cs.AI 62%

Challenges and Trends in Egocentric Vision: A Survey

Xiang Li, Heqian Qiu, Lanxiao Wang, Hanwen Zhang, Chenghao Qi, Linfeng Han, Huiyu Xiong, Hongliang Li

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV、cs.AI

Comments This article was accepted by Machine Intelligence Research

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.14480 2025-09-25 cs.RO cs.CL cs.CV cs.LG 62%

GraphEQA: Using 3D Semantic Scene Graphs for Real-time Embodied Question Answering

Saumya Saxena, Blake Buchanan, Chris Paxton, Peiqi Liu, Bingqing Chen, Narunas Vaskevicius, Luigi Palmieri, Jonathan Francis, Oliver Kroemer

机构 * Carnegie Mellon University(卡内基梅隆大学) Neya Systems(Neya系统) Agility Robotics(敏捷机器人) Hello Robot Bosch Center for AI(博世人工智能中心)

专题命中 多模态Agent :multi-modal(abstract);分类 cs.CV、cs.CL

Comments Project website: https://saumyasaxena.github.io/grapheqa

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.16803 2025-09-25 cs.RO cs.CV cs.NI eess.IV 57%

RG-Attn: Radian Glue Attention for Multi-modality Multi-agent Cooperative Perception

Lantao Li, Kang Yang, Wenqi Zhang, Xiaoxue Wang, Chen Sun

机构 * Sony (China) Limited(索尼(中国)有限公司) Renmin University of China(中国人民大学)

专题命中 多模态Agent :cross-modal(abstract);分类 cs.CV

Comments Accepted by ICCV 2025 DriveX workshop (Final Version)

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 多模态训练与对齐 13 篇

2509.19875 2025-09-25 cs.CV cs.AI 84%

Adaptive Guidance Semantically Enhanced via Multimodal LLM for Edge-Cloud Object Detection

Yunqing Hu, Zheming Yang, Chang Zhao, Wen Ji

机构 * Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所) Institute of AI for Industries(工业人工智能研究所) University of Chinese Academy of Sciences(中国科学院大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20022 2025-09-25 cs.CV 83%

PS3: A Multimodal Transformer Integrating Pathology Reports with Histology Images and Biological Pathways for Cancer Survival Prediction

Manahil Raza, Ayesha Azam, Talha Qaiser, Nasir Rajpoot

机构 * University of Warwick, UK(沃里克大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments Accepted at ICCV 2025. Copyright 2025 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19628 2025-09-25 cs.CE cs.CL q-fin.CP 83%

Multimodal Language Models with Modality-Specific Experts for Financial Forecasting from Interleaved Sequences of Text and Time Series

Ross Koval, Nicholas Andrews, Xifeng Yan

机构 * University of California, Santa Barbara(加州大学圣巴巴拉分校) Johns Hopkins University(约翰霍普金斯大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20225 2025-09-25 cs.IR cs.AI 79%

Multimodal Representation-disentangled Information Bottleneck for Multimodal Recommendation

Hui Wang, Jinghui Qin, Wushao Wen, Qingling Li, Shanshan Zhong, Zhongzhan Huang

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19719 2025-09-25 cs.CV 79%

Frequency-domain Multi-modal Fusion for Language-guided Medical Image Segmentation

Bo Yu, Jianhua Yang, Zetao Du, Yan Huang, Chenglong Li, Liang Wang

机构 * School of Computer Science and Technology, Anhui University(安徽大学计算机科学与技术学院) NLPR, MAIS, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) School of Information Science and Technology, ShanghaiTech University(上海科技大学信息科学与技术学院) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) School of Artificial Intelligence, Anhui University(安徽大学人工智能学院)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted by MICCAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20306 2025-09-25 cs.AI eess.IV q-bio.QM 79%

Multi-Modal Artificial Intelligence of Embryo Grading and Pregnancy Prediction in Assisted Reproductive Technology: A Review

Xueqiang Ouyang, Jia Wei

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.10679 2025-09-25 cs.CV 70%

SMLNet: A SPD Manifold Learning Network for Infrared and Visible Image Fusion

Huan Kang, Hui Li, Tianyang Xu, Xiao-Jun Wu, Rui Wang, Chunyang Cheng, Josef Kittler

机构 * School of Artificial Intelligence and Computer Science(人工智能与计算机科学学院) Jiangnan University(江南大学) Centre for Vision, Speech and Signal Processing(视觉、语音与信号处理中心) University of Surrey(Surrey大学)

专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

Comments 23 pages, 17 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20240 2025-09-25 cs.LG cs.AI 57%

A HyperGraphMamba-Based Multichannel Adaptive Model for ncRNA Classification

Xin An, Ruijie Li, Qiao Ning, Hui Li, Qian Ma, Shikai Guo

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

Comments 9 pages, 17 figures (including subfigures), 1 table. Xin An and Ruijie Li contributed equally to this work and should be considered co-first authors

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13739 2025-09-25 cs.CV 57%

Enhancing Targeted Adversarial Attacks on Large Vision-Language Models via Intermediate Projector

Yiming Cao, Yanjie Li, Kaisheng Liang, Bin Xiao

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19733 2025-09-25 cs.CV 57%

Robust RGB-T Tracking via Learnable Visual Fourier Prompt Fine-tuning and Modality Fusion Prompt Generation

Hongtao Yang, Bineng Zhong, Qihua Liang, Zhiruo Zhu, Yaozong Zheng, Ning Li

机构 * Key Laboratory of Education Blockchain and Intelligent Technology, Ministry of Education, Guangxi Normal University(教育区块链与智能技术重点实验室,教育部,广西师范大学) Guangxi Key Lab of Multi-Source Information Mining and Security, Guangxi Normal University(多源信息挖掘与安全广西重点实验室,广西师范大学)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

Comments Accepted by TMM2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16549 2025-09-25 cs.CV 57%

Efficient Rectified Flow for Image Fusion

Zirui Wang, Jiayi Zhang, Tianwei Guan, Yuhan Zhou, Xingyuan Li, Minjing Dong, Jinyuan Liu

机构 * City University of Hong Kong(香港城市大学) Dalian University of Technology(大连理工大学) Chinese University of Hong Kong(香港中文大学) Zhejiang University(浙江大学)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.19860 2025-09-25 cs.CV cs.LG 57%

SpaRC: Sparse Radar-Camera Fusion for 3D Object Detection

Philipp Wolters, Johannes Gilg, Torben Teepe, Fabian Herzog, Felix Fent, Gerhard Rigoll

机构 * Technical University of Munich(慕尼黑技术大学)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

Comments 18 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19521 2025-09-25 cs.RO 50%

A Bimanual Gesture Interface for ROS-Based Mobile Manipulators Using TinyML and Sensor Fusion

Najeeb Ahmed Bhuiyan, M. Nasimul Huq, Sakib H. Chowdhury, Rahul Mangharam

机构 * †‡Department of Mechatronics Engineering, Rajshahi University of Engineering \& Technology, Kazla, Rajshahi-6204, Bangladesh §Department of Electrical \& Systems Engineering, School of Engineering \& Applied Science, University of Pennsylvania, Philadelphia, PA 19104, United States Email: , †, ‡, §

专题命中 多模态训练与对齐 :multimodal(abstract)

Comments 12 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 其他多模态 7 篇

2509.19352 2025-09-25 cs.CL cs.AI 81%

TriSPrompt: A Hierarchical Soft Prompt Model for Multimodal Rumor Detection with Incomplete Modalities

Jiajun Chen, Yangyang Wu, Xiaoye Miao, Mengying Zhu, Meng Xi

机构 * Center for Data Science, Zhejiang University(数据科学中心,浙江大学) School of Software Technology, Zhejiang University(软件技术学院,浙江大学)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21976 2025-09-25 cs.AI 79%

Compression Strategies for Efficient Multimodal LLMs in Medical Contexts

Tanvir A. Khan, Aranya Saha, Ismam N. Swapnil, Mohammad A. Haque

机构 * Department of Electrical and Electronic Engineering, Bangladesh University of Engineering and Technology (BUET)(电子与电气工程系,孟加拉国工程与技术大学)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.AI

Comments 16 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.07401 2025-09-25 cs.HC cs.AI 79%

Enhancing Higher Education with Generative AI: A Multimodal Approach for Personalised Learning

Johnny Chan, Yuming Li

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.AI

Comments 9 pages, 4 figures, accepted and presented in the 2025 6th International Conference on Advances in Education and Information Technology (AEIT)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19657 2025-09-25 cs.CL cs.AI cs.SI 62%

Large Language Models for Pedestrian Safety: An Application to Predicting Driver Yielding Behavior at Unsignalized Intersections

Yicheng Yang, Zixian Li, Jean Paul Bizimana, Niaz Zafri, Yongfeng Dong, Tianyi Li

专题命中 其他多模态 :multimodal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19954 2025-09-25 cs.RO 50%

Robot Trajectron V2: A Probabilistic Shared Control Framework for Navigation

Pinhao Song, Yurui Du, Ophelie Saussus, Sofie De Schrijver, Irene Caprara, Peter Janssen, Renaud Detry

机构 * KU Leuven(卢森堡大学) University of Washington(华盛顿大学)

专题命中 其他多模态 :multimodal(abstract)

Comments 26 pages, 20 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19652 2025-09-25 stat.AP 50%

Quality-Ensured In-Situ Process Monitoring with Deep Canonical Correlation Analysis

Xiaoyang Song, Wenbo Sun, Metin Kayitmazbatir, Jionghua, Jin

专题命中 其他多模态 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19473 2025-09-25 cs.RO cs.SY eess.SY 50%

Crater Observing Bio-inspired Rolling Articulator (COBRA)

Adarsh Salagame, Henry Noyes, Alireza Ramezani, Eric Sihite, Arash Kalantari

专题命中 其他多模态 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏