arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-10-06 至 2025-10-06 共收录 40 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 6 篇

2510.02922 2025-10-06 cs.CV cs.AI 81%

Multimodal Carotid Risk Stratification with Large Vision-Language Models: Benchmarking, Fine-Tuning, and Clinical Insights

Daphne Tsolissou, Theofanis Ganitidis, Konstantinos Mitsis, Stergios CHristodoulidis, Maria Vakalopoulou, Konstantina Nikita

机构 * National Technical University of Athens(希腊国家技术大学) CentraleSupélec, Université Paris-Saclay(巴黎-萨克雷大学中央理工学院)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02815 2025-10-06 cs.CV 70%

Med-K2N: Flexible K-to-N Modality Translation for Medical Image Synthesis

Feng Yuan, Yifan Gao, Yuehua Ye, Haoyue Li, Xin Gao

机构 * University of Science and Technology of China(中国科学技术大学) Suzhou Institute of Biomedical Engineering and Technology(苏州生物医学工程与技术研究所) Chinese Academy of Sciences(中国科学院) The Third Affiliated Hospital of Sun Yat-sen University(中山大学第三附属医院)

专题命中 图文多模态 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

Comments ICLR2026 under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02780 2025-10-06 cs.CV 57%

Reasoning Riddles: How Explainability Reveals Cognitive Limits in Vision-Language Models

Prahitha Movva

机构 * University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Journal ref COLM 2025: First Workshop on the Application of LLM Explainability to Reasoning and Planning

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.09047 2025-10-06 cs.CL 57%

Same Task, Different Circuits: Disentangling Modality-Specific Mechanisms in VLMs

Yaniv Nikankin, Dana Arad, Yossi Gandelsman, Yonatan Belinkov

机构 * Technion – Israel Institute of Technology(技术学院 – 以色列理工学院) UC Berkeley(加州大学伯克利分校)

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.18035 2025-10-06 cs.CV 57%

Vehicle-Scene Interaction: A Text-Driven 3D Lidar Place Recognition Method for Autonomous Driving

Tianyi Shang, Zhenyu Li, Pengjie Xu, Zhaojun Deng

专题命中 图文多模态 :cross-modal(abstract);分类 cs.CV

Comments 13 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.01534 2025-10-06 cs.CV 57%

Toward a Holistic Evaluation of Robustness in CLIP Models

Weijie Tu, Weijian Deng, Tom Gedeon

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted to IEEE TPAMI, extension of NeurIPS'23 work: A Closer Look at the Robustness of Contrastive Language-Image Pre-Training (CLIP)

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 音频语音多模态 1 篇

2510.02322 2025-10-06 eess.AS cs.CL 81%

SpeechCT-CLIP: Distilling Text-Image Knowledge to Speech for Voice-Native Multimodal CT Analysis

Lukas Buess, Jan Geier, David Bani-Harouni, Chantal Pellegrini, Matthias Keicher, Paula Andrea Perez-Toro, Nassir Navab, Andreas Maier, Tomas Arias-Vergara

机构 * Computer Aided Medical Procedures, Technical University of Munich(慕尼黑技术大学计算机辅助医学程序)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL、eess.AS

Comments Submitted to ICASSP 2026; under review

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频多模态 6 篇

2507.02790 2025-10-06 cs.CV cs.CL 81%

From Long Videos to Engaging Clips: A Human-Inspired Video Editing Framework with Multimodal Narrative Understanding

Xiangfeng Wang, Xiao Li, Yadong Wei, Xueyu Song, Yang Song, Xiaoqiang Xia, Fangrui Zeng, Zaiyi Chen, Liu Liu, Gu Xu, Tong Xu

机构 * University of Science and Technology of China(中国科学技术大学) ByteDance China(字节跳动中国)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments Accepted by EMNLP 2025 Industry Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.03569 2025-10-06 cs.CV cs.AI 81%

Representation Learning for Compressed Video Action Recognition via Attentive Cross-modal Interaction with Motion Enhancement

Bing Li, Jiaxin Chen, Dongming Zhang, Xiuguo Bao, Di Huang

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted to IJCAI 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.07963 2025-10-06 cs.RO cs.AI cs.SY eess.SY 74%

Hierarchical Contact-Rich Trajectory Optimization for Multi-Modal Manipulation using Tight Convex Relaxations

Yuki Shirai, Arvind Raghunathan, Devesh K. Jha

机构 * Mitsubishi Electric Research Laboratories(三菱电机研究实验室)

专题命中 视频多模态 :multi-modal(title);分类 cs.AI

Comments 2025 IEEE International Conference on Robotics and Automation (2025 ICRA)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02778 2025-10-06 cs.CV 57%

AdaRD-key: Adaptive Relevance-Diversity Keyframe Sampling for Long-form Video understanding

Xian Zhang, Zexi Wu, Zinuo Li, Hongming Xu, Luqi Gong, Farid Boussaid, Naoufel Werghi, Mohammed Bennamoun

机构 * The University of Western Australia(西澳大学) Dalian University of Technology(大连理工大学) Khalifa University(卡利夫大学) Zhejiang Lab(浙江实验室)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.04470 2025-10-06 cs.CV 57%

Gate-Shift-Pose: Enhancing Action Recognition in Sports with Skeleton Information

Edoardo Bianchi, Oswald Lanz

机构 * Free University of Bozen-Bolzano(博洛尼亚-博兹纳自由大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted at the 2025 Winter Conference on Applications of Computer Vision (WACV) Workshops. Visit the project page at https://edowhite.github.io/Gate-Shift-Pose

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02702 2025-10-06 cs.CE cs.SI stat.ML 50%

VisitHGNN: Heterogeneous Graph Neural Networks for Modeling Point-of-Interest Visit Patterns

Lin Pang, Jidong J. Yang

专题命中 视频多模态 :multimodal(abstract)

Comments 16 pages, 9 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 跨模态检索 6 篇

2508.00579 2025-10-06 cs.MM cs.IR 79%

MHier-RAG: Multi-Modal RAG for Visual-Rich Document Question-Answering via Hierarchical and Multi-Granularity Reasoning

Ziyu Gong, Chengcheng Mai, Yihua Huang

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.MM

Comments Comments: Update Title, Author, Abstract, etc

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02580 2025-10-06 cs.AI 79%

V2X-UniPool: Unifying Multimodal Perception and Knowledge Reasoning for Autonomous Driving

Xuewen Luo, Fengze Yang, Fan Ding, Xiangbo Gao, Shuo Xing, Yang Zhou, Zhengzhong Tu, Chenxi Liu

机构 * University of Utah(犹他大学) Monash University(莫纳什大学) Texas A&M University(德克萨斯农工大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.06461 2025-10-06 cs.CV 79%

Ranked from Within: Ranking Large Multimodal Models Without Labels

Weijie Tu, Weijian Deng, Dylan Campbell, Yu Yao, Jiyang Zheng, Tom Gedeon, Tongliang Liu

机构 * Australian National University Sydney AI Centre, The University of Sydney Curtin University University of \'OBuda

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

Comments ICML 2025 Camera Ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02790 2025-10-06 cs.CV cs.AI cs.CL cs.MM 70%

MaskCD: Mitigating LVLM Hallucinations by Image Head Masked Contrastive Decoding

Jingyuan Deng, Yujiu Yang

机构 * Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院,清华大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments accepted to emnlp2025 findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02328 2025-10-06 cs.CL cs.AI cs.MA 62%

AMANDA: Agentic Medical Knowledge Augmentation for Data-Efficient Medical Visual Question Answering

Ziqing Wang, Chengsheng Mao, Xiaole Wen, Yuan Luo, Kaize Ding

机构 * Northwestern University(西北大学) Microsoft(微软公司)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL、cs.AI

Comments EMNLP Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03200 2025-10-06 cs.CV 57%

MonSTeR: a Unified Model for Motion, Scene, Text Retrieval

Luca Collorone, Matteo Gioia, Massimiliano Pappa, Paolo Leoni, Giovanni Ficarra, Or Litany, Indro Spinelli, Fabio Galasso

机构 * Sapienza University of Rome(罗马萨皮恩扎大学) Technion, NVIDIA(技术学院与NVIDIA) WSense

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 多模态生成 7 篇

2504.20629 2025-10-06 cs.CV cs.AI cs.MM 82%

AlignDiT: Multimodal Aligned Diffusion Transformer for Synchronized Speech Generation

Jeongsoo Choi, Ji-Hoon Kim, Kim Sung-Bin, Tae-Hyun Oh, Joon Son Chung

机构 * Korea Advanced Institute of Science and Technology(韩国科学技术院) Pohang University of Science and Technology(釜山科学技术大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI、cs.MM

Comments ACM Multimedia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02403 2025-10-06 q-bio.QM cs.AI cs.CV 81%

Glaucoma Detection and Structured OCT Report Generation via a Fine-tuned Multimodal Large Language Model

Jalil Jalili, Yashraj Gavhane, Evan Walker, Anna Heinke, Christopher Bowd, Akram Belghith, Massimo A. Fazio, Christopher A. Girkin, C. Gustavo De Moraes, Jeffrey M. Liebmann, Sally L. Baxter, Robert N. Weinreb, Linda M. Zangwill, Mark Christopher

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17121 2025-10-06 cs.CL cs.AI 81%

NeSyGeo: A Neuro-Symbolic Framework for Multimodal Geometric Reasoning Data Generation

Weiming Wu, Jin Ye, Zi-kang Wang, Zhi Zhou, Yu-Feng Li, Lan-Zhe Guo

机构 * School of Intelligence Science and Technology, Nanjing University(智能科学与技术学院,南京大学) National Key Laboratory for Novel Software Technology, Nanjing University(新型软件技术国家实验室,南京大学) School of Artificial Intelligence, Nanjing University(人工智能学院,南京大学)

专题命中 多模态生成 :multimodal(title);multi-modal(abstract);分类 cs.CL、cs.AI

Comments 29 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02880 2025-10-06 cs.AI 79%

Consolidating Reinforcement Learning for Multimodal Discrete Diffusion Models

Tianren Ma, Mu Zhang, Yibing Wang, Qixiang Ye

机构 * University of Chinese Academy of Sciences(中国科学院大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.AI

Comments Project Page: https://github.com/martian422/MaskGRPO

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02987 2025-10-06 cs.CV 57%

TIT-Score: Evaluating Long-Prompt Based Text-to-Image Alignment via Text-to-Image-to-Text Consistency

Juntong Wang, Huiyu Duan, Jiarui Wang, Ziheng Jia, Guangtao Zhai, Xiongkuo Min

机构 * Institute of Image Communication and Network Engineering(图像通信与网络工程研究所) MoE Key Lab of Artificial Intelligence, AI Institute(人工智能关键实验室) Shanghai Jiao Tong University(上海交通大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02722 2025-10-06 cs.CV 57%

MoGIC: Boosting Motion Generation via Intention Understanding and Visual Context

Junyu Shi, Yong Sun, Zhiyuan Zhang, Lijiang Liu, Zhengjie Zhang, Yuxin He, Qiang Nie

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.18406 2025-10-06 cs.CV cs.AI cs.CL 56%

RACCooN: A Versatile Instructional Video Editing Framework with Auto-Generated Narratives

Jaehong Yoon, Shoubin Yu, Mohit Bansal

机构 * UNC Chapel Hill(北卡罗来纳大学教堂山分校) Nanyang Technological University(南洋理工大学)

专题命中 多模态生成 :分类 cs.CV、cs.CL、cs.AI;MLLM(comments)

Comments EMNLP 2025 main; The first two authors contribute equally. Project Page: https://raccoon-mllm-gen.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 多模态评测 4 篇

2510.02876 2025-10-06 cs.CV cs.LG 79%

ELMF4EggQ: Ensemble Learning with Multimodal Feature Fusion for Non-Destructive Egg Quality Assessment

Md Zahim Hassan, Md. Osama, Muhammad Ashad Kabir, Md. Saiful Islam, Zannatul Naim

机构 * organization= Department of Computer Science Engineering, Bangladesh Army University of Science Engineering, Rajshahi University of Engineering \& Technology , state= Rajshahi , postcode= 6204 , country= Bangladesh organization= School of Computing, Mathematics Engineering, Charles Sturt University , city= Bathurst , state= NSW , postcode= 2795 , country= Australia organization= Gulbali Institute for Agriculture, Water Environment, Charles Sturt University , city= Wagga Wagga , state= NSW , postcode= 2678 , country= Australia organization= Department of Animal Production Management, Sher-e-Bangla Agricultural University , city= Dhaka , country= Bangladesh

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Comments 30 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02726 2025-10-06 cs.CL 79%

PGMEL: Policy Gradient-based Generative Adversarial Network for Multimodal Entity Linking

KM Pooja, Cheng Long, Aixin Sun

机构 * Department of Information Technology, Indian Institute of Information Technology, Allahabad India(信息科技系,印度信息科技学院,阿勒颇印度) School of Computer Science and Engineering, Nanyang Technological University(计算机科学与工程学院,南洋理工大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02423 2025-10-06 cs.AI 57%

RefineShot: Rethinking Cinematography Understanding with Foundational Skill Evaluation

Hang Wu, Yujun Cai, Haonan Ge, Hongkai Chen, Ming-Hsuan Yang, Yiwei Wang

机构 * University of California, Merced(加州大学梅尔德分校) The University of Queensland(昆士兰大学) vivo Mobile Communication Co., Ltd(vivo移动通信有限公司)

专题命中 多模态评测 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02870 2025-10-06 math.OC 50%

Wasserstein crossover for evolutionary algorithm-based topology optimization

Taisei Kii, Kentaro Yaji, Hiroshi Teramoto, Kikuo Fujita

专题命中 多模态评测 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏