arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-09-10 至 2025-09-10 共收录 41 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 2 篇

2509.07213 2025-09-10 cs.CV cs.AI 81%

XBusNet: Text-Guided Breast Ultrasound Segmentation via Multimodal Vision-Language Learning

Raja Mallina, Bryar Shareef

机构 * Department of Computer Science University of Nevada, Las Vegas(计算机科学系 纳瓦拉大学拉斯维加斯分校)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 15 pages, 3 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04378 2025-09-10 cs.CV 70%

Aesthetic Image Captioning with Saliency Enhanced MLLMs

Yilin Tao, Jiashui Huang, Huaze Xu, Ling Shao

专题命中 图文多模态 :multimodal(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 音频语音多模态 1 篇

2505.20511 2025-09-10 cs.CL 79%

Multimodal Emotion Recognition in Conversations: A Survey of Methods, Trends, Challenges and Prospects

Chengyan Wu, Yiqiang Cai, Yang Liu, Pengxu Zhu, Yun Xue, Ziwei Gong, Julia Hirschberg, Bolei Ma

机构 * Guangdong Provincial Key Laboratory of Quantum Engineering and Quantum Materials(广东省量子工程与量子材料重点实验室) School of Electronic Science and Engineering (School of Microelectronics), South China Normal University(华南师范大学电子科学学院(微电子学院)) North Carolina Central University(北卡罗来纳中央大学) Georgia Institute of Technology(佐治亚理工学院) Columbia University(哥伦比亚大学) LMU Munich & Munich Center for Machine Learning(慕尼黑大学及慕尼黑机器学习中心)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL

Comments EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频多模态 3 篇

2506.18407 2025-09-10 cs.GR cs.CV 83%

IntuiTF: MLLM-Guided Transfer Function Optimization for Direct Volume Rendering

Yiyao Wang, Bo Pan, Ke Wang, Han Liu, Jinyuan Mao, Yuxin Liu, Minfeng Zhu, Xiuqi Huang, Weifeng Chen, Bo Zhang, Wei Chen

机构 * State Key Lab of CAD&CG, Zhejiang University(浙江大学CAD与CG国家重点实验室) Laboratory of Art and Archaeology Image (Zhejiang University), Ministry of Education, China(浙江大学艺术与考古图像实验室(教育部,中国)) Zhejiang University of Finance&Economics(浙江财经大学) Zhejiang University(浙江大学)

专题命中 视频多模态 :MLLM(title,abstract);multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.02074 2025-09-10 cs.CV cs.AI 62%

Large Language Models for Crash Detection in Video: A Survey of Methods, Datasets, and Challenges

Sanjeda Akter, Ibne Farabi Shihab, Anuj Sharma

机构 * Department of Computer Science, Iowa State University(计算机科学系,爱荷华州立大学) Department of Civil, Construction and Environmental Engineering, Iowa State University(土木、建设与环境工程系,爱荷华州立大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.16013 2025-09-10 cs.RO cs.CV 57%

GraspCoT: Integrating Physical Property Reasoning for 6-DoF Grasping under Flexible Language Instructions

Xiaomeng Chu, Jiajun Deng, Guoliang You, Wei Liu, Xingchen Li, Jianmin Ji, Yanyong Zhang

机构 * University of Science and Technology of China(中国科学技术大学) The University of Adelaide(阿德莱德大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted to ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 跨模态检索 2 篇

2509.07666 2025-09-10 cs.CL cs.IR 79%

MoLoRAG: Bootstrapping Document Understanding via Multi-modal Logic-aware Retrieval

Xixi Wu, Yanchao Tan, Nan Hou, Ruiyang Zhang, Hong Cheng

机构 * The Chinese University of Hong Kong(香港中文大学) Fuzhou University(福州大学) University of Macau(澳门大学)

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CL

Comments EMNLP Main 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10559 2025-09-10 cs.CV cs.AI cs.ET 62%

From Images to Insights: Explainable Biodiversity Monitoring with Plain Language Habitat Explanations

Yutong Zhou, Masahiro Ryo

机构 * Leibniz Centre for Agricultural Landscape Research (ZALF)(莱比锡农业景观研究中心(ZALF)) Brandenburg University of Technology Cottbus–Senftenberg(勃兰登堡技术大学科特博斯-森芬根堡)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.AI

Comments AISE workshop camera-ready version @ ECAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 多模态生成 4 篇

2509.07817 2025-09-10 cs.CL cs.MM 81%

Dual Knowledge-Enhanced Two-Stage Reasoner for Multimodal Dialog Systems

Xiaolin Chen, Xuemeng Song, Haokun Wen, Weili Guan, Xiangyu Zhao, Liqiang Nie

机构 * National University of Singapore(新加坡国立大学) Southern University of Science and Technology(南方科技大学) Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳)) City University of Hong Kong(香港城市大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07473 2025-09-10 cs.AI 79%

SheetDesigner: MLLM-Powered Spreadsheet Layout Generation with Rule-Based and Vision-Based Reflection

Qin Chen, Yuanyi Ren, Xiaojun Ma, Mugeng Liu, Han Shi, Dongmei Zhang

机构 * Peking University(北京大学) Microsoft(微软)

专题命中 多模态生成 :MLLM(title);multimodal(abstract);分类 cs.AI

Comments Accepted to EMNLP 2025 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06945 2025-09-10 cs.CV cs.AI cs.CL cs.LG 67%

Interleaving Reasoning for Better Text-to-Image Generation

Wenxuan Huang, Shuang Chen, Zheyong Xie, Shaosheng Cao, Shixiang Tang, Yufan Shen, Qingyu Yin, Wenbo Hu, Xiaoman Wang, Yuntian Tang, Junbo Qiao, Yue Guo, Yao Hu, Zhenfei Yin, Philip Torr, Yu Cheng, Wanli Ouyang, Shaohui Lin

机构 * East China Normal University(华东师范大学) The Chinese University of Hong Kong(香港中文大学) Xiaohongshu Inc.(小红书公司) University of California, Los Angeles(加州大学洛杉矶分校) Zhejiang University(浙江大学) University of Oxford(牛津大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04444 2025-09-10 cs.CV 57%

One Flight Over the Gap: A Survey from Perspective to Panoramic Vision

Xin Lin, Xian Ge, Dizhe Zhang, Zhaoliang Wan, Xianshun Wang, Xiangtai Li, Wenjie Jiang, Bo Du, Dacheng Tao, Ming-Hsuan Yang, Lu Qi

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments Project Page: https://insta360-research-team.github.io/Survey-of-Panorama/

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 多模态评测 14 篇

2508.20670 2025-09-10 cs.CV cs.MM 84%

"Humor, Art, or Misinformation?": A Multimodal Dataset for Intent-Aware Synthetic Image Detection

Anastasios Skoularikis, Stefanos-Iordanis Papadopoulos, Symeon Papadopoulos, Panagiotis C. Petrantonakis

机构 * Department of Electrical & Computer Engineering, Aristotle University of Thessaloniki(电气与计算机工程系,阿基米德大学塞萨洛尼基分校) Information Technology Institute, Centre for Research & Technology Hellas(信息科技研究所,希腊研究中心)

专题命中 多模态评测 :multimodal(title,abstract);image-text(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.13923 2025-09-10 cs.AI cs.CL cs.CV cs.MM 83%

PIN: A Knowledge-Intensive Dataset for Paired and Interleaved Multimodal Documents

Junjie Wang, Yuxiang Zhang, Minghao Liu, Yin Zhang, Yatai Ji, Weihao Xuan, Nie Lin, Kang Zhu, Zhiqiang Lin, Yiming Ren, Chunyang Jiang, Yiyao Yu, Zekun Wang, Tiezhen Wang, Wenhao Huang, Jie Fu, Qunshu Lin, Yujiu Yang, Ge Zhang, Ruibin Yuan, Bei Chen, Wenhu Chen

机构 * M-A-P

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Technical report v1.0

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00785 2025-09-10 cs.AI cs.CV cs.LG 81%

GeoChain: Multimodal Chain-of-Thought for Geographic Reasoning

Sahiti Yerramilli, Nilay Pande, Rynaa Grover, Jayant Sravan Tamarapalli

机构 * Google(谷歌) Waymo

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07050 2025-09-10 cs.CV cs.AI cs.CY 81%

Automated Evaluation of Gender Bias Across 13 Large Multimodal Models

Juan Manuel Contreras

机构 * Aymara

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06011 2025-09-10 cs.CV 79%

Light-Weight Cross-Modal Enhancement Method with Benchmark Construction for UAV-based Open-Vocabulary Object Detection

Zhenhai Weng, Xinjie Li, Can Wu, Weijie He, Jianfeng Lv, Dong Zhou, Zhongliang Yu

机构 * School of Automation, Chongqing University(重庆大学自动化学院) School of Information Science and Engineering, Lanzhou University(兰州大学信息科学与工程学院) Department of Control Science and Engineering, Harbin Institute of Technology(哈尔滨工业大学控制科学与工程系) Department of Mechanical and Automation Engineering, The Chinese University of Hong Kong(香港中文大学机械与自动化工程系)

专题命中 多模态评测 :cross-modal(title);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17471 2025-09-10 cs.CL 79%

FinRAGBench-V: A Benchmark for Multimodal RAG with Visual Citation in the Financial Domain

Suifeng Zhao, Zhuoran Jin, Sujian Li, Jun Gao

机构 * Key Laboratory of High Confidence Software Technologies, CS, Peking University, China(北京大学高可信软件技术重点实验室) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) State Key Laboratory of Multimedia Information Processing, School of Computer Sciences, Peking University(北京大学多媒体信息处理国家重点实验室)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07362 2025-09-10 cs.RO 71%

Aerial-ground Cross-modal Localization: Dataset, Ground-truth, and Benchmark

Yandi Yang, Jianping Li, Youqi Liao, Yuhao Li, Yizhe Zhang, Zhen Dong, Bisheng Yang, Naser El-Sheimy

机构 * Department of Geomatics Engineering(测绘工程系) School of Electrical and Electronic Engineering(电子电气工程学院) State Key Laboratory of Information Engineering in Surveying, Mapping and Remote Sensing(测绘遥感信息工程国家重点实验室) School of Mechanical Engineering(机械工程学院)

专题命中 多模态评测 :cross-modal(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07969 2025-09-10 cs.CV cs.AI cs.CL 67%

Mini-o3: Scaling Up Reasoning Patterns and Interaction Turns for Visual Search

Xin Lai, Junyi Li, Wei Li, Tao Liu, Tianjian Li, Hengshuang Zhao

机构 * The University of Hong Kong(香港大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Code, datasets, models are available at https://github.com/Mini-o3/Mini-o3. Project Page: https://mini-o3.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07772 2025-09-10 cs.CV cs.AI 62%

XSRD-Net: EXplainable Stroke Relapse Detection

Christian Gapp, Elias Tappeiner, Martin Welk, Karl Fritscher, Stephanie Mangesius, Constantin Eisenschink, Philipp Deisl, Michael Knoflach, Astrid E. Grams, Elke R. Gizewski, Rainer Schubert

机构 * Institute of Biomedical Image Analysis UMIT TIROL -- Private University for Health Sciences(生物医学影像分析研究所)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.AI

Comments Contribution to MICAD 2025 conference, Nov. 19-21, 2025 | London, UK

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17180 2025-09-10 cs.AI cs.CV cs.LG 62%

MaRVL-QA: A Benchmark for Mathematical Reasoning over Visual Landscapes

Nilay Pande, Sahiti Yerramilli, Jayant Sravan Tamarapalli, Rynaa Grover

机构 * Waymo Google(谷歌)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06585 2025-09-10 cs.AI cs.CV 62%

CountQA: How Well Do MLLMs Count in the Wild?

Jayant Sravan Tamarapalli, Rynaa Grover, Nilay Pande, Sahiti Yerramilli

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.18046 2025-09-10 cs.CV cs.AI 62%

DMS-Net:Dual-Modal Multi-Scale Siamese Network for Binocular Fundus Image Classification

Guohao Huo, Zibo Lin, Zitong Wang, Ruiting Dai, Hao Tang

机构 * University of Electronic Science and Technology of China(电子科技大学) School of Computer Science, Peking University(北京大学计算机学院)

专题命中 多模态评测 :cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07010 2025-09-10 cs.CV cs.AI cs.ET 62%

Human-in-the-Loop: Quantitative Evaluation of 3D Models Generation by Large Language Models

Ahmed R. Sadik, Mariusz Bujny

机构 * Honda Research Institute Europe - Germany(本田欧洲研究机构)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07504 2025-09-10 cs.CR 50%

Backdoor Attacks and Defenses in Computer Vision Domain: A Survey

Bilal Hussain Abbasi, Yanjun Zhang, Leo Zhang, Shang Gao

专题命中 多模态评测 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

7. 多模态训练与对齐 10 篇

2407.20454 2025-09-10 cs.LG cs.CL 83%

CoMMIT: Coordinated Multimodal Instruction Tuning

Xintong Li, Junda Wu, Tong Yu, Yu Wang, Xiang Chen, Jiuxiang Gu, Lina Yao, Julian McAuley, Jingbo Shang

机构 * University of California, San Diego(加州大学圣地亚哥分校) Adobe Research(Adobe研究) The University of New South Wales(新南威尔士大学) CSIRO’s Data61(澳大利亚联邦科学与工业研究组织的数据61)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CL

Comments 9 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06976 2025-09-10 cs.LG cs.AI 83%

A Knowledge-Guided Cross-Modal Feature Fusion Model for Local Traffic Demand Prediction

Lingyu Zhang, Pengfei Xu, Guobin Wu, Jian Liang, Ruiyang Dong, Yunhai Wang, Xuan Song

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07923 2025-09-10 cs.CV cs.AI 81%

Multimodal Contrastive Pretraining of CBCT and IOS for Enhanced Tooth Segmentation

Moo Hyun Son, Juyoung Bae, Zelin Qiu, Jiale Peng, Kai Xin Li, Yifan Lin, Hao Chen

机构 * Department of Computer Science and Engineering(计算机科学与工程系) The Hong Kong University of Science and Technology(香港科学与技术大学) Division of Paediatric Dentistry and Orthodontics, Faculty of Dentistry(牙科学院儿童牙科与正畸科) The University of Hong Kong(香港大学) Delun Dental Hospital(德尔伦牙科医院) Department of Chemical and Biological Engineering(化学与生物工程系) Division of Life Science(生命科学系) HKUST Shenzhen-Hong Kong Collaborative Innovation Research Institute(香港科技大学深圳-香港协同创新研究院) State Key Laboratory of Nervous System Disorders(神经系统紊乱国家重点实验室)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06987 2025-09-10 cs.CV cs.AI 81%

FusWay: Multimodal hybrid fusion approach. Application to Railway Defect Detection

Alexey Zhukov, Jenny Benois-Pineau, Amira Youssef, Akka Zemmari, Mohamed Mosbah, Virginie Taillandier

机构 * University Bordeaux, CNRS, Bordeaux INP, INRIA, LaBRI(波尔多大学、国家科学研究中心、波尔多国立理工学院、INRIA、LaBRI) SNCF, DIR TECHNOLOGIES INNOVATION ET PROJETS GROUPE, IR - DPISF TECH4RAIL - TLI(法国国家铁路公司、技术与项目集团、IR - DPISF TECH4RAIL - TLI)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏