arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-12-18 至 2025-12-18 共收录 38 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 6 篇

2512.14989 2025-12-18 cs.CL cs.AI cs.CV 82%

Evaluating Large Language Models on Multimodal Chemistry Olympiad Exams

评估大型语言模型在多模态化学奥林匹克考试中的表现

Yiming Cui, Xin Yao, Yuxuan Qin, Xin Li, Shijin Wang, Guoping Hu

机构 * State Key Laboratory of Cognitive Intelligence(认知智能国家重点实验室) iFLYTEK AI Research(iFLYTEK人工智能研究院)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 本文评估了多种多模态LLM在化学奥林匹克考试中的表现,发现其在多模态融合和科学推理方面存在显著局限,提出通过链式思维提示提升模型性能的方法。

Comments Published at Communications Chemistry

Journal ref Commun. Chem. 8 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.15611 2025-12-18 cs.CV 79%

If you can describe it, they can see it: Cross-Modal Learning of Visual Concepts from Textual Descriptions

如果你能描述它,他们就能看到它:从文本描述中跨模态学习视觉概念

Carlo Alberto Barbano, Luca Molinaro, Massimiliano Ciranni, Emanuele Aiello, Vito Paolo Pastore, Marco Grangetto

机构 * University of Turin(都灵大学) University of Genoa(热那亚大学) Politecnico di Torino(都灵理工大学)

专题命中 图文多模态 :cross-modal(title);image-text(abstract);分类 cs.CV

AI总结 本文提出通过文本描述跨模态学习视觉概念的方法,利用知识迁移技术提升视觉-语言模型的零样本性能。

Comments 27 pages. Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.09172 2025-12-18 cs.CV cs.AI 62%

Prompt-Based Continual Compositional Zero-Shot Learning

基于提示的持续组成零样本学习

Sauda Maryam, Sara Nadeem, Faisal Qureshi, Mohsen Ali

机构 * Intelligent Machines Lab(智能机器实验室) Information Technology University(信息技术大学) Visual Computing Lab(视觉计算实验室) Ontario Tech University(安大略技术大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 PromptCCZSL通过多教师蒸馏和正交投影损失,实现持续组成零样本学习中的知识保留与泛化提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.15531 2025-12-18 cs.CV 57%

An Efficient and Effective Encoder Model for Vision and Language Tasks in the Remote Sensing Domain

一种用于遥感领域的高效且有效的编码器模型

João Daniel Silva, Joao Magalhaes, Devis Tuia, Bruno Martins

机构 * INESC-ID, Instituto Superior Técnico, University of Lisbon(INESC-ID,理工学院,里斯本大学) Department of Computer Science, Faculty of Science and Technology, Universidade NOVA de Lisboa(计算机科学系,科学与技术学院,诺瓦大学) School of Architecture, Civil and Environmental Engineering, EPFL(建筑、土木和环境工程学院,EPFL)

专题命中 图文多模态 :cross-modal(abstract);分类 cs.CV

AI总结 本文提出了一种高效且紧凑的编码器模型,用于处理遥感领域的多任务学习,包括图像文本生成和跨模态检索。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.15433 2025-12-18 cs.CV 57%

CLIP-FTI: Fine-Grained Face Template Inversion via CLIP-Driven Attribute Conditioning

CLIP-FTI: 通过CLIP驱动的属性条件化实现细粒度面部模板逆向

Longchen Dai, Zixuan Shen, Zhiheng Zhou, Peipeng Yu, Zhihua Xia

专题命中 图文多模态 :cross-modal(abstract);分类 cs.CV

AI总结 CLIP-FTI通过CLIP驱动的属性条件化实现细粒度面部模板逆向,提升识别准确性和属性相似性,增强跨模型攻击的可转移性。

Comments Accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.15254 2025-12-18 cs.CV cs.LG 57%

Assessing the Visual Enumeration Abilities of Specialized Counting Architectures and Vision-Language Models

评估专用计数架构和视觉-语言模型的视觉计数能力

Kuinan Hou, Jing Mi, Marco Zorzi, Lamberto Ballan, Alberto Testolin

机构 * University of Padova(帕多瓦大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

AI总结 本文比较了专用计数架构与VLMs在物体计数任务中的性能,发现VLMs在生成中间表示时计数准确率显著提高,但复杂场景计数仍需进一步研究。

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 音频语音多模态 2 篇

2512.15262 2025-12-18 eess.IV cs.MM 89%

Audio-Visual Cross-Modal Compression for Generative Face Video Coding

音频-视觉跨模态压缩用于生成式面部视频编码

Youmin Xu, Mengxi Guo, Shijie Zhao, Weiqi Li, Junlin Li, Li Zhang, Jian Zhang

专题命中 音频语音多模态 :cross-modal(title,abstract);audio-visual(title,abstract);multimodal(abstract);分类 cs.MM

AI总结 本文提出AVCC框架,通过联合压缩音频和视频流,提升生成式面部视频编码的率-失真性能。

Comments Accepted as a PAPER and for publication in the DCC 2026 proceedings

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26569 2025-12-18 cs.CV cs.IR cs.MM 81%

AdSum: Two-stream Audio-visual Summarization for Automated Video Advertisement Clipping

AdSum:用于自动化视频广告剪辑的双流音频视觉摘要

Wen Xie, Yanjun Zhu, Gijs Overgoor, Yakov Bart, Agata Lapedriza Garcia, Sarah Ostadabbas

机构 * Northeastern University(东北大学) Southern Methodist University(南方 Methodist 大学)

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV、cs.MM

AI总结 AdSum提出了一种双流音频视觉融合模型,用于自动化视频广告剪辑,通过重要性预测提升广告剪辑效率。

Comments Accepted at 32nd International Conference on MultiMedia Modeling

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频多模态 1 篇

2512.15681 2025-12-18 eess.IV 50%

Radiomics and Clinical Features in Predictive Modelling of Brain Metastases Recurrence

放射组学与临床特征在脑转移瘤复发预测模型中的应用

Ines Faria, Matheus Silva, Crystian Saraiva, Jose Soares, Victor Alves

专题命中 视频多模态 :multimodal(abstract)

AI总结 本研究利用放射组学和临床数据构建模型,预测脑转移瘤复发并探讨辐射剂量差异与复发风险的关系。

Comments 14 pages, 6 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 跨模态检索 3 篇

2508.20209 2025-12-18 physics.med-ph 78%

Low-exposure, high-quality multimodal speckle X-ray imaging via an intrinsic gradient-flow approach

低曝光高质量多模态散射X射线成像:基于内在梯度流方法

Jayvan Liu, Samantha J. Alloo, Max Langer, Konstantin M. Pavlov

专题命中 跨模态检索 :multimodal(title,abstract)

AI总结 本研究提出了一种基于内在梯度流方法的改进算法,用于高效获取高质量的多模态散射X射线成像数据。

Comments 10 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16674 2025-12-18 cs.CV 57%

FitPro: A Zero-Shot Framework for Interactive Text-based Pedestrian Retrieval in Open World

FitPro: 一个用于开放世界中基于文本的交互行人检索的零样本框架

Zengli Luo, Canlong Zhang, Xiaochun Lu, Zhixin Li

机构 * Key Lab of Education Blockchain and Intelligent Technology, Ministry of Education, Guangxi Normal University(教育区块链与智能技术重点实验室,教育部,广西师范大学) Guangxi Key Lab of Multi-source Information Mining & Security, Guangxi Normal University(广西多源信息挖掘与安全重点实验室,广西师范大学)

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV

AI总结 FitPro通过引入特征对比解码、增量语义挖掘和查询感知分层检索,提升开放世界中基于文本的交互行人检索的语义理解和跨场景适应能力。

Comments 12pages,6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.14878 2025-12-18 cs.CV 57%

Visual-textual Dermatoglyphic Animal Biometrics: A First Case Study on Panthera tigris

视觉-文本皮肤纹路动物生物特征:对豹属(Panthera tigris)的首次案例研究

Wenshuo Li, Majid Mirmehdi, Tilo Burghardt

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

AI总结 本研究通过引入皮肤纹路文本描述符,提出了一种视觉-文本结合的动物Re-ID方法,提升跨模态身份检索的准确性并缓解数据稀缺问题。

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 多模态生成 5 篇

2512.15261 2025-12-18 cs.CV 83%

MMMamba: A Versatile Cross-Modal In Context Fusion Framework for Pan-Sharpening and Zero-Shot Image Enhancement

MMMamba: 一种多功能跨模态在上下文融合框架用于全色锐化和零样本图像增强

Yingying Wang, Xuanhua He, Chen Wu, Jialing Huang, Suiyun Zhang, Rui Liu, Xinghao Ding, Haoxuan Che

专题命中 多模态生成 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV

AI总结 MMMamba通过跨模态在上下文融合框架实现全色锐化和零样本图像增强,结合Mamba架构和多模态交错扫描机制,提升跨模态交互能力与计算效率。

Comments \link{Code}{https://github.com/Gracewangyy/MMMamba}

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.15069 2025-12-18 cs.CV cs.AI 81%

PMMD: A pose-guided multi-view multi-modal diffusion for person generation

PMMD: 一种基于姿态的多视角多模态扩散用于人物生成

Ziyu Shang, Haoran Liu, Rongchao Zhang, Zhiqian Wei, Tongtong Feng

机构 * Harbin Institute of Technology, Shenzhen, China(哈尔滨工业大学(深圳)) City University of Hong Kong, Hong Kong, China(香港城市大学) Peking University, Beijing, China(北京大学) Tsinghua University, Beijing, China(清华大学)

专题命中 多模态生成 :multi-modal(title);multimodal(abstract);分类 cs.CV、cs.AI

AI总结 PMMD通过多视角多模态扩散框架,实现可控姿态和外观的人像生成,提升一致性、细节保留和可控性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.14938 2025-12-18 cs.CV cs.AI cs.MM cs.SD 75%

TalkVerse: Democratizing Minute-Long Audio-Driven Video Generation

TalkVerse:民主化分钟级音频驱动视频生成

Zhenzhi Wang, Jian Wang, Ke Ma, Dahua Lin, Bing Zhou

机构 * The Chinese University of Hong Kong(香港中文大学) Snap Inc.(Snap公司)

专题命中 多模态生成 :MLLM(abstract);audio-visual(abstract);分类 cs.CV、cs.AI、cs.MM

AI总结 TalkVerse通过大规模开放数据和高效模型,实现了分钟级音频驱动视频生成,降低研究门槛,提升生成质量与效率。

Comments open-sourced single-person full-body talking video generation dataset, training code and checkpoints

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.14855 2025-12-18 cs.CE cs.AI 57%

A Roadmap for Applying Graph Neural Networks to Numerical Data: Insights from Cementitious Materials

将图神经网络应用于数值数据的路线图:从水泥材料中获得的见解

Mahmuda Sharmin, Taihao Han, Jie Huang, Narayanan Neithalath, Gaurav Sant, Aditya Kumar

专题命中 多模态生成 :multi-modal(abstract);分类 cs.AI

AI总结 本文提出利用图神经网络(GNN)设计混凝土,通过K-NN方法将表格数据转化为图表示,并优化参数以实现可解释的预测。

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.05005 2025-12-18 cs.CV cs.LG cs.RO 57%

Diff-2-in-1: Bridging Generation and Dense Perception with Diffusion Models

Diff-2-in-1:通过扩散模型弥合生成与密集感知的鸿沟

Shuhong Zheng, Zhipeng Bao, Ruoyu Zhao, Martial Hebert, Yu-Xiong Wang

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Carnegie Mellon University(卡内基梅隆大学) Tsinghua University(清华大学)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

AI总结 Diff-2-in-1通过结合扩散模型的生成与感知能力,实现多模态数据生成与密集视觉感知的统一框架,提升视觉感知的判别能力。

Comments 26 pages, 14 figures

Journal ref ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 多模态评测 10 篇

2510.01899 2025-12-18 cs.LG cs.AI cs.HC 89%

Multimodal Foundation Models for Early Disease Detection

多模态基础模型用于早期疾病检测

Md Talha Mohsin, Ismail Abdulrashid

机构 * The University of Tulsa(塔尔萨大学)

专题命中 多模态评测 :multimodal(title,abstract);multimodal foundation model(title,abstract);cross-modal(abstract);分类 cs.AI

AI总结 本文提出一种多模态基础模型,通过整合电子病历、医学影像、基因组学和可穿戴传感器数据,提升早期疾病检测的准确性和鲁棒性。

Comments 6 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.15171 2025-12-18 cs.CV 85%

Cross-modal ultra-scale learning with tri-modalities of renal biopsy images for glomerular multi-disease auxiliary diagnosis

跨模态超尺度学习:利用三种模态的肾活检图像进行肾小球多病种辅助诊断

Kaixing Long, Danyi Weng, Yun Mi, Zhentai Zhang, Yanmeng Lu, Jian Geng, Zhitao Zhou, Liming Zhong, Qianjin Feng, Wei Yang, Lei Cao

机构 * School of Biomedical Engineering, Southern Medical University(生物医学工程学院,南方医科大学) Guangdong Provincial Key Laboratory of Medical Image Processing(广东省医学影像处理重点实验室) Guangdong Province Engineering Laboratory for Medical Imaging and Diagnostic Technology(广东省医学影像与诊断技术工程实验室) Central Laboratory, Southern Medical University(中心实验室,南方医科大学) Department of Pathology, School of Basic Medical Sciences, Southern Medical University(病理学部,基础医学学院,南方医科大学) Guangzhou Huayin Medical Laboratory Center(广州华英医学实验室中心)

专题命中 多模态评测 :cross-modal(title,abstract);multimodal(abstract);multi-modal(abstract);分类 cs.CV

AI总结 本文提出CMUS-Net,通过跨模态超尺度学习,利用三种模态的肾活检图像实现肾小球多病种的自动分类,提升了诊断准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.09603 2025-12-18 cs.CV 83%

Do MLLMs Exhibit Human-like Perceptual Behaviors? HVSBench: A Benchmark for MLLM Alignment with Human Perceptual Behavior

多模态大语言模型是否表现出类人感知行为?HVSBench:一个用于多模态大语言模型对齐人类视觉行为的基准测试

Jiaying Lin, Shuquan Ye, Dan Xu, Wanli Ouyang, Rynson W. H. Lau

专题命中 多模态评测 :MLLM(title,abstract);multimodal(abstract);分类 cs.CV

AI总结 HVSBench通过测试多模态大语言模型与人类视觉行为的对齐性,揭示了其在感知任务上的显著差距,并强调了开发更符合人类认知的AI的重要性。

Comments Project page: https://jiaying.link/HVSBench/

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.14691 2025-12-18 cs.CL cs.CV 81%

MMGR: Multi-Modal Generative Reasoning

MMGR:多模态生成推理

Zefan Cai, Haoyi Qiu, Tianyi Ma, Haozhe Zhao, Gengze Zhou, Kung-Hsiang Huang, Parisa Kordjamshidi, Minjia Zhang, Wen Xiao, Jiuxiang Gu, Nanyun Peng, Junjie Hu

机构 * University of Wisconsin–Madison(威斯康星大学麦迪逊分校) University of California, Los Angeles(加州大学洛杉矶分校) Michigan State University(密歇根州立大学) University of Illinois Urbana–Champaign(伊利诺伊大学厄巴纳-香槟分校) University of Adelaide(阿德莱德大学) Salesforce AI Research(Salesforce人工智能研究) Microsoft(微软) Adobe Research(Adobe研究)

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV、cs.CL

AI总结 MMGR提出一个多模态生成推理评估框架,通过物理、逻辑、空间和时间五个能力评估视频和图像生成模型,揭示其在抽象推理和具身导航中的性能不足。

Comments work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01140 2025-12-18 stat.ML cs.AI cs.CV cs.LG eess.IV 81%

Few-Shot Multimodal Medical Imaging: A Theoretical Framework

少样本多模态医学影像:一个理论框架

Md Talha Mohsin, Ismail Abdulrashid

机构 * The University of Tulsa 800 S Tucker Dr, Tulsa, OK 74104, USA(塔尔萨大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 本文提出少样本多模态医学影像的理论框架,通过样本复杂性、不确定性量化和可解释性分析,为低资源环境下的高效、可靠诊断模型设计提供理论基础。

Comments 6 Pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.03197 2025-12-18 cs.CL 79%

Explain with Visual Keypoints Like a Real Mentor! A Benchmark for Multimodal Solution Explanation

像真实导师一样用视觉关键点解释吧!一个多模态解决方案解释基准

Jaewoo Park, Jungyang Park, Dongju Jang, Jiwan Chung, Byungwoo Yoo, Jaewoo Shin, Seonjoon Park, Taehyeong Kim, Youngjae Yu

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL

AI总结 本文提出多模态解决方案解释任务和ME2数据集,旨在评估模型识别视觉关键点并生成相关解释的能力,揭示当前LLM在数学视觉推理和教育应用中的不足。

Comments 14 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.14750 2025-12-18 q-bio.QM cs.AI 79%

Multiscale Cross-Modal Mapping of Molecular, Pathologic, and Radiologic Phenotypes in Lipid-Deficient Clear Cell Renal CellCarcinoma

多尺度跨模态映射分子、病理和放射表型在脂质缺乏型清细胞癌肾癌中的应用

Ying Cui, Dongzhe Zheng, Ke Yu, Xiyin Zheng, Xiaorui Wang, Xinxiang Li, Yan Gu, Lin Fu, Xinyi Chen, Wenjie Mei, Xin-Gui Peng

机构 * Nurturing Center of Jiangsu Province for State Laboratory of AI Imaging & Interventional Radiology, Department of Radiology, Zhongda Hospital, School of Medicine, Southeast University(江苏省人工智能影像与介入放射学 state laboratory 育成中心,放射科,中大医院,东南大学) Department of Mechanical and Aerospace Engineering, Princeton University(机械与航空航天工程系,普林斯顿大学) School of Automation, Southeast University(自动化学院,东南大学) Department of Biomedical Engineering, Columbia University(生物医学工程系,哥伦比亚大学) Automation and Electrical Engineering College, Lanzhou University of Technology(自动化与电气工程学院,兰州理工大学) The First Affiliated Hospital of USTC, Division of Life Sciences and Medicine, University of Science and Technology of China(中国科学技术大学附属第一医院,生命科学与医学学院,中国科学技术大学) Lianyungang First People’s Hospital(连云港第一人民医院) School of Robotics and Automation, Suzhou Campus, Nanjing University(机器人与自动化学院,南京大学苏智校区)

专题命中 多模态评测 :cross-modal(title,abstract);分类 cs.AI

AI总结 本研究提出多尺度跨模态框架,用于术前识别脂质缺乏型清细胞肾癌的分子、病理和放射表型,实现非侵入性分子表型分析。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.15298 2025-12-18 cs.AI cs.CL cs.CY 62%

ChatGPT and Gemini participated in the Korean College Scholastic Ability Test -- Earth Science I

ChatGPT 和 Gemini 参与韩国大学入学考试——地球科学I

Seok-Hyun Ga, Chun-Yen Chang

专题命中 多模态评测 :multimodal(abstract);分类 cs.CL、cs.AI

AI总结 本研究分析了ChatGPT和Gemini在地球科学I考试中的多模态推理能力,发现模型在感知与认知之间存在差距,揭示了AI在科学评估中的局限性。

Comments 23 pages, 9 tables, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.09677 2025-12-18 cs.CV cs.AI 62%

Benchmarking Gaslighting Negation Attacks Against Reasoning Models

对抗性否定攻击下推理模型的基准测试

Bin Zhu, Hailong Yin, Jingjing Chen, Yu-Gang Jiang

机构 * Singapore Management University, Singapore College of Computer Science(新加坡国立管理学院计算机科学学院) Artificial Intelligence, Fudan University, China(人工智能,复旦大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 本文评估了三种顶级推理模型在对抗性否定攻击下的表现,发现其准确率显著下降,并提出了GaslightingBench-R基准以进一步研究模型对这类攻击的防御能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18929 2025-12-18 cs.CV 57%

Human-Centric Open-Future Task Discovery: Formulation, Benchmark, and Scalable Tree-Based Search

以人为中心的开放未来任务发现: formulation、基准和可扩展的树基搜索

Zijian Song, Xiaoxin Lin, Tao Pu, Zhenlong Yuan, Guangrun Wang, Liang Lin

机构 * Zijian Song 1(Song 研究所) Xiaoxin Lin 1(Lin 研究所) Tao Pu 1(Pu 研究所) Zhenlong Yuan 4(Yuan 研究所) Guangrun Wang 1,2,3(Wang 研究所) Liang Lin 1,2,3(Lin 研究所)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

AI总结 本研究提出HOTD问题,通过CMAST框架在开放未来任务发现中实现最佳性能,显著超越现有LMMs。

Comments accepted to AAAI 2026, 10 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏

7. 多模态Agent 2 篇

2312.09245 2025-12-18 cs.CV 83%

DriveMLM: Aligning Multi-Modal Large Language Models with Behavioral Planning States for Autonomous Driving

DriveMLM: 通过行为规划状态对齐多模态大语言模型以实现自动驾驶

Erfei Cui, Wenhai Wang, Zhiqi Li, Jiangwei Xie, Haoming Zou, Hanming Deng, Gen Luo, Lewei Lu, Xizhou Zhu, Jifeng Dai

机构 * Department of Electronic Engineering, Tsinghua University(清华大学电子工程系) Beijing National Research Center for Information Science and Technology(北京信息科学与技术国家研究中心)

专题命中 多模态Agent :multi-modal(title);multimodal(abstract);MLLM(abstract);分类 cs.CV

AI总结 DriveMLM通过多模态大语言模型对齐行为规划状态,提升自动驾驶系统的决策能力与闭环性能。

Comments Accepted to Visual Intelligence

Journal ref Visual Intelligence, Volume 3, article number 22, (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07400 2025-12-18 cs.MA cs.AI cs.CV cs.LG 76%

MedChat: A Multi-Agent Framework for Multimodal Diagnosis with Large Language Models

MedChat: 一种基于大语言模型的多智能体多模态诊断框架

Philip R. Liu, Sparsh Bansal, Jimmy Dinh, Aditya Pawar, Ramani Satishkumar, Shail Desai, Neeraj Gupta, Xin Wang, Shu Hu

机构 * University at Albany, State University of New York(纽约州立大学阿布扎克分校)

专题命中 多模态Agent :multimodal(title);分类 cs.CV、cs.AI

AI总结 MedChat通过多智能体框架结合专用视觉模型和角色特定LLM代理,提升多模态诊断的可靠性与交互性。

Journal ref Proc. 2025 IEEE 8th International Conference on Multimedia Information Processing and Retrieval (MIPR), pp. 456-462, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

8. 多模态训练与对齐 5 篇

2512.15707 2025-12-18 cs.CV 83%

GateFusion: Hierarchical Gated Cross-Modal Fusion for Active Speaker Detection

GateFusion:用于活动说话检测的分层门控跨模态融合

Yu Wang, Juhyung Ha, Frangil M. Ramirez, Yuchen Wang, David J. Crandall

机构 * Indiana University(印第安纳大学)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV

AI总结 GateFusion通过分层门控融合解码器提升活动说话检测的跨模态融合效果,实现新的SOTA结果。

Comments accepted by WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏