arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-02-19 至 2026-02-19 共收录 38 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 9 篇

2508.02669 2026-02-19 cs.CV 83%

MedVLThinker: Simple Baselines for Multimodal Medical Reasoning

MedVLThinker: 多模态医疗推理的简单基线

Xiaoke Huang, Juncheng Wu, Hui Liu, Xianfeng Tang, Yuyin Zhou

机构 * UC Santa Cruz(加州大学圣克ruz分校) Amazon Research(亚马逊研究院)

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CV

AI总结 MedVLThinker通过简单基线和RLVR方法,在医疗多模态推理中实现新突破,超越现有开源模型并接近专有模型性能。

Comments Project page: https://ucsc-vlaa.github.io/MedVLThinker/ ; Code: https://github.com/UCSC-VLAA/MedVLThinker ; Model and Data: https://huggingface.co/collections/UCSC-VLAA/medvlthinker-688f52224fb7ff7d965d581d ; Accepted by ML4H'25

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.15918 2026-02-19 cs.CV cs.AI 81%

EarthSpatialBench: Benchmarking Spatial Reasoning Capabilities of Multimodal LLMs on Earth Imagery

EarthSpatialBench: 多模态大语言模型在地球影像上空间推理能力的基准测试

Zelin Xu, Yupu Zhang, Saugat Adhikari, Saiful Islam, Tingsong Xiao, Zibo Liu, Shigang Chen, Da Yan, Zhe Jiang

机构 * University of Florida(佛罗里达大学) Indiana University Bloomington(印第安纳大学布卢明顿分校)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 EarthSpatialBench是一个用于评估多模态大语言模型在地球影像上空间推理能力的综合基准测试,涵盖空间距离、方向、拓扑关系及复杂几何查询。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.15903 2026-02-19 cs.CV 79%

Detecting Deepfakes with Multivariate Soft Blending and CLIP-based Image-Text Alignment

利用多变量软融合和基于CLIP的图像-文本对齐检测深度伪造

Jingwei Li, Jiaxin Tong, Pengfei Wu

机构 * Zhejiang Gongshang University(浙江工商大学)

专题命中 图文多模态 :image-text(title);multimodal(abstract);分类 cs.CV

AI总结 本文提出MSBA-CLIP框架,通过多变量软融合和CLIP引导的伪造强度估计,提升深度伪造检测的准确性和泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.16019 2026-02-19 cs.CV cs.AI 73%

MedProbCLIP: Probabilistic Adaptation of Vision-Language Foundation Model for Reliable Radiograph-Report Retrieval

MedProbCLIP: 基于概率适应的视觉-语言基础模型用于可靠放射影像-报告检索

Ahmad Elallaf, Yu Zhang, Yuktha Priya Masupalli, Jeong Yang, Young Lee, Zechun Cao, Gongbo Liang

机构 * Texas A&M University-San Antonio(德克萨斯A&M大学-圣安东尼奥分校) Boise State University(博伊州立大学)

专题命中 图文多模态 :multimodal(abstract);image-text(abstract);分类 cs.CV、cs.AI

AI总结 MedProbCLIP 通过概率适应提升放射影像与报告检索的可靠性,优于现有基线方法。

Comments Accepted to the 2026 Winter Conference on Applications of Computer Vision (WACV) Workshops

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.15873 2026-02-19 cs.RO cs.AI 70%

Test-Time Adaptation for Tactile-Vision-Language Models

测试时适应用于触觉-视觉-语言模型

Chuyang Ye, Haoxian Jing, Qinting Jiang, Yixi Lin, Qiang Li, Xing Tang, Jingyan Jiang

机构 * Shenzhen Technology University(深圳技术大学) New York University(纽约大学) Shenzhen University(深圳大学) Tsinghua University(清华大学)

专题命中 图文多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.AI

AI总结 本文提出了一种可靠性感知框架,用于提升触觉-视觉-语言模型在测试时的适应能力,通过估计各模态可靠性来过滤不可靠样本并优化融合过程,有效提升了在模态损坏情况下的准确率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25867 2026-02-19 cs.LG 67%

Synthesizing High-Quality Visual Question Answering from Medical Documents with Generator-Verifier LMMs

从医学文献合成高质量的视觉问答系统:基于生成-验证框架的大型多模态模型

Xiaoke Huang, Ningsen Wang, Hui Liu, Xianfeng Tang, Yuyin Zhou

机构 * UC Santa Cruz(加州大学圣克ruz分校) Fudan University(复旦大学) Amazon Research(亚马逊研究)

专题命中 图文多模态 :multimodal(abstract);image-text(abstract)

AI总结 MedVLSynther通过生成-验证框架从开放文献合成高质量医学VQA数据,提升六个基准测试的准确率,达到77.57的VQA-RAD表现。

Comments Project page, code, data, and models: https://ucsc-vlaa.github.io/MedVLSynther/ ; Accepted by ICLR'26

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.16430 2026-02-19 cs.CV cs.AI 62%

Designing Production-Scale OCR for India: Multilingual and Domain-Specific Systems

为印度设计生产级OCR:多语言和领域特定系统

Ali Faraz, Raja Kolla, Ashish Kulkarni, Shubham Agarwal

机构 * Krutrim AI

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出两种多语言OCR训练策略,通过Chitrapathak系列在印度多语言文档中实现SOTA性能,并展示了Parichay在政府文档中的高效提取能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13494 2026-02-19 cs.CV cs.AI 62%

Language-Guided Invariance Probing of Vision-Language Models

语言引导的视觉-语言模型不变性探测

Jae Joong Lee

机构 * Department of Computer Science, Purdue University(普渡大学计算机科学系)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV、cs.AI

AI总结 本文提出LGIP基准,用于评估视觉-语言模型对语言扰动的鲁棒性,发现EVA02-CLIP和OpenCLIP表现优异,而SigLIP存在显著缺陷。

Comments Pattern Recognition Letters 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.05023 2026-02-19 cs.CR cs.AI 57%

Do Vision-Language Models Respect Contextual Integrity in Location Disclosure?

视觉-语言模型在位置披露中是否尊重上下文完整性?

Ruixin Yang, Ethan Mendes, Arthur Wang, James Hays, Sauvik Das, Wei Xu, Alan Ritter

专题命中 图文多模态 :multimodal(abstract);分类 cs.AI

AI总结 本文研究了视觉-语言模型在位置披露中的隐私问题,指出现有模型在隐私保护方面存在不足,提出VLM-GEOPRIVACY基准以评估模型对上下文完整性的尊重程度。

Comments Accepted by ICLR 2026. Code and data can be downloaded via https://github.com/99starman/VLM-GeoPrivacyBench

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 音频语音多模态 4 篇

2602.16687 2026-02-19 cs.SD cs.CL eess.AS 62%

Scaling Open Discrete Audio Foundation Models with Interleaved Semantic, Acoustic, and Text Tokens

通过交错语义、音频和文本标记扩展开放离散音频基础模型

Potsawee Manakul, Woody Haosheng Gan, Martijn Bartelds, Guangzhi Sun, William Held, Diyi Yang

机构 * Stanford University(斯坦福大学) University of Southern California(南加州大学) University of Cambridge(剑桥大学)

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.CL、eess.AS

AI总结 本文提出SODA模型,通过交错语义、音频和文本标记扩展开放离散音频基础模型,实现通用音频生成和跨模态能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.16008 2026-02-19 cs.SD cs.AI cs.CL cs.LG 62%

MAEB: Massive Audio Embedding Benchmark

MAEB:大规模音频嵌入基准

Adnan El Assadi, Isaac Chung, Chenghao Xiao, Roman Solomatin, Animesh Jha, Rahul Chand, Silky Singh, Kaitlyn Wang, Ali Sartaz Khan, Marc Moussa Nasser, Sufen Fong, Pengfei He, Alan Xiao, Ayush Sunil Munot, Aditya Shrivastava, Artem Gazizov, Niklas Muennighoff, Kenneth Enevoldsen

机构 * Carleton University(卡尔顿大学) Durham University(杜伦大学) Stanford University(斯坦福大学) Aarhus University(奥胡斯大学) Indian Institute of Technology, Kharagpur(印度理工学院,卡里格普) Harvard University(哈佛大学)

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.CL、cs.AI

AI总结 MAEB是一个涵盖多语言音频任务的基准,揭示了不同模型在音频与语言任务上的性能差异,展示了音频编码器在跨模态任务中的表现关联性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.16334 2026-02-19 cs.SD cs.AI 57%

Spatial Audio Question Answering and Reasoning on Dynamic Source Movements

空间音频问答与动态声源运动推理

Arvind Krishna Sridhar, Yinyi Guo, Erik Visser

机构 * Qualcomm Technologies Inc.(高通技术公司)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI

AI总结 本文提出了一种空间音频问答方法,通过运动推理和多模态微调提升动态声源识别与理解能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.14172 2026-02-19 cs.SD cs.CL cs.LG 57%

Investigation for Relative Voice Impression Estimation

相对语音印象估计研究

Kenichi Fujita, Yusuke Ijima

机构 * NTT, Inc., Japan(日本NTT公司)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL

AI总结 本研究探讨了相对语音印象估计方法,发现自监督语音模型在捕捉复杂感知差异方面表现更优,而多模态大语言模型在此任务中效果不佳。

Comments 5 pages,3 figures, Accepted to Speech Prosody 2026

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频多模态 5 篇

2602.16147 2026-02-19 cs.LG cs.AI cs.HC eess.SP 70%

ASPEN: Spectral-Temporal Fusion for Cross-Subject Brain Decoding

ASPEN:跨受试者脑解码的频谱-时间融合

Megan Lee, Seung Ha Hwang, Inhyeok Choi, Shreyas Darade, Mengchun Zhang, Kateryna Shapovalenko

机构 * Carnegie Mellon University(卡内基梅隆大学) Kyung Hee University(Kyung Hee大学) Korea Advanced Institute of Science and Technology(韩国科学技术院) University of Pittsburgh(匹兹堡大学)

专题命中 视频多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.AI

AI总结 ASPEN通过频谱-时间融合实现跨受试者脑解码,利用乘法融合提升跨模态一致性,有效提升未见受试者准确率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.05523 2026-02-19 cs.CV cs.AI 62%

Prompt When the Animal is: Temporal Animal Behavior Grounding with Positional Recovery Training

提示当动物是:基于位置恢复训练的时序动物行为 grounding

Sheng Yan, Xin Du, Zongying Li, Yi Wang, Hongcang Jin, Mengyuan Liu

机构 * School of Artificial Intelligence, Chongqing University of Technology(重庆理工大学人工智能学院) National Key Laboratory of General Artificial Intelligence, Shenzhen Graduate School, Peking University(北京大学深圳研究生院国家通用人工智能实验室)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出 Port 框架,通过位置恢复训练提升动物行为时序 grounding 的准确性,在 Animal Kingdom 数据集上取得 38.52 的 IoU@0.3 成绩,并在 ICME 2024 挑战赛中表现突出。

Comments Accepted by ICMEW 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.15862 2026-02-19 cs.CL cs.AI 62%

Enhancing Action and Ingredient Modeling for Semantically Grounded Recipe Generation

增强语义 grounded 的动作和成分建模以实现语义 grounded 的食谱生成

Guoshan Liu, Bin Zhu, Yian Li, Jingjing Chen, Chong-Wah Ngo, Yu-Gang Jiang

专题命中 视频多模态 :multimodal(abstract);分类 cs.CL、cs.AI

AI总结 本文提出语义 grounded 的框架,通过两阶段流程结合监督微调和强化微调,提升食谱生成的语义准确性与成分预测能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.17091 2026-02-19 cs.CV cs.AI cs.LG 62%

Ctrl-GenAug: Controllable Generative Augmentation for Medical Sequence Classification

Ctrl-GenAug: 可控生成增强用于医学序列分类

Xinrui Zhou, Yuhao Huang, Haoran Dou, Shijing Chen, Ao Chang, Jia Liu, Weiran Long, Jian Zheng, Erjiao Xu, Jie Ren, Alejandro F. Frangi, Ruobing Huang, Jun Cheng, Xiaomeng Li, Wufeng Xue, Dong Ni

机构 * National-Regional Key Technology Engineering Laboratory for Medical Ultrasound(国家级区域医疗超声关键技术工程实验室) School of Biomedical Engineering(生物医学工程学院) Medical School(医学院) Shenzhen University(深圳大学) Medical UltraSound Image Computing (MUSIC) Lab(医学超声图像计算(MUSIC)实验室) School of Artificial Intelligence(人工智能学院) The Hong Kong University of Science and Technology(香港科学与技术大学) Department of Electronic and Computer Engineering(电子与计算机工程系) The University of Manchester(曼彻斯特大学) The Third Affiliated Hospital of Sun Yat-sen University(中山大学第三附属医院) Longgang District People’s Hospital of Shenzhen(深圳龙岗区人民医院) The Second Affiliated Hospital of The Chinese University of Hong Kong(香港中文大学第二附属医院) School of Biomedical Engineering and Informatics(生物医学工程与信息学学院) NIHR Manchester Biomedical Research Centre(NIHR曼彻斯特生物医学研究中心)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 Ctrl-GenAug通过可控生成增强方法提升医学序列分类性能,解决语义和序列可控性不足及噪声样本问题。

Comments Accepted by International Journal of Computer Vision, 30 pages, 11 figures, 11 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.08822 2026-02-19 cs.RO cs.AI 57%

FreqPolicy: Efficient Flow-based Visuomotor Policy via Frequency Consistency

FreqPolicy: 通过频率一致性实现高效的基于流的视觉-运动策略

Yifei Su, Ning Liu, Dong Chen, Zhen Zhao, Kun Wu, Meng Li, Zhiyuan Xu, Zhengping Che, Jian Tang

机构 * Beijing Innovation Center of Humanoid Robotics(北京人形机器人创新中心) NLPR, MAIS, Institute of Automation of Chinese Academy of Sciences(神经语言处理实验室、人工智能研究所、中国科学院自动化研究所)

专题命中 视频多模态 :multimodal(abstract);分类 cs.AI

AI总结 FreqPolicy通过频率一致性约束提升基于流的视觉-运动策略的效率与质量,实现高效、高质量的动作生成。

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 多模态生成 2 篇

2602.16006 2026-02-19 cs.CV 57%

BTReport: A Framework for Brain Tumor Radiology Report Generation with Clinically Relevant Features

BTReport: 一种用于脑肿瘤放射学报告生成的框架,结合临床相关特征

Juampablo E. Heras Rivera, Dickson T. Chen, Tianyi Ren, Daniel K. Low, Asma Ben Abacha, Alberto Santamaria-Pang, Mehmet Kurt

机构 * University of Washington(华盛顿大学) University of Washington School of Medicine(华盛顿大学医学院) Microsoft Health AI(微软健康人工智能) Johns Hopkins School of Medicine(约翰霍普金斯医学院)

专题命中 多模态生成 :image-text(abstract);分类 cs.CV

AI总结 BTReport通过确定性特征提取和大语言模型结合,生成可解释的脑肿瘤放射学报告,并提供配套数据集提升临床应用

Comments Accepted to Medical Imaging with Deep Learning (MIDL) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22007 2026-02-19 cs.LG 50%

Stage-wise Dynamics of Classifier-Free Guidance in Diffusion Models

扩散模型中分类器自由引导的分阶段动态

Cheng Jin, Qitan Shi, Yuantao Gu

机构 * Department of Electronic Engineering, Tsinghua University(电子工程系,清华大学)

专题命中 多模态生成 :multimodal(abstract)

AI总结 本文研究了扩散模型中分类器自由引导的分阶段动态,揭示了引导对采样过程三个阶段的影响,并提出时间变化的引导策略以提升质量和多样性。

Comments 24 pages, 10 figures, accepted by ICLR26

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 多模态评测 5 篇

2509.17544 2026-02-19 cs.AI 79%

A Multimodal Conversational Assistant for the Characterization of Agricultural Plots from Geospatial Open Data

一种多模态对话助手,用于从遥感开放数据中表征农业地块

Juan Cañada, Raúl Alonso, Julio Molleda, Fidel Díez

机构 * Dept. AI \& Data Analytics CTIC Technology Center Gijón, Spain Dept. IT CTIC Technology Center Gijón, Spain Eng. University of Oviedo Gijón, Spain Dept. R \& D CTIC Technology Center Gijón, Spain

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

AI总结 本研究提出一种多模态对话助手,通过自然语言交互整合遥感和农业数据,降低专业信息访问门槛,实现农业地块表征。

Comments Accepted at 2025 4th International Conference on Geographic Information and Remote Sensing Technology

Journal ref Proc. Int. Conf. GIRST 2025 (2025) 127-132

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09103 2026-02-19 cs.RO 78%

IMPACT: Behavioral Intention-aware Multimodal Trajectory Prediction with Adaptive Context Trimming

IMPACT: 基于适应性上下文修剪的多模态轨迹预测:行为意图感知

Jiawei Sun, Xibin Yue, Jiahui Li, Tianle Shen, Chengran Yuan, Shuo Sun, Sheng Guo, Quanyun Zhou, Marcelo H Ang

机构 * National University of Singapore(新加坡国立大学) Xiaomi EV(小米电动车)

专题命中 多模态评测 :multimodal(title,abstract)

AI总结 IMPACT提出了一种基于适应性上下文修剪的多模态轨迹预测框架,通过联合预测行为意图和轨迹,提升预测精度、可解释性和效率,并在Waymo数据集上取得优异成绩。

Comments accepted by IEEE Robotics and Automation Letters

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18318 2026-02-19 cs.AI 74%

Earth AI: Unlocking Geospatial Insights with Foundation Models and Cross-Modal Reasoning

Earth AI:利用基础模型和跨模态推理解锁地理空间洞察

Aaron Bell, Amit Aides, Amr Helmy, Arbaaz Muslim, Aviad Barzilai, Aviv Slobodkin, Bolous Jaber, David Schottlander, George Leifman, Joydeep Paul, Mimi Sun, Nadav Sherman, Natalie Williams, Per Bjornsson, Roy Lee, Ruth Alcantara, Thomas Turnbull, Tomer Shekel, Vered Silverman, Yotam Gigi, Adam Boulanger, Alex Ottenwess, Ali Ahmadalipour, Anna Carter, Behzad Vahedi, Charles Elliott, David Andre, Elad Aharoni, Gia Jung, Hassler Thurston, Jacob Bien, Jamie McPike, Jessica Sapick, Juliet Rothenberg, Kartik Hegde, Kel Markert, Kim Philipp Jablonski, Luc Houriez, Monica Bharel, Phing VanLee, Reuven Sayag, Sebastian Pilarski, Shelley Cazares, Shlomi Pasternak, Siduo Jiang, Thomas Colthurst, Yang Chen, Yehonathan Refael, Yochai Blau, Yuval Carny, Yael Maguire, Avinatan Hassidim, James Manyika, Tim Thelin, Genady Beryozkin, Gautam Prasad, Luke Barrington, Yossi Matias, Niv Efron, Shravya Shetty

机构 * Google Research(谷歌研究) Google X(谷歌X) Google Cloud(谷歌云)

专题命中 多模态评测 :cross-modal(title);分类 cs.AI

AI总结 Earth AI通过基础模型和跨模态推理,提升地理空间数据分析能力,实现对地球的深入洞察。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.15958 2026-02-19 cs.CL cs.AI cs.CV 67%

DocSplit: A Comprehensive Benchmark Dataset and Evaluation Approach for Document Packet Recognition and Splitting

DocSplit: 一个全面的基准数据集和评估方法用于文档包识别与分割

Md Mofijul Islam, Md Sirajus Salekin, Nivedha Balakrishnan, Vincil C. Bishop, Niharika Jain, Spencer Romo, Bob Strahan, Boyi Xie, Diego A. Socolinsky

机构 * Amazon Web Services(亚马逊网络服务)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 DocSplit提出首个全面的文档包识别与分割基准数据集及评估方法,涵盖多类型文档和复杂场景,揭示当前模型在处理复杂文档分割任务中的性能差距。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15242 2026-02-19 cs.CV 57%

Trustworthy and Fair SkinGPT-R1 for Democratizing Dermatological Reasoning across Diverse Ethnicities

可信且公平的SkinGPT-R1:用于在不同种族中普及皮肤病推理

Yuhao Shen, Zhangtianyi Chen, Yuanhao He, Yan Xu, Shuping Zhang, Liyuan Sun, Zijian Wang, Yinghao Zhu, Yuyuan Yang, Jiahe Qian, Ziwen Wang, Xinyuan Zhang, Wenbin Liu, Zongyuan Ge, Tao Lu, Siyuan Yan, Juexiao Zhou

机构 * School of Data Science, The Chinese University of Hong Kong, Shenzhen (CUHK-Shenzhen)(数据科学学院,香港中文大学(深圳)) Faculty of Information Technology, Monash University(信息技术学院,莫纳什大学) Department of Dermatology, Tianjin Institute of Integrative Dermatology, Tianjin Academy of Traditional Chinese Medicine Affiliated Hospital(皮肤科,天津整合皮肤科研究所,天津中医研究院附属医院) Department of Dermatology, The First Affiliated Hospital, Shantou University Medical College(皮肤科,汕头大学医学院第一附属医院) Department of Dermatology, Beijing AnZhen Hospital, Capital Medical University(皮肤科,北京安贞医院,首都医科大学) School of Computing and Data Science, The University of Hong Kong(计算与数据科学学院,香港大学) Institute of Automation, Chinese Academy of Sciences(自动化研究所,中国科学院) Department of Dermatology, Beijing Aerospace General Hospital(皮肤科,北京航天总医院)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

AI总结 SkinGPT-R1通过公平性意识的专家混合架构实现可解释且公平的皮肤病诊断,提升不同种族间的诊断准确性与安全性。

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 多模态Agent 3 篇

2602.16493 2026-02-19 cs.CV 79%

MMA: Multimodal Memory Agent

MMA:多模态记忆代理

Yihao Lu, Wanru Cheng, Zeyu Zhang, Hao Tang

机构 * School of Computer Science, Peking University(北京大学计算机学院)

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV

AI总结 MMA通过动态可靠性评分和冲突感知网络共识,提升多模态代理在长horizon任务中的记忆检索与决策能力,同时在多个基准测试中展现优越性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.15864 2026-02-19 cs.RO cs.CV 70%

ReasonNavi: Human-Inspired Global Map Reasoning for Zero-Shot Embodied Navigation

ReasonNavi: 人类启发的全局地图推理用于零样本具身导航

Yuzhuo Ao, Anbang Wang, Yu-Wing Tai, Chi-Keung Tang

机构 * The Hong Kong University of Science and Technology(香港科学与技术大学) Dartmouth College(达特茅斯学院)

专题命中 多模态Agent :multimodal(abstract);MLLM(abstract);分类 cs.CV

AI总结 ReasonNavi通过结合多模态大语言模型与确定性规划器,实现人类启发的全局地图推理,提供无需微调的零样本具身导航解决方案。

Comments 18 pages, 6 figures, Project page: https://reasonnavi.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12707 2026-02-19 cs.LG cs.AI cs.MA 57%

PLAICraft: Large-Scale Time-Aligned Vision-Speech-Action Dataset for Embodied AI

PLAICraft: 大规模时间对齐的视觉-语音-动作数据集用于具身人工智能

Yingchen He, Christian D. Weilbach, Martyna E. Wojciechowska, Yuxuan Zhang, Frank Wood

机构 * University of British Columbia(不列颠哥伦比亚大学)

专题命中 多模态Agent :multi-modal(abstract);分类 cs.AI

AI总结 PLAICraft通过大规模时间对齐的视觉-语音-动作数据集,推动具身人工智能的研究与评估。

Comments 9 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏

7. 多模态训练与对齐 7 篇

2602.16245 2026-02-19 cs.CV 79%

HyPCA-Net: Advancing Multimodal Fusion in Medical Image Analysis

HyPCA-Net:在医学图像分析中推进多模态融合

J. Dhar, M. K. Pandey, D. Chakladar, M. Haghighat, A. Alavi, S. Mistry, N. Zaidi

机构 * Indian Institute of Technology Ropar(印度理工学院罗帕尔分校) RoentGen Health(RoentGen健康公司) Lulea University of Technology(卢勒奥大学) QUT(昆士兰科技大学) RMIT University(皇家墨尔本理工大学) Curtin University(Curtin大学) Deakin University(德肯大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 HyPCA-Net通过高效残差注意力模块和双视角级联注意力模块,提升多模态医学图像分析的性能与效率。

Comments Accepted at the IEEE/CVF Winter Conference on Applications of Computer Vision 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.15896 2026-02-19 cs.CL 79%

Every Little Helps: Building Knowledge Graph Foundation Model with Fine-grained Transferable Multi-modal Tokens

每一点帮助都有价值:通过细粒度可迁移多模态标记构建知识图谱基础模型

Yichi Zhang, Zhuo Chen, Lingbing Guo, Wen Zhang, Huajun Chen

机构 * Zhejiang University, Zhejiang, China(浙江大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CL

AI总结 本文提出TOFU模型,通过细粒度可迁移多模态标记提升多模态知识图谱推理的跨图谱迁移能力。

Comments Work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏