arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-12-11 至 2025-12-11 共收录 14 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 14 篇

2512.09872 2025-12-11 cs.CR cs.AI 83%

FlipLLM: Efficient Bit-Flip Attacks on Multimodal LLMs using Reinforcement Learning

FlipLLM: 通过强化学习高效攻击多模态大语言模型的位翻转攻击

Khurram Khalil, Khaza Anuarul Hoque

机构 * Department of Electrical Engineering and Computer Science(电气工程与计算机科学系) University of Missouri-Columbia(密苏里大学哥伦比亚分校)

专题命中 多模态评测 :multimodal(title,abstract);multimodal foundation model(abstract);分类 cs.AI

AI总结 FlipLLM通过强化学习高效识别多模态大语言模型的位翻转攻击漏洞,显著提升攻击检测速度和防御指导价值。

Comments Accepted in IEEE HOST 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02022 2025-12-11 cs.CV 83%

Do You See Me : A Multidimensional Benchmark for Evaluating Visual Perception in Multimodal LLMs

你看见我:一个多维基准用于评估多模态大语言模型的视觉感知

Aditya Kanade, Tanuja Ganu

机构 * Microsoft Research India(微软印度研究院)

专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

AI总结 该研究提出‘你看见我’基准,评估多模态大语言模型的视觉感知能力,发现其在复杂任务中表现远低于人类。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19366 2025-12-11 cs.LG cs.AI 79%

Grounding the Ungrounded: A Spectral-Graph Framework for Quantifying Hallucinations in Multimodal LLMs

未接地的接地:一种基于谱-图框架的多模态大语言模型幻觉量化方法

Supratik Sarkar, Swagatam Das

机构 * Morgan Stanley(摩根士丹利) Indian Statistical Institute (Kolkata)(印度统计研究所(加尔各答))

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

AI总结 本文提出了一种基于谱-图框架的多模态大语言模型幻觉量化方法,通过谱分解和RKHS本征模式,提供可解释的语义失真度量。

Comments 49 pages, 3 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.16954 2025-12-11 q-bio.QM 78%

Deep Multi-modal Breast Cancer Detection Network

深度多模态乳腺癌检测网络

Noor Ul Huda Shah, Tanveer Hussain, Amr Ahmed, Yonghuai Liu, Usman Ali, Ardhendu Behera

专题命中 多模态评测 :multi-modal(title,abstract)

AI总结 本文提出深度多模态乳腺癌检测网络,整合视觉与临床数据提升诊断准确率,实验显示在Mini-DDSM数据集上准确率提升至90.87%。

Comments The paper is withdrawn because we identified an error in the dataset splitting procedure used in the experiments for MMCDNET. The training/validation split was incorrectly implemented, leading to data leakage and invalid performance results

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.09251 2025-12-11 cs.CV cs.AI 76%

GLACIA: Instance-Aware Positional Reasoning for Glacial Lake Segmentation via Multimodal Large Language Model

GLACIA:基于多模态大语言模型的冰湖分割实例感知位置推理

Lalit Maurya, Saurabh Kaushik, Beth Tellman

机构 * Portsmouth AI and Data Science Centre (PAIDS), School of Computing, University of Portsmouth(普森大学计算学院) Center for Sustainability and the Global Environment (SAGE), University of Wisconsin–Madison(威斯康星大学麦迪逊分校可持续性与全球环境中心)

专题命中 多模态评测 :multimodal(title);分类 cs.CV、cs.AI

AI总结 GLACIA通过多模态大语言模型实现冰湖分割的实例感知位置推理,提升分割精度与可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.09085 2025-12-11 cs.LG 71%

Preoperative Prognosis Assessment of Lumbar Spinal Surgery for Low Back Pain and Sciatica Patients based on Multimodalities and Multimodal Learning

基于多模态与多模态学习的腰椎手术术前预后评估

Li-Chin Chen, Jung-Nien Lai, Hung-En Lin, Hsien-Te Chen, Kuo-Hsuan Hung, Yu Tsao

专题命中 多模态评测 :multimodal(title)

AI总结 本研究通过结合东方医学和机器学习,开发了一种基于多模态数据的术前预后评估工具,准确率高达0.81,为腰椎手术患者提供术前预后预测。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23981 2025-12-11 cs.CV 70%

TeleEgo: Benchmarking Egocentric AI Assistants in the Wild

TeleEgo:在真实场景中评估第一人称AI助手的基准测试

Jiaqi Yan, Ruilong Ren, Jingren Liu, Shuning Xu, Ling Wang, Yiheng Wang, Xinlin Zhong, Yun Wang, Long Zhang, Xiangyu Chen, Changzhi Sun, Jixiang Luo, Dell Zhang, Hao Sun, Chi Zhang, Xuelong Li

机构 * Institute of Artificial Intelligence (TeleAI), China Telecom(人工智能研究院(TeleAI),中国电信)

专题命中 多模态评测 :multi-modal(abstract);omni-modal(abstract);分类 cs.CV

AI总结 TeleEgo是一个用于评估第一人称AI助手在真实场景中多模态处理能力的基准测试,通过长持续时间流媒体数据和严格评估指标,推动未来具有更强流媒体记忆能力的助手发展。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.09663 2025-12-11 cs.CV 57%

IF-Bench: Benchmarking and Enhancing MLLMs for Infrared Images with Generative Visual Prompting

IF-Bench:基于生成视觉提示的多模态大语言模型在红外图像上的基准测试与增强

Tao Zhang, Yuyang Hong, Yang Xia, Kun Ding, Zeyu Zhang, Ying Wang, Shiming Xiang, Chunhong Pan

机构 * MAIS, Institute of Automation(自动化研究所信息处理中心) School of Artificial Intelligence, UCAS(中国科学院大学人工智能学院) Research Center of Aerospace Information, Institute of Automation(自动化研究所航空信息研究中心)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

AI总结 IF-Bench通过生成视觉提示方法提升多模态大语言模型对红外图像的理解能力,提供首个高质量基准测试及实验验证。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.09591 2025-12-11 cs.LG cs.AI 57%

Stanford Sleep Bench: Evaluating Polysomnography Pre-training Methods for Sleep Foundation Models

斯坦福睡眠基准:评估多导睡眠图预训练方法用于睡眠基础模型

Magnus Ruud Kjaer, Rahul Thapa, Gauri Ganjoo, Hyatt Moore, Poul Joergen Jennum, Brandon M. Westover, James Zou, Emmanuel Mignot, Bryan He, Andreas Brink-Kjaer

机构 * Stanford University(斯坦福大学) Technical University of Denmark(技术大学) Danish Center for Sleep Medicine(丹麦睡眠医学中心) University of Copenhagen(哥本哈根大学) Harvard Medical School(哈佛医学院)

专题命中 多模态评测 :multimodal(abstract);分类 cs.AI

AI总结 斯坦福睡眠基准通过大规模多导睡眠图数据集评估多种自监督预训练方法,提升睡眠分析的准确性和可重复性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20209 2025-12-11 cs.LG cs.AI 57%

Assessing the Feasibility of Early Cancer Detection Using Routine Laboratory Data: An Evaluation of Machine Learning Approaches on an Imbalanced Dataset

利用常规实验室数据评估早期癌症检测的可行性:对机器学习方法在不平衡数据集上的评估

Shumin Li

专题命中 多模态评测 :multi-modal(abstract);分类 cs.AI

AI总结 本研究评估了利用常规实验室数据通过机器学习检测癌症的可行性,发现其在临床分类上表现欠佳,需整合多模态数据以提升效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.08934 2025-12-11 cs.HC cs.AI 57%

Motion2Meaning: A Clinician-Centered Framework for Contestable LLM in Parkinson's Disease Gait Interpretation

Motion2Meaning: 一种以临床为中心的可挑战LLM框架,用于帕金森病步态解读

Loc Phuc Truong Nguyen, Hung Thanh Do, Hung Truong Thanh Nguyen, Hung Cao

机构 * Friedrich-Alexander-Universität Erlangen-Nürnberg(埃朗根-纽伦堡弗里德里希-亚历山大大学) University of New Brunswick(新 Brunswick大学)

专题命中 多模态评测 :cross-modal(abstract);分类 cs.AI

AI总结 Motion2Meaning通过结合可解释AI和可挑战LLM,为帕金森病步态解读提供透明、可审计的临床系统。

Comments Accepted at the 9th International Symposium on Chatbots and Human-Centered AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.09205 2025-12-11 physics.plasm-ph cond-mat.mtrl-sci physics.bio-ph 50%

Activation of Polylactic Acid and Polycarbonate Surfaces with Non-Thermal Plasma

非热等离子体激活聚乳酸和聚碳酸酯表面

Jairo Rondón, Ginger Urrutia, Angel Gonzalez-Lizardo

专题命中 多模态评测 :multimodal(abstract)

AI总结 非热等离子体活化聚乳酸和聚碳酸酯表面,通过多模式分析提升细胞-材料相互作用,为下一代生物材料设计提供指导。

Comments 4 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04842 2025-12-11 cs.HC 50%

Charts-of-Thought: Enhancing LLM Visualization Literacy Through Structured Data Extraction

图表思维:通过结构化数据提取提升大语言模型的可视化素养

Amit Kumar Das, Mohammad Tarun, Klaus Mueller

专题命中 多模态评测 :multimodal(abstract)

AI总结 通过结构化数据提取方法,图表思维显著提升了大语言模型在可视化任务上的表现,使其超越人类基准,同时为复杂视觉解释任务提供了新的基准。

Comments 11 pages, 8 figures. Accepted at IEEE VIS: Visualization & Visual Analytics 2025 conference, November 2-7, 2025, Vienna, Austria

Journal ref IEEE Transactions on Visualization and Computer Graphics (TVCG), PrePrints 5555, pp. 1-11, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17353 2025-12-11 cs.CE 50%

RoadBench: A Vision-Language Foundation Model and Benchmark for Road Damage Understanding

RoadBench: 一种用于道路损伤理解的视觉-语言基础模型和基准

Xi Xiao, Yunbei Zhang, Janet Wang, Lin Zhao, Yuxiang Wei, Hengjia Li, Yanshu Li, Xinyuan Song, Xiao Wang, Swalpa Kumar Roy, Hao Xu, Tianyang Wang

专题命中 多模态评测 :multimodal(abstract)

AI总结 RoadBench通过整合视觉与文本信息,提出RoadCLIP模型,显著提升道路损伤识别性能,为基础设施监测提供新基准。

Comments Accepted by WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏