arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-02-16 至 2026-02-16 共收录 51 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 1 篇

2602.12461 2026-02-16 cs.CV 57%

Semantic-aware Adversarial Fine-tuning for CLIP

语义感知对抗微调用于CLIP

Jiacheng Zhang, Jinhao Li, Hanxun Huang, Sarah M. Erfani, Benjamin I. P. Rubinstein, Feng Liu

机构 * School of Computing and Information Systems(计算机与信息系统学院) The University of Melbourne(墨尔本大学)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

AI总结 本文提出语义感知对抗微调方法,通过生成语义感知的对抗性示例提升CLIP模型在零样本分类任务中的对抗鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 音频语音多模态 7 篇

2602.12597 2026-02-16 cs.RO cs.HC 78%

PISHYAR: A Socially Intelligent Smart Cane for Indoor Social Navigation and Multimodal Human-Robot Interaction for Visually Impaired People

PISHYAR:一种具有社会智能的智能拐杖,用于室内社交导航和多模态人机交互,以支持视障人士

Mahdi Haghighat Joo, Maryam Karimi Jafari, Alireza Taheri

专题命中 音频语音多模态 :multimodal(title,abstract)

AI总结 PISHYAR是一款结合社会智能导航与多模态交互的智能拐杖,通过多模态技术提升视障人士的移动辅助与社交互动能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07969 2026-02-16 eess.AS cs.AI cs.LG cs.SD 62%

Tuberculosis Screening from Cough Audio: Baseline Models, Clinical Variables, and Uncertainty Quantification

咳嗽音频中肺结核筛查:基线模型、临床变量和不确定性量化

George P. Kafentzis, Efstratios Selisios

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI、eess.AS

AI总结 本文提出了一种标准化框架,用于通过咳嗽音频和临床数据检测肺结核,通过建立基线模型和量化不确定性,提供公平比较的参考点。

Comments Updated to published version in Sensors; DOI: 10.3390/s26041223

Journal ref Sensors 2026, 26(4), 1223

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.12302 2026-02-16 cs.CL cs.CV 62%

Grandes Modelos de Linguagem Multimodais (MLLMs): Da Teoria à Prática

大规模多模态语言模型(MLLMs):从理论到实践

Neemias da Silva, Júlio C. W. Scholz, John Harrison, Marina Borges, Paulo Ávila, Frances A Santos, Myriam Delgado, Rodrigo Minetto, Thiago H Silva

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV、cs.CL

AI总结 本文介绍了多模态大语言模型(MLLMs)的理论基础和实践方法,探讨了预处理、提示工程及多模态管道构建技术,并提供了补充材料供进一步研究。

Comments in Portuguese language. Accepted book chapter - Webmedia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.12918 2026-02-16 cs.RO 50%

Adding internal audio sensing to internal vision enables human-like in-hand fabric recognition with soft robotic fingertips

添加内部音频感知至内部视觉可使软机器人指尖实现类人手持织物识别

Iris Andrussow, Jans Solano, Benjamin A. Richardson, Georg Martius, Katherine J. Kuchenbecker

机构 * Haptic Intelligence Department, Max Planck Institute for Intelligent Systems(马克斯·普朗克智能系统研究所人机交互部门) Department for Distributed Intelligence / Autonomous Learning, University of Tübingen(图宾根大学分布式智能/自主学习部门)

专题命中 音频语音多模态 :audio-visual(abstract)

AI总结 本研究通过结合内部视觉和音频感知,使软机器人指尖实现类人织物识别,利用变压器方法在 20 种常见织物上达到 97% 分类准确率。

Journal ref 2025 IEEE-RAS 24th International Conference on Humanoid Robots (Humanoids)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13148 2026-02-16 cs.SD 50%

Can Large Audio Language Models Understand Audio Well? Speech, Scene and Events Understanding Benchmark for LALMs

大音频语言模型能理解音频吗?LALMs的语音、场景和事件理解基准

Han Yin, Jung-Woo Choi

机构 * School of Electrical Engineering, KAIST, Daejeon, Republic of Korea(韩国成均馆大学电气工程学院)

专题命中 音频语音多模态 :cross-modal(abstract)

AI总结 本文提出SSEU-Bench,首个考虑语音与非语音音频能量差异的音频理解基准,通过链式思维提升LALMs在联合理解任务中的性能。

Comments Accepted by ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11906 2026-02-16 cs.RO 50%

Palpation Alters Auditory Pain Expressions with Gender-Specific Variations in Robopatients

触诊改变与性别相关的听觉疼痛表达在仿生患者中

Chapa Sirithunge, Yue Xie, Saitarun Nadipineni, Fumiya Iida, Thilina Dulantha Lalitharatne

机构 * Department of Engineering, University of Cambridge(剑桥大学工程系) School of Engineering and Materials Science in Queen Mary University of London(伦敦大学玛丽女王学院工程与材料科学学院)

专题命中 音频语音多模态 :multimodal(abstract)

AI总结 本文提出了一种基于人类在环强化学习的仿生患者听觉疼痛表达生成方法,通过动态调整触诊力与声音反馈的关系,实现性别差异的个性化疼痛表达,以提升医疗模拟训练效果。

Comments 12 pages, 9 figures, journal

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.12416 2026-02-16 cs.RO cs.SY eess.SY 50%

Control Barrier Functions with Audio Risk Awareness for Robot Safe Navigation on Construction Sites

具有音频风险意识的控制障碍函数用于建筑工地机器人安全导航

Johannes Mootz, Reza Akhavian

专题命中 音频语音多模态 :multimodal(abstract)

AI总结 本文提出一种基于控制障碍函数的机器人安全导航方法,通过音频风险提示增强避障能力,有效提升建设工地环境下的自主导航安全性。

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频多模态 3 篇

2602.12641 2026-02-16 cs.NI cs.AI cs.HC cs.MM 84%

Artic: AI-oriented Real-time Communication for MLLM Video Assistant

Artic: 面向AI的实时通信用于多模态大语言模型视频助手

Jiangkai Wu, Zhiyuan Ren, Junquan Zhong, Liming Liu, Xinggong Zhang

机构 * Peking University(北京大学)

专题命中 视频多模态 :MLLM(title,abstract);multimodal(abstract);分类 cs.AI、cs.MM

AI总结 Artic提出了一种面向AI的实时通信框架,通过自适应比特率、上下文感知流式传输和退化视频理解基准提升多模态大语言模型视频助手的准确性和降低延迟。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.12593 2026-02-16 cs.IR cs.AI 79%

RQ-GMM: Residual Quantized Gaussian Mixture Model for Multimodal Semantic Discretization in CTR Prediction

RQ-GMM:用于CTR预测的多模态语义离散化残差量化高斯混合模型

Ziye Tong, Jiahao Liu, Weimin Zhang, Hongji Ruan, Derick Tang, Zhanpeng Zeng, Qinsong Zeng, Peng Zhang, Tun Lu, Ning Gu

机构 * Tencent(腾讯) Fudan University(复旦大学) Beijing Jiaotong University(北京交通大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

AI总结 RQ-GMM通过残差量化高斯混合模型提升CTR预测中多模态语义离散化效果,实现代码本利用和重建准确性的显著提升。

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13082 2026-02-16 cs.SI 50%

Revealing Process Structure in Urban Mobility Networks

揭示城市交通网络中的过程结构

Khristina Filonchik, Jose Pedro Pinto, Flávio L. Pinheiro, Fernando Bacao

专题命中 视频多模态 :multimodal(abstract)

AI总结 本文通过过程挖掘技术分析城市交通数据,揭示交通行为的结构化特征,并提出以对象为中心的模型以提升多模式交通规划的效率。

Comments This paper was presented at the Fourteenth International Conference on Complex Networks & Their Applications (Complex Networks 2025), Binghamton, USA, and appears in the conference proceedings

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 跨模态检索 8 篇

2602.12936 2026-02-16 cs.CV 85%

Unleashing MLLMs on the Edge: A Unified Framework for Cross-Modal ReID via Adaptive SVD Distillation

在边缘侧释放MLLMs:通过自适应SVD蒸馏实现跨模态重识别的统一框架

Hongbo Jiang, Jie Li, Xinqi Cai, Tianyu Xie, Yunhang Shen, Pingyang Dai, Liujuan Cao

机构 * Tencent Youtu Lab, Shanghai, China(腾讯优图实验室,上海,中国) Xiamen University, Media Analytics(厦门大学,媒体分析) Computing Lab, Department of Artificial Intelligence, School of Informatics, Xiamen, China(计算实验室,人工智能系,信息学院,厦门,中国)

专题命中 跨模态检索 :cross-modal(title,abstract);multi-modal(abstract);MLLM(abstract);分类 cs.CV

AI总结 本文提出MLLMEmbed-ReID框架,通过自适应SVD蒸馏实现跨模态重识别的统一部署,结合云-边缘架构提升性能与效率。

Comments Equal contribution by Jie Li

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.07909 2026-02-16 cs.LG cs.AI cs.CV 84%

Explaining and Mitigating the Modality Gap in Contrastive Multimodal Learning

解释并缓解对比多模态学习中的模态差距

Can Yaras, Siyi Chen, Peng Wang, Qing Qu

机构 * Department of Electrical Engineering \& Computer Science, University of Michigan

专题命中 跨模态检索 :multimodal(title,abstract);image-text(abstract);分类 cs.CV、cs.AI

AI总结 本文通过分析对比多模态学习中的模态差距成因,提出通过温度调度和模态交换等策略缓解模态差距,从而提升多模态任务性能。

Comments The first two authors contributed equally to this work

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.10271 2026-02-16 cs.IR cs.CL 83%

MLDocRAG: Multimodal Long-Context Document Retrieval Augmented Generation

MLDocRAG: 多模态长上下文文档检索增强生成

Yongyue Zhang, Yaxiong Wu

机构 * Independent Researcher(独立研究者)

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL

AI总结 MLDocRAG通过构建多模态片段-查询图,实现多模态长上下文文档的检索增强生成,提升问答的准确性和连贯性。

Comments 15 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01390 2026-02-16 cs.CV cs.AI cs.LG 81%

Multimodal Doctor-in-the-Loop: A Clinically-Guided Explainable Framework for Predicting Pathological Response in Non-Small Cell Lung Cancer

多模态医生在环:一种临床指导的可解释框架,用于预测非小细胞肺癌的病理反应

Alice Natalina Caragliano, Claudia Tacconi, Carlo Greco, Lorenzo Nibid, Edy Ippolito, Michele Fiore, Giuseppe Perrone, Sara Ramella, Paolo Soda, Valerio Guarrasi

机构 * Research Unit of Computer Systems and Bioinformatics, Department of Engineering(计算机系统与生物信息学研究单位,工程系) Research Unit of Radiation Oncology, Department of Medicine and Surgery(放射肿瘤学研究单位,医学与外科系) Research Unit of Anatomical Pathology, Department of Medicine and Surgery(病理学研究单位,医学与外科系)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 本研究提出多模态医生在环框架,结合影像与临床数据,通过嵌入医生知识提升肺癌病理反应预测的准确性和可解释性。

Comments arXiv admin note: substantial text overlap with arXiv:2502.17503

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.12774 2026-02-16 cs.CV 79%

Bootstrapping MLLM for Weakly-Supervised Class-Agnostic Object Counting

基于MLLM的弱监督类无关物体计数

Xiaowen Zhang, Zijie Yue, Yong Luo, Cairong Zhao, Qijun Chen, Miaojing Shi

机构 * College of Electronic and Information Engineering, Tongji University(电子信息工程学院,同济大学) College of Computer Science, Wuhan University(计算机科学学院,武汉大学) College of Computer Science, Tongji University(计算机科学学院,同济大学) State Key Laboratory of Autonomous Intelligent Unmanned Systems(自主智能无人系统国家重点实验室)

专题命中 跨模态检索 :MLLM(title,abstract);分类 cs.CV

AI总结 本文提出 WS-COC,基于 MLLM 的弱监督框架,通过三种策略提升类无关物体计数性能,有效降低标注成本。

Comments Accepted at ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.05096 2026-02-16 cs.CV cs.LG 79%

Visual concept ranking uncovers medical shortcuts used by large multimodal models

视觉概念排名揭示大型多模态模型使用的医学捷径

Joseph D. Janizek, Sonnet Xu, Junayd Lateef, Roxana Daneshjou

机构 * Stanford University(斯坦福大学) University of California Berkeley(加州大学伯克利分校)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出视觉概念排名方法,用于揭示大型多模态模型在医疗任务中表现的潜在视觉特征依赖性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13179 2026-02-16 cs.IR 78%

Fix Before Search: Benchmarking Agentic Query Visual Pre-processing in Multimodal Retrieval-augmented Generation

在搜索前修复:多模态检索增强生成中代理查询视觉预处理的基准测试

Jiankun Zhang, Shenglai Zeng, Kai Guo, Xinnan Dai, Hui Liu, Jiliang Tang, Yi Chang

专题命中 跨模态检索 :multimodal(title,abstract)

AI总结 本文提出V-QPP-Bench,通过代理决策任务评估视觉查询预处理的改进,揭示视觉不完美对检索性能的严重影响及训练方法的优化潜力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.12883 2026-02-16 eess.IV cs.CV 74%

Dual-Phase Cross-Modal Contrastive Learning for CMR-Guided ECG Representations for Cardiovascular Disease Assessment

双相跨模态对比学习用于CMR引导的ECG表示以评估心血管疾病

Laura Alvarez-Florez, Angel Bujalance-Gomez, Femke Raijmakers, Samuel Ruiperez-Campillo, Maarten Z. H. Kolk, Jesse Wiers, Julia Vogt, Erik J. Bekkers, Ivana Išgum, Fleur V. Y. Tjong

机构 * Amsterdam University Medical Center(阿姆斯特丹大学医学中心) University of Amsterdam(阿姆斯特丹大学) ETH Zurich(苏黎世联邦理工学院) Department of Radiology and Nuclear Medicine, Amsterdam University Medical Center(放射医学与核医学系,阿姆斯特丹大学医学中心) Department of Radiology, Mayo Clinic, Rochester(放射医学系,梅奥诊所,罗切斯特)

专题命中 跨模态检索 :cross-modal(title);分类 cs.CV

AI总结 本文提出双相跨模态对比学习方法,通过结合ECG和CMR数据提升心血管疾病评估的ECG表示能力。

Comments Paper accepted at SPIE Medical Imaging 2026 Conference

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 多模态生成 7 篇

2505.01091 2026-02-16 cs.CV cs.AI 87%

Any-to-Any Vision-Language Model for Multimodal X-ray Imaging and Radiological Report Generation

任意到任意的视觉-语言模型用于多模态X射线成像与放射学报告生成

Daniele Molino, Francesco di Feola, Linlin Shen, Paolo Soda, Valerio Guarrasi

机构 * Department of Diagnostics and Intervention, Biomedical Engineering and Radiation Physics, Umeå University(诊断与介入部门,生物医学工程与放射物理,乌梅大学) College of Computer Science and Software Engineering, Shenzhen University(计算机科学与软件工程学院,深圳大学)

专题命中 多模态生成 :multimodal(title,abstract);any-to-any(title);分类 cs.CV、cs.AI

AI总结 本文提出了一种任意到任意的视觉-语言模型,用于多模态X射线成像和放射学报告生成,通过生成高质量图像和语义连贯的报告,提升了医疗领域生成模型的临床应用价值。

Comments arXiv admin note: substantial text overlap with arXiv:2501.04614

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.12205 2026-02-16 cs.CV cs.AI 84%

DeepGen 1.0: A Lightweight Unified Multimodal Model for Advancing Image Generation and Editing

DeepGen 1.0: 一种轻量级统一多模态模型,用于推进图像生成与编辑

Dianyi Wang, Ruihang Li, Feng Han, Chaofan Ma, Wei Song, Siyuan Wang, Yibin Wang, Yi Xin, Hongjian Liu, Zhixiong Zhang, Shengyuan Ding, Tianhang Wang, Zhenglin Cheng, Tao Lin, Cheng Jin, Kaicheng Yu, Jingjing Chen, Wenjie Wang, Zhongyu Wei, Jiaqi Wang

机构 * Shanghai Innovation Institute(上海创新研究院) Fudan University(复旦大学) University of Science and Technology of China(中国科学技术大学) Shanghai Jiao Tong University(上海交通大学) Zhejiang University(浙江大学) Westlake University(西湖大学) Nanjing University(南京大学) University of Southern California(南加州大学)

专题命中 多模态生成 :multimodal(title,abstract);image-text(abstract);分类 cs.CV、cs.AI

AI总结 DeepGen 1.0通过轻量级统一多模态模型在图像生成与编辑领域实现高性能,采用SCB框架和数据驱动训练策略,超越大参数模型表现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.15052 2026-02-16 cs.CL cs.AI 84%

SGM: Safety Glasses for Multimodal Large Language Models via Neuron-Level Detoxification

SGM: 通过神经层面的净化为多模态大语言模型提供安全眼镜

Hongbo Wang, MaungMaung AprilPyone, Isao Echizen

专题命中 多模态生成 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL、cs.AI

AI总结 SGM通过神经层面的净化技术,有效降低多模态大语言模型的毒性,同时保持生成流畅性和推理能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.04614 2026-02-16 cs.AI cs.LG 83%

XGeM: A Multi-Prompt Foundation Model for Multimodal Medical Data Generation

XGeM:一种多提示基础模型,用于多模态医学数据生成

Daniele Molino, Francesco Di Feola, Eliodoro Faiella, Deborah Fazzini, Domiziana Santucci, Linlin Shen, Valerio Guarrasi, Paolo Soda

机构 * Unit of Artificial Intelligence and Computer Systems, Department of Engineering, Università Campus Bio-Medico di Roma(人工智能与计算机系统单位,工程系,罗马生物医学大学) Department of Diagnostics and Intervention, Biomedical Engineering and Radiation Physics, Umeå University(诊断与介入部门,生物医学工程与辐射物理,乌梅大学) Department of Diagnostic Imaging and Stereotactic Radiosurgey, Centro Diagnostico Italiano S.p.A.(诊断影像与立体放射外科部门,意大利诊断中心股份有限公司) Department of Radiology and Interventional Radiology, Fondazione Policlinico Universitario Campus Bio-Medico(放射科与介入放射科,大学医学中心生物医学校园基金会) Research Unit of Radiology and Interventional Radiology, Department of Medicine and Surgery, Università Campus Bio-Medico di Roma(放射科与介入放射科研究单位,医学与外科系,罗马生物医学大学) College of Computer Science and Software Engineering, Shenzhen University(计算机科学与软件工程学院,深圳大学)

专题命中 多模态生成 :multimodal(title,abstract);any-to-any(abstract);分类 cs.AI

AI总结 XGeM是一种多模态生成模型,通过多提示训练策略实现灵活的医学数据合成,解决多模态数据生成中的临床一致性与数据稀缺问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.12987 2026-02-16 cs.HC 50%

GroundLink: Exploring How Contextual Meeting Snippets Can Close Common Ground Gaps in Editing 3D Scenes for Virtual Production

GroundLink:探索上下文会议片段如何在虚拟制作3D场景编辑中弥合常见共识差距

Gun Woo, Park, Frederik Brudy, George Fitzmaurice, Fraser Anderson

专题命中 多模态生成 :cross-modal(abstract)

AI总结 GroundLink通过在虚拟制作3D场景编辑中引入会议知识仪表板和跨模态同步功能,帮助专业人员更高效地建立团队共识。

Comments 27 pages, 13 figures, to appear at CHI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.12932 2026-02-16 stat.ML cs.LG 50%

TFTF: Training-Free Targeted Flow for Conditional Sampling

TFTF:无训练目标流用于条件采样

Qianqian Qu, Jun S. Liu

机构 * Zhili College, Tsinghua University, Beijing, China(清华大学紫荆学院) Department of Statistics and Data Science, Tsinghua University, Beijing, China(清华大学统计与数据科学系) Department of Statistics, Harvard University, Cambridge, MA, USA(哈佛大学统计系)

专题命中 多模态生成 :multimodal(abstract)

AI总结 TFTF通过引入随机流和重采样技术,实现无训练条件采样,提升高维和多模态场景下的生成效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01504 2026-02-16 cs.LG 50%

Instruction-based Time Series Editing

基于指令的时间序列编辑

Jiaxing Qiu, Dongliang Guo, Brynne Sullivan, Teague R. Henry, Thomas Hartvigsen

机构 * University of Virginia(弗吉尼亚大学)

专题命中 多模态生成 :multi-modal(abstract)

AI总结 InstructTime是一种基于指令的时间序列编辑器,通过自然语言指令实现灵活且可控的编辑,能够生成高质量的编辑结果并适应新条件。

Comments (KDD 26) Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.1

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 多模态评测 12 篇

2602.07621 2026-02-16 cs.CL 85%

SciClaimEval: Cross-modal Claim Verification in Scientific Papers

SciClaimEval: 科学论文中的跨模态声明验证

Xanh Ho, Yun-Ang Wu, Sunisth Kumar, Tian Cheng Xia, Florian Boudin, Andre Greiner-Petter, Akiko Aizawa

专题命中 多模态评测 :cross-modal(title,abstract);multimodal(abstract);multimodal foundation model(abstract);分类 cs.CL

AI总结 SciClaimEval是一个包含真实科学声明和跨模态证据的数据集,用于验证科学论文中的声明,通过专家标注和多模态模型测试,揭示了基于图表验证的挑战。

Comments Accepted at LREC 2026; 12 pages; data is available at https://sciclaimeval.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.12892 2026-02-16 cs.CV cs.AI cs.CL 85%

RADAR: Revealing Asymmetric Development of Abilities in MLLM Pre-training

RADAR: 揭示多模态大语言模型预训练中能力的非对称发展

Yunshuang Nie, Bingqian Lin, Minzhe Niu, Kun Xiang, Jianhua Han, Guowei Huang, Xingyue Quan, Hang Xu, Bokui Chen, Xiaodan Liang

机构 * Shenzhen Campus of Sun Yat-sen University(中山大学深圳校区) Peng Cheng Laboratory(鹏城实验室) Guangdong Key Laboratory of Big Data Analysis and Processing(广东大数据分析与处理重点实验室) Tsinghua Shenzhen International Graduate School(清华大学深圳国际 Graduate School) Tsinghua University(清华大学) Shanghai Jiao Tong University(上海交通大学) Yinwang Intelligent Technology Co., Ltd.(亿纬智能科技有限公司) Huawei’s 2012 Lab(华为2012实验室)

专题命中 多模态评测 :MLLM(title,abstract);multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 RADAR提出了一种高效的以能力为中心的评估框架,用于揭示多模态大语言模型预训练中感知和推理能力的非对称发展。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13028 2026-02-16 cs.CV cs.CL 84%

Human-Aligned MLLM Judges for Fine-Grained Image Editing Evaluation: A Benchmark, Framework, and Analysis

面向细粒度图像编辑评估的人类对齐MLLM评判:一个基准、框架和分析

Runzhou Liu, Hailey Weingord, Sejal Mittal, Prakhar Dungarwal, Anusha Nandula, Bo Ni, Samyadeep Basu, Hongjie Chen, Nesreen K. Ahmed, Li Li, Jiayi Zhang, Koustava Goswami, Subhojyoti Mukherjee, Branislav Kveton, Puneet Mathur, Franck Dernoncourt, Yue Zhao, Yu Wang, Ryan A. Rossi, Zhengzhong Tu, Hongru Du

机构 * University of Virginia(弗吉尼亚大学) Columbia University(哥伦比亚大学) Vanderbilt University(范德比大学) Adobe Research(Adobe研究) Dolby Laboratories(杜比实验室) Cisco Research(思科研究) University of Southern California(南加州大学) University of Wisconsin-Madison(威斯康星大学麦迪逊分校) University of Oregon(俄勒冈大学) Texas A&M University(德克萨斯大学)

专题命中 多模态评测 :MLLM(title,abstract);multimodal(abstract);分类 cs.CV、cs.CL

AI总结 本文提出细粒度MLLM评判框架,通过分解十二个可解释因素,提升图像编辑评估的精度与实用性,为研究和改进图像编辑方法提供实用基础。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08275 2026-02-16 cs.CL cs.AI 84%

MLLM-CTBench: A Benchmark for Continual Instruction Tuning with Reasoning Process Diagnosis

MLLM-CTBench: 一个用于持续指令微调的基准,包含推理过程诊断

Haiyun Guo, Zhiyan Hou, Yandu Sun, Jinghan He, Yu Chen, Yuzhe Zhou, Yuheng Jia, Jinqiao Wang, Tat-Seng Chua

机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Southeast University(东南大学) School of Computing, National University of Singapore(新加坡国立大学计算机学院)

专题命中 多模态评测 :MLLM(title,abstract);multimodal(abstract);分类 cs.CL、cs.AI

AI总结 MLLM-CTBench提出一个用于持续指令微调的基准,通过多维评估框架和强化微调方法,分析跨任务知识保留和灾难性遗忘问题。

Comments under review

详情

展开后加载摘要…

URL PDF HTML 收藏