arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-02-11 至 2026-02-11 共收录 9 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 9 篇

2512.19379 2026-02-11 cs.LG cs.AI cs.MM 84%

OmniMER: Auxiliary-Enhanced LLM Adaptation for Indonesian Multimodal Emotion Recognition

OmniMER: 增辅增强的LLM适应用于印度尼西亚多模态情感识别

Xueming Yan, Boyan Xu, Yaochu Jin, Lixian Xiao, Wenlong Ye, Runyang Cai, Zeqi Zheng, Jingfa Liu, Aimin Yang, Yongduan Song

机构 * School of Information Science and Technology, Guangdong University of Foreign Studies(广东外语外贸大学信息科学与技术学院) School of Computer Science, Guangdong University of Technology(广东工业大学计算机学院) Faculty of Asian Languages and Cultures, Guangdong University of Foreign Studies(广东外语外贸大学亚洲语言文化学院) School of Engineering, Westlake University(西湖大学工程学院) School of Computer Science and Intelligence Education, Lingnan Normal University(岭南师范学院计算机科学与智能教育学院) School of Automation, Chongqing University(重庆大学自动化学院)

专题命中 多模态评测 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI、cs.MM

AI总结 OmniMER通过三种辅助任务提升印度尼西亚多模态情感识别性能,实现情感分类和识别的显著提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09315 2026-02-11 cs.CV cs.AI 81%

A Deep Multi-Modal Method for Patient Wound Healing Assessment

一种用于患者伤口愈合评估的深度多模态方法

Subba Reddy Oota, Vijay Rowtula, Shahid Mohammed, Jeffrey Galitz, Minghsun Liu, Manish Gupta

机构 * Woundtech Innovative Healthcare Solutions(Woundtech创新医疗解决方案) Microsoft AI Research(微软人工智能研究院)

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV、cs.AI

AI总结 本文提出了一种深度多模态方法,通过整合伤口变量和图像数据,预测患者住院风险,以提高伤口诊断效率和早期发现愈合问题。

Comments 4 pages, 2 figures

Journal ref Medical Imaging Meets NeurIPS Workshop, 33rd Conference on Neural Information Processing Systems (NeurIPS 2019), Vancouver, Canada

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11362 2026-02-11 cs.LG cs.CV 79%

PersonaX: Multimodal Datasets with LLM-Inferred Behavior Traits

PersonaX: 多模态数据集与LLM推断行为特征

Loka Li, Wong Yu Kang, Minghao Fu, Guangyi Chen, Zhenhao Chen, Gongxu Luo, Yuewen Sun, Salman Khan, Peter Spirtes, Kun Zhang

机构 * Mohamed bin Zayed University of Artificial Intelligence(莫扎德·本·扎耶德人工智能大学) Carnegie Mellon University(卡内基梅隆大学) University of California San Diego(加州大学圣地亚哥分校) Australian National University(澳大利亚国立大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

AI总结 PersonaX通过多模态数据集结合LLM推断行为特征,推动多模态特征分析与因果推理发展。

Comments ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09294 2026-02-11 cs.CE 67%

BrainTAP: Brain Disorder Prediction with Adaptive Distill and Selective Prior Integration

BrainTAP: 基于自适应蒸馏和选择性先验整合的脑部疾病预测

Zhenyu Lei, Aiying Zhang, Song Wang, Han Fan, Jundong Li

专题命中 多模态评测 :multimodal(abstract);cross-modal(abstract)

AI总结 BrainTAP通过自适应蒸馏和选择性先验整合,提升脑部疾病预测性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09214 2026-02-11 cs.CV 57%

VLM-UQBench: A Benchmark for Modality-Specific and Cross-Modality Uncertainties in Vision Language Models

VLM-UQBench:一种用于视觉语言模型中模态特定和跨模态不确定性的基准

Chenyu Wang, Tianle Chen, H. M. Sabbir Ahmad, Kayhan Batmanghelich, Wenchao Li

机构 * Boston University(波士顿大学)

专题命中 多模态评测 :cross-modal(abstract);分类 cs.CV

AI总结 VLM-UQBench通过评估不同UQ方法在模态特定和跨模态不确定性上的表现,揭示了现有方法在细粒度不确定性检测上的不足。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09732 2026-02-11 physics.optics 50%

The mixture of glycerin with tartrazine: a solution to reversibly increase tissue transparency for in vitro quantitative phase imaging

甘油与焦糖色素的混合物:一种可逆性增加组织透明度用于体外定量相位成像的解决方案

Mikolaj Krysa, Anna Chwastowicz, Malgorzata Lenarcik, Pawel Matrybak, Piotr Zdankowski, Maciej Trusiak

专题命中 多模态评测 :multimodal(abstract)

AI总结 本研究提出了一种低成本且安全的甘油与焦糖色素混合物,用于提高组织透明度,从而实现高通量无标记定量相位成像。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09540 2026-02-11 cs.SE 50%

SWE-Bench Mobile: Can Large Language Model Agents Develop Industry-Level Mobile Applications?

SWE-Bench Mobile: 大语言模型代理能否开发行业级移动应用?

Muxin Tian, Zhe Wang, Blair Yang, Zhenwei Tang, Kunlun Zhu, Honghua Dong, Hanchen Li, Xinni Xie, Guangjing Wang, Jiaxuan You

专题命中 多模态评测 :multi-modal(abstract)

AI总结 SWE-Bench Mobile评估大语言模型代理在开发行业级移动应用中的能力,发现商业代理表现优于开源替代品,且简单提示策略更有效。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17400 2026-02-11 stat.ME 50%

Variational autoencoder for inference of nonlinear mixed effect models based on ordinary differential equations

基于常微分方程的非线性混合效应模型推断的变分自编码器

Zhe Li, Mélanie Prague, Rodolphe Thiébaut, Quentin Clairon

专题命中 多模态评测 :multimodal(abstract)

AI总结 本文提出基于常微分方程的非线性混合效应模型参数估计方法,利用变分自编码器替代传统MCMC方法,通过最大化ELBO提升参数估计效率与不确定性量化。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.02883 2026-02-11 q-bio.BM cs.LG 50%

DISPROTBENCH: Uncovering the Functional Limits of Protein Structure Prediction Models in Intrinsically Disordered Regions

DISPROTBENCH: 揭示蛋白质结构预测模型在内源性无序区域的功能限制

Xinyue Zeng, Tuo Wang, Adithya Kulkarni, Alexander Lu, Alexandra Ni, Phoebe Xing, Junhan Zhao, Siwei Chen, Dawei Zhou

机构 * Institute for Clarity in Documentation(清晰文档研究所) Inria Paris-Rocquencourt(巴黎-罗克琴堡研究所) Rajiv Gandhi University(拉贾·甘地大学) Tsinghua University(清华大学) Palmer Research Laboratories(帕勒研究实验室)

专题命中 多模态评测 :multi-modal(abstract)

AI总结 DISPROTBENCH通过引入功能不确定性敏感度度量,揭示蛋白质结构预测模型在内源性无序区域的功能限制及预测不确定性对下游任务的影响。

详情

展开后加载摘要…

URL PDF HTML 收藏