arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-02-09 至 2026-02-09 共收录 13 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 13 篇

2510.19585 2026-02-09 cs.CL cs.AI cs.CV cs.DL 82%

Detecting Latin in Historical Books with Large Language Models: A Multimodal Benchmark

利用大语言模型检测历史书籍中的拉丁文:一个多模态基准测试

Yu Wu, Ke Shu, Jonas Fischer, Lidia Pivovarova, David Rosson, Eetu Mäkelä, Mikko Tolonen

机构 * University of Helsinki(赫尔辛基大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 本文提出利用大语言模型检测历史文献中的拉丁文片段,并通过多模态数据集评估模型性能,揭示了零样本模型在拉丁文识别中的有效性及局限性。

Comments Accepted by the EACL 2026 main conference. Code and data available at https://github.com/COMHIS/EACL26-detect-latin

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.06743 2026-02-09 cs.CV 79%

Clinical-Prior Guided Multi-Modal Learning with Latent Attention Pooling for Gait-Based Scoliosis Screening

具有潜在注意力池的临床先验引导多模态学习用于基于步态的脊柱侧弯筛查

Dong Chen, Zizhuang Wei, Jialei Xu, Xinyang Sun, Zonglin He, Meiru An, Huili Peng, Yong Hu, Kenneth MC Cheung

机构 * Orthopaedic Centre, The University of Hong Kong - Shenzhen Hospital, Shenzhen, China(香港大学深圳医院骨科中心) Translational Medicine Centre, The University of Hong Kong - Shenzhen Hospital, China(香港大学深圳医院转化医学中心) Department of Orthopaedics and Traumatology, Li Ka Shing Faculty of Medicine, The University of Hong Kong, Hong Kong, China(香港大学李嘉成医学院骨科及创伤外科学院) Huawei, China(华为)

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV

AI总结 本文提出了一种基于步态的脊柱侧弯筛查方法,通过多模态学习和潜在注意力池化机制,结合临床先验知识,实现了可解释的特征表示和高精度的筛查性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.06158 2026-02-09 cs.CV 79%

MGP-KAD: Multimodal Geometric Priors and Kolmogorov-Arnold Decoder for Single-View 3D Reconstruction in Complex Scenes

MGP-KAD:多模态几何先验与科莫戈罗夫-阿诺德解码器用于复杂场景中的单视图3D重建

Luoxi Zhang, Chun Xie, Itaru Kitahara

机构 * Doctoral Program in Empowerment Informatics, University of Tsukuba, Japan(赋能信息学博士项目,茨ukuba大学,日本) Center for Computational Science, Tsukuba, Ibaraki, Japan(计算科学中心,茨ukuba,Ibaraki,日本)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

AI总结 MGP-KAD通过整合多模态特征与几何先验,结合科莫戈罗夫-阿诺德网络解码器,提升复杂场景中单视图3D重建的精度与细节保持能力。

Comments 6 pages. Published in IEEE International Conference on Image Processing (ICIP) 2025

Journal ref Proc. IEEE International Conference on Image Processing (ICIP), 2025, pp. 1564-1569

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21355 2026-02-09 cs.LG 78%

SMMILE: An Expert-Driven Benchmark for Multimodal Medical In-Context Learning

SMMILE:一个多模态医学上下文学习的专家驱动基准

Melanie Rieff, Maya Varma, Ossian Rabow, Subathra Adithan, Julie Kim, Ken Chang, Hannah Lee, Nidhi Rohatgi, Christian Bluethgen, Mohamed S. Muneer, Jean-Benoit Delbrouck, Michael Moor

机构 * ETH Zurich(苏黎世联邦理工学院) Stanford University(斯坦福大学) Lund University(隆德大学) Jawaharlal Institute of Postgraduate Medical Education and Research(贾瓦哈拉尔·尼赫鲁医学教育与研究学院) UCSF(加州大学旧金山分校) University of Zurich(苏黎世大学) University Hospital Zurich(苏黎世大学医院) HOPPR

专题命中 多模态评测 :multimodal(title,abstract)

AI总结 SMMILE是一个专家驱动的多模态医学上下文学习基准,评估了15个MLLMs在医学任务中的多模态ICL能力,发现其表现有限且易受无关示例和最近性偏见影响。

Comments NeurIPS 2025 (Datasets & Benchmarks Track)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.06369 2026-02-09 cs.CV cs.AI 62%

Revisiting Salient Object Detection from an Observer-Centric Perspective

重新审视以观察者为中心的显著物体检测

Fuxi Zhang, Yifan Wang, Hengrun Zhao, Zhuohan Sun, Changxing Xia, Lijun Wang, Huchuan Lu, Yangrui Shao, Chen Yang, Long Teng

机构 * Dalian University of Technology(大连理工大学)

专题命中 多模态评测 :multi-modal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出OC-SOD方法,通过考虑观察者偏好和意图,实现个性化显著性预测,并构建首个OC-SOD数据集和代理基线。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.04692 2026-02-09 cs.CV cs.AI 62%

DRMOT: A Dataset and Framework for RGBD Referring Multi-Object Tracking

DRMOT:用于RGBD参照多目标跟踪的数据集和框架

Sijia Chen, Lijuan Ma, Yanqiu Yu, En Yu, Liman Liu, Wenbing Tao

专题命中 多模态评测 :MLLM(abstract);分类 cs.CV、cs.AI

AI总结 本文提出DRMOT任务,通过融合RGB、深度和语言信息,构建DRSet数据集并提出DRTrack框架,提升多目标跟踪的3D感知能力。

Comments https://github.com/chen-si-jia/DRMOT

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.06203 2026-02-09 cs.CV cs.AI cs.LG cs.RO 62%

AnyThermal: Towards Learning Universal Representations for Thermal Perception

AnyThermal:迈向热感知的通用表示学习

Parv Maheshwari, Jay Karhade, Yogesh Chawla, Isaiah Adu, Florian Heisen, Andrew Porco, Andrew Jong, Yifei Liu, Santosh Pitla, Sebastian Scherer, Wenshan Wang

机构 * Robotics Institute, Carnegie Mellon University(卡内基梅隆大学机器人研究所) Biological Systems Engineering, University of Nebraska-Lincoln(内布拉斯加大学林肯分校生物系统工程系) Mechanical Engineering, Penn State University(宾夕法尼亚州立大学机械工程系) School of Engineering and Design, Technical University of Munich(慕尼黑技术大学工程与设计学院) Mechanical Engineering, Carnegie Mellon University(卡内基梅隆大学机械工程系)

专题命中 多模态评测 :cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 AnyThermal通过蒸馏视觉基础模型的特征表示,实现跨环境和任务的通用热感知表示学习,取得显著性能提升。

Comments Accepted at IEEE ICRA (International Conference on Robotics & Automation) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.06363 2026-02-09 cs.CV 57%

Robust Pedestrian Detection with Uncertain Modality

具有不确定模态的鲁棒行人检测

Qian Bie, Xiao Wang, Bin Yang, Zhixi Yu, Jun Chen, Xin Xu

专题命中 多模态评测 :cross-modal(abstract);分类 cs.CV

AI总结 本文提出AUNet,通过自适应不确定性感知网络在不确定输入下实现鲁棒行人检测,结合UMVR和MAI模块提升跨模态信息融合效果。

Comments Due to the limitation "The abstract field cannot be longer than 1,920 characters", the abstract here is shorter than that in the PDF file

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.00734 2026-02-09 eess.SP cs.AI cs.LG 57%

EEG-MACS: Manifold Attention and Confidence Stratification for EEG-based Cross-Center Brain Disease Diagnosis under Unreliable Annotations

EEG-MACS: 用于基于EEG的跨中心脑部疾病诊断的流形注意力与置信度分层

Zhenxi Song, Ruihan Qin, Huixia Ren, Zhen Liang, Yi Guo, Min Zhang, Zhiguo Zhang

机构 * Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳)) Shenzhen People's Hospital(深圳人民医院) Shenzhen University(深圳大学) Shenzhen Bay Laboratory(深圳湾实验室)

专题命中 多模态评测 :multimodal(abstract);分类 cs.AI

AI总结 EEG-MACS通过流形注意力与置信度分层技术,提升跨中心脑部疾病诊断的准确性与鲁棒性。

Comments 15 pages, 9 figures. Oral presentation at ACM MM 2024

Journal ref In Proceedings of the 32nd ACM International Conference on Multimedia (ACM MM '24), pp. 340-349, 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.06904 2026-02-09 astro-ph.CO 50%

Non-spherical BUFFALOs: a weak lensing view of the Frontier Field clusters and associated systematics

非球形 BUFFALOs:从弱透镜视图看前沿场簇及相关系统误差

A. Niemiec, A. Acebron, B. Beauchesne, M. Jauzac, J. M. Diego, D. Eckert, D. Harvey, A. M. Koekemoer, D. J. Lagattuta, M. Limousin, G. Mahler, N. Patel, S. Tam, J. F. V. Allingham, R. Cen, A. Faisst, D. Perera, M. Sereno

专题命中 多模态评测 :multi-modal(abstract)

AI总结 本研究通过高分辨率成像和光谱学分析,揭示了前沿场星系团中弱透镜测量的系统误差来源,特别是复杂结构对质量估计的影响。

Comments 19 pages, 12 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.06631 2026-02-09 cs.CY 50%

Estimating Exam Item Difficulty with LLMs: A Benchmark on Brazil's ENEM Corpus

利用LLM估算考试题目难度:在巴西ENEM语料库上的基准测试

Thiago Brant, Julien Kühn, Jun Pang

专题命中 多模态评测 :multimodal(abstract)

AI总结 本文研究了LLM在估算考试题目难度上的表现,发现其在单模态问题上表现中等,但对多模态问题和人口统计数据提示的适应性有限,建议采用评估后再生成的流程进行负责任的评估设计。

Comments 42 pages, 19 figures. Appendix included

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.06451 2026-02-09 cs.LG 50%

BrokenBind: Universal Modality Exploration beyond Dataset Boundaries

BrokenBind: 超越数据集边界的通用模态探索

Zhuo Huang, Runnan Chen, Bo Han, Gang Niu, Masashi Sugiyama, Tongliang Liu

机构 * Sydney AI Centre, The University of Sydney(悉尼人工智能中心,悉尼大学) Hong Kong Baptist University(香港 Baptist 大学) RIKEN Center for Advanced Intelligence Project(理化学研究所先进智能项目中心) The University of Tokyo(东京大学)

专题命中 多模态评测 :multi-modal(abstract)

AI总结 BrokenBind提出了一种超越数据集边界的通用模态探索方法,通过结合不同数据集的模态实现灵活且通用的多模态学习。

Comments 17 pages, 8 figures, and 10 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08076 2026-02-09 cs.DB 50%

Heterogeneity in Entity Matching: A Survey and Experimental Analysis

实体匹配中的异质性:调查与实验分析

Mohammad Hossein Moslemi, Amir Mousavi, Behshid Behkamal, Mostafa Milani

专题命中 多模态评测 :multimodal(abstract)

AI总结 本文调查了实体匹配中的异质性问题,分析了不同异质性类型对匹配任务的影响,并评估了现有方法的局限性,提出了多模态匹配、人机协作等未来研究方向。

Comments Accepted at Data & Knowledge Engineering (DKE)

详情

展开后加载摘要…

URL PDF HTML 收藏