arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-12-17 至 2025-12-17 共收录 4 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4 篇

2503.09143 2025-12-17 cs.CV 83%

Exo2Ego: Exocentric Knowledge Guided MLLM for Egocentric Video Understanding

Exo2Ego:基于外部知识引导的多模态大语言模型用于第一人称视频理解

Haoyu Zhang, Qiaohui Chu, Meng Liu, Haoxiang Shi, Yaowei Wang, Liqiang Nie

专题命中 视频多模态 :MLLM(title,abstract);multimodal(abstract);分类 cs.CV

AI总结 Exo2Ego通过迁移学习提升内向视频理解能力,利用外向知识增强模型性能。

Comments This paper is accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.14058 2025-12-17 cs.CV cs.AI 81%

Real-time prediction of workplane illuminance distribution for daylight-linked controls using non-intrusive multimodal deep learning

基于非侵入式多模态深度学习的实时工作平面照度分布预测用于日光联动控制

Zulin Zhuang, Yu Bian

机构 * School of Architecture(建筑学院) State Key Laboratory of Subtropical Building(亚热带建筑科学国家重点实验室) South China University of Technology(华南理工大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 本研究提出了一种基于非侵入式多模态深度学习的实时工作平面照度分布预测方法,用于提高日光联动控制的能效

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.10021 2025-12-17 cs.LG cs.AI 79%

Online Multi-modal Root Cause Identification in Microservice Systems

微服务系统中的在线多模态根本原因识别

Lecheng Zheng, Zhengzhang Chen, Haifeng Chen

机构 * University of Illinois Urbana-Champagin(伊利诺伊大学厄巴纳-香槟分校) NEC Labs America(NEC美洲实验室)

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.AI

AI总结 本文提出OCEAN方法,通过多模态因果结构学习实现微服务系统中的在线根本原因识别,结合扩张卷积神经网络和图神经网络,提升因果图学习的准确性与效率。

Comments Accepted by BigData 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.14017 2025-12-17 cs.CV cs.AI 62%

KFS-Bench: Comprehensive Evaluation of Key Frame Sampling in Long Video Understanding

KFS-Bench: 长视频理解中关键帧采样的全面评估

Zongyao Li, Kengo Ishida, Satoshi Yamazaki, Xiaotong Ji, Jianquan Liu

机构 * Visual Intelligence Research Laboratories, NEC Corporation(NEC公司视觉智能研究实验室)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 KFS-Bench通过多场景标注评估关键帧采样策略,提出新度量标准和方法提升问答性能。

Comments WACV2026

详情

展开后加载摘要…

URL PDF HTML 收藏