arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-12-24 至 2025-12-24 共收录 3 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 3 篇

2512.17601 2025-12-24 cs.CV 83%

HeadHunt-VAD: Hunting Robust Anomaly-Sensitive Heads in MLLM for Tuning-Free Video Anomaly Detection

HeadHunt-VAD: 在MLLM中寻找鲁棒的异常敏感头部以实现无调优视频异常检测

Zhaolin Cai, Fan Li, Ziwei Zheng, Haixia Bi, Lijun He

专题命中 视频多模态 :MLLM(title,abstract);multimodal(abstract);分类 cs.CV

AI总结 HeadHunt-VAD通过直接在冻结的MLLM中寻找鲁棒的异常敏感头部,实现高效的无调优视频异常检测。

Comments AAAI 2026 Oral

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.20501 2025-12-24 cs.CV 79%

Bridging Modalities and Transferring Knowledge: Enhanced Multimodal Understanding and Recognition

弥合模态与知识转移:增强多模态理解和识别

Gorjan Radevski

机构 * University of Bristol(布里斯托大学) University of Würzburg(乌尔姆大学) Processing Speech and Images (PSI)(语音与图像处理组)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出多模态对齐、翻译、融合和转移方法,提升复杂输入的理解与识别能力,涵盖空间语言、医学文本、知识图谱和动作识别等多个领域。

Comments Ph.D. manuscript; Supervisors/Mentors: Marie-Francine Moens and Tinne Tuytelaars

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20538 2025-12-24 cs.RO 78%

Uni-Mapper: Unified Mapping Framework for Multi-modal LiDARs in Complex and Dynamic Environments

Uni-Mapper:多模态激光雷达在复杂动态环境中的统一映射框架

Gilhwan Kang, Hogyun Kim, Byunghee Choi, Seokhwan Jeong, Young-Sik Shin, Younggun Cho

机构 * Hyundai Motor Company(现代汽车公司) Inha University(inha大学) Korea Institute of Machinery and Materials(韩国机械材料研究院)

专题命中 视频多模态 :multi-modal(title,abstract)

AI总结 Uni-Mapper通过动态感知和多模态激光雷达融合技术,实现复杂动态环境下的统一地图构建与回环检测。

Comments 18 pages, 14 figures

Journal ref 2025

详情

展开后加载摘要…

URL PDF HTML 收藏