arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-01-09 至 2026-01-09 共收录 9 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 9 篇

2512.00949 2026-01-09 cs.LG cs.AI 83%

Multi-Modal AI for Remote Patient Monitoring in Cancer Care

多模态AI在癌症护理中的远程患者监测

Yansong Liu, Ronnie Stafford, Pramit Khetrapal, Huriye Kocadag, Graça Carvalho, Patricia de Winter, Maryam Imran, Amelia Snook, Adamos Hadjivasiliou, D. Vijay Anand, Weining Lin, John Kelly, Yukun Zhou, Ivana Drobnjak

机构 * University College London(伦敦大学学院) Ethera Health LTD(Ethera健康有限公司) Centro Algoritmi, Universidade do Minho(阿尔戈里米中心,明霍大学)

专题命中 视频多模态 :multi-modal(title,abstract);multimodal(abstract);分类 cs.AI

AI总结 本研究提出多模态AI框架用于癌症护理中的远程患者监测,通过整合多源数据预测不良事件风险,准确率达83.9%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.04709 2026-01-09 cs.AI 79%

Bridging Temporal and Textual Modalities: A Multimodal Framework for Automated Cloud Failure Root Cause Analysis

弥合时序与文本模态:一种多模态框架用于自动化云故障根本原因分析

Gijun Park

机构 * Okestro

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

AI总结 本文提出一种多模态框架,通过融合时间序列与文本数据,提升云故障根本原因分析的自动化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.00123 2026-01-09 cs.CV 79%

A Spatially Masked Adaptive Gated Network for multimodal post-flood water extent mapping using SAR and incomplete multispectral data

一种空间掩码自适应门控网络用于利用SAR和不完整多光谱数据的多模态洪水后水 extent 映射

Hyunho Lee, Wenwen Li

机构 * School of Geographical Sciences and Urban Planning, Arizona State University(地理科学与城市规划学院,亚利桑那州立大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出SMAGNet,一种多模态深度学习模型,通过融合SAR和不完整多光谱数据,提升洪水后水 extent 映射的准确性和鲁棒性。

Comments 50 pages, 12 figures, 6 tables

Journal ref ISPRS Journal of Photogrammetry and Remote Sensing, 232, 492-508, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.04299 2026-01-09 cs.LG q-bio.QM 78%

Transformer-Based Multi-Modal Temporal Embeddings for Explainable Metabolic Phenotyping in Type 1 Diabetes

基于Transformer的多模态时间嵌入用于1型糖尿病的可解释代谢表型分析

Pir Bakhsh Khokhar, Carmine Gravino, Fabio Palomba, Sule Yildrim Yayilgan, Sarang Shaikh

专题命中 视频多模态 :multi-modal(title);multimodal(abstract)

AI总结 本研究提出基于Transformer的多模态时间嵌入框架,用于1型糖尿病的可解释代谢表型分析,识别出5种代谢亚组并揭示其与心血管风险的关联。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19850 2026-01-09 cs.CV 74%

FALCONEye: Finding Answers and Localizing Content in ONE-hour-long videos with multi-modal LLMs

FALCONEye: 在一小时内视频中寻找答案并定位内容的多模态大语言模型

Carlos Plou, Cesar Borja, Ruben Martinez-Cantin, Ana C. Murillo

机构 * DIIS-I3A, University of Zaragoza(DIIS-I3A,西班牙阿利坎特大学)

专题命中 视频多模态 :multi-modal(title);分类 cs.CV

AI总结 FALCONEye通过多模态大语言模型在小时长视频中高效定位内容并回答问题,超越现有模型性能并显著降低推理成本。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.04778 2026-01-09 cs.CV cs.AI cs.CL cs.MM 70%

CounterVid: Counterfactual Video Generation for Mitigating Action and Temporal Hallucinations in Video-Language Models

CounterVid: 通过生成反事实视频缓解视频-语言模型中动作和时间幻觉

Tobia Poppi, Burak Uzkent, Amanmeet Garg, Lucas Porto, Garin Kessler, Yezhou Yang, Marcella Cornia, Lorenzo Baraldi, Rita Cucchiara, Florian Schiffers

机构 * University of Modena and Reggio Emilia(摩德纳和雷吉奥艾米利亚大学) University of Pisa(帕尔马大学) Amazon Prime Video(亚马逊Prime视频)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 CounterVid通过生成反事实视频缓解视频-语言模型中动作和时间幻觉问题,提出反事实视频生成框架和MixDPO方法,提升时间推理和动作识别性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11399 2026-01-09 cs.CL cs.CV 62%

Minimal Clips, Maximum Salience: Long Video Summarization via Key Moment Extraction

最短片段,最大显著性:通过关键时刻提取实现长视频摘要

Galann Pennec, Zhengyuan Liu, Nicholas Asher, Philippe Muller, Nancy F. Chen

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.CL

AI总结 本文提出了一种通过关键时刻提取实现长视频多模态摘要的方法,利用轻量级模型和大型语言模型实现高效摘要生成。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.04982 2026-01-09 cs.RO cs.AI 57%

When to Act: Calibrated Confidence for Reliable Human Intention Prediction in Assistive Robotics

何时行动:校准信心以在辅助机器人中实现可靠的意图预测

Johannes A. Gaus, Winfried Ilg, Daniel Haeufle

机构 * Hertie Institute for Clinical Brain Research & Center for Integrative Neuroscience, University of Tübingen(海德堡临床脑研究所及整合神经科学中心,图宾根大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.AI

AI总结 本文提出了一种基于校准概率的安全触发框架,用于提升辅助机器人中意图预测的可靠性,通过校准信心减少误校准并提高安全性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.04891 2026-01-09 cs.CV cs.LG 57%

Scaling Vision Language Models for Pharmaceutical Long Form Video Reasoning on Industrial GenAI Platform

在工业生成式AI平台上扩展视觉语言模型用于制药长格式视频推理

Suyash Mishra, Qiang Li, Srikanth Patil, Satyanarayan Pati, Baddu Narendra

机构 * Roche(罗氏) Accenture(埃森哲) Involead

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

AI总结 本文提出工业GenAI框架,解决制药领域长格式视频推理问题,通过多模态架构和实证分析提升效率并揭示现有VLMs的限制。

Comments Submitted to the Industry Track of Top Tier Conference; currently under peer review

详情

展开后加载摘要…

URL PDF HTML 收藏