arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-01-19 至 2026-01-19 共收录 4 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4 篇

2601.11359 2026-01-19 cs.CV cs.AI 62%

Think-Clip-Sample: Slow-Fast Frame Selection for Video Understanding

Think-Clip-Sample: 慢-快帧选择用于视频理解

Wenhui Tan, Ruihua Song, Jiaze Li, Jianzhong Ju, Zhenbo Luo

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

AI总结 Think-Clip-Sample通过多查询推理和慢-快采样提升长视频理解的效率和效果

Comments Accepted by ICASSP2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11220 2026-01-19 cs.CL 57%

MultiCaption: Detecting disinformation using multilingual visual claims

多 caption:利用多语言视觉主张检测虚假信息

Rafael Martins Frade, Rrubaa Panchendrarajan, Arkaitz Zubiaga

专题命中 视频多模态 :multimodal(abstract);分类 cs.CL

AI总结 MultiCaption通过多语言视觉主张数据集提升多模态虚假信息检测性能

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.02937 2026-01-19 cs.LG cs.AI 57%

Towards Explainable Traffic Flow Prediction with Large Language Models

面向大语言模型的可解释交通流预测

Xusen Guo, Qiming Zhang, Junyue Jiang, Mingxing Peng, Meixin Zhu, Hao, Yang

机构 * Hong Kong University of Science and Technology (Guangzhou)(香港理工大学(广州)) Johns Hopkins University(约翰霍普金斯大学) Department of Civil and System Engineering(土木与系统工程系)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.AI

AI总结 本文提出xTP-LLM模型,利用大语言模型生成可解释的交通流预测,首次将LLM应用于交通预测的可解释性研究。

Comments 31pages, 16 figures

Journal ref Communications in Transportation Research, vol. 4, 100150, 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.10986 2026-01-19 cs.LG 50%

StellarF: A Physics-Informed LoRA Framework for Stellar Flare Forecasting with Historical & Statistical Data

StellarF:一种结合物理知识的LoRA框架,用于利用历史和统计数据进行恒星耀斑预测

Tianyu Su, Zhiqiang Zou, Qingyu Lu, Feng Zhang, Ali Luo, Xiao Kong, Min Li

机构 * School of Computer Science, Nanjing University of Posts and Telecommunications(南京邮电大学计算机科学学院) Jiangsu Key Laboratory of Big Data Security and Intelligent Processing(江苏大数据安全与智能处理重点实验室) University of Chinese Academy of Sciences(中国科学院大学) CAS Key Laboratory of Optical Astronomy, National Astronomical Observatories(中国科学院国家天文台光学天文重点实验室) School of Astronomy and Space Science, University of Chinese Academy of Sciences(中国科学院大学天文与空间科学学院)

专题命中 视频多模态 :multimodal(abstract)

AI总结 StellarF通过结合物理知识和LoRA框架,利用历史和统计数据提升恒星耀斑预测的准确性与物理可解释性。

Comments 12 pages, 8 figures (5 main, 3 appendix), 7 tables (2 main, 5 appendix)

详情

展开后加载摘要…

URL PDF HTML 收藏