Think-Clip-Sample: Slow-Fast Frame Selection for Video Understanding
Think-Clip-Sample: 慢-快帧选择用于视频理解
专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV、cs.AI
AI总结 Think-Clip-Sample通过多查询推理和慢-快采样提升长视频理解的效率和效果
Comments Accepted by ICASSP2026
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
Think-Clip-Sample: 慢-快帧选择用于视频理解
专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV、cs.AI
AI总结 Think-Clip-Sample通过多查询推理和慢-快采样提升长视频理解的效率和效果
Comments Accepted by ICASSP2026
多 caption:利用多语言视觉主张检测虚假信息
专题命中 视频多模态 :multimodal(abstract);分类 cs.CL
AI总结 MultiCaption通过多语言视觉主张数据集提升多模态虚假信息检测性能
面向大语言模型的可解释交通流预测
机构 * Hong Kong University of Science and Technology (Guangzhou)(香港理工大学(广州)) ; Johns Hopkins University(约翰霍普金斯大学) ; Department of Civil and System Engineering(土木与系统工程系)
专题命中 视频多模态 :multi-modal(abstract);分类 cs.AI
AI总结 本文提出xTP-LLM模型,利用大语言模型生成可解释的交通流预测,首次将LLM应用于交通预测的可解释性研究。
Comments 31pages, 16 figures
Journal ref Communications in Transportation Research, vol. 4, 100150, 2024
StellarF:一种结合物理知识的LoRA框架,用于利用历史和统计数据进行恒星耀斑预测
机构 * School of Computer Science, Nanjing University of Posts and Telecommunications(南京邮电大学计算机科学学院) ; Jiangsu Key Laboratory of Big Data Security and Intelligent Processing(江苏大数据安全与智能处理重点实验室) ; University of Chinese Academy of Sciences(中国科学院大学) ; CAS Key Laboratory of Optical Astronomy, National Astronomical Observatories(中国科学院国家天文台光学天文重点实验室) ; School of Astronomy and Space Science, University of Chinese Academy of Sciences(中国科学院大学天文与空间科学学院)
专题命中 视频多模态 :multimodal(abstract)
AI总结 StellarF通过结合物理知识和LoRA框架,利用历史和统计数据提升恒星耀斑预测的准确性与物理可解释性。
Comments 12 pages, 8 figures (5 main, 3 appendix), 7 tables (2 main, 5 appendix)