StreamMem: Query-Agnostic KV Cache Memory for Streaming Video Understanding
机构 * Meta AI ; New York University(纽约大学)
专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI
Comments 15 pages, 3 figures
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
机构 * Meta AI ; New York University(纽约大学)
专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI
Comments 15 pages, 3 figures
机构 * Yale University(耶鲁大学)
专题命中 视频多模态 :multimodal(abstract);分类 cs.CL、cs.MM
Comments 12 pages, 4 figures, 2 tables. Extends our earlier framework on hierarchical narrative graphs with a semantic normalization module
机构 * Singapore University of Technology and Design(新加坡科技设计大学)
专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI
Comments 6 pages, 3 figures
机构 * Zhipu AI(智谱AI) ; Tsinghua University(清华大学)
专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI
机构 * School of Automotive Studies, Tongji University(同济大学汽车学院) ; College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院) ; School of Automobile, Chang'an University(长安大学汽车学院) ; School of Information and Electrical Engineering, Hangzhou City University(杭州城市学院信息与电气工程学院) ; College of Science and Engineering, James Cook University(詹姆斯库克大学科学与工程学院)
专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV、cs.AI
专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI
机构 * University Of Reading(阅读大学)
专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI
Comments 35 pages, 18 figures, Manuscript submitted to ACM
机构 * Universitat de Barcelona(巴塞罗那大学)
专题命中 视频多模态 :multimodal(abstract);分类 cs.CL、cs.AI
Comments 55 pages, 16 figures, 3 tables
机构 * Center for Research in Computer Vision, University of Central Florida(计算机视觉研究中心,中央佛罗里达大学) ; Center for Vision Technology, SRI International(视觉技术中心,SRI国际)
专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI
Comments Accepted to ICCV 2025 main conference
机构 * S-Lab, Nanyang Technological University(南洋理工大学S实验室) ; BUPT(北京邮电大学) ; ByteDance(字节跳动)
专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.CL
Comments Project page: https://llava-vl.github.io/blog/2024-09-30-llava-video/; Accepted at TMLR
机构 * UCLA(加州大学洛杉矶分校) ; HKU(香港大学) ; ISTA(因斯布鲁克大学)
专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV、cs.AI
机构 * Wangxuan Institute of Computer Technology, Peking University, Beijing, China(王轩计算机技术研究所,北京大学,北京,中国) ; State Key Laboratory of General Artificial Intelligence, Peking Universitys, Beijing, China(通用人工智能国家重点实验室,北京大学,北京,中国)
专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV、cs.MM
Journal ref ICME 2025
机构 * Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳)) ; Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院) ; Research Center of Artificial Intelligence, Peng Cheng Laboratory(鹏城实验室人工智能研究中心) ; The Hong Kong University of Science and Technology(香港科学与技术大学)
专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV、cs.MM
Comments Accepted by ICCV'25. 13 pages, 6 figures, 4 tables
专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV、cs.AI
机构 * Department of Machine Learning(机器学习系) ; Mohamed bin Zayed University of Artificial Intelligence(Mohamed bin Zayed人工智能大学)
专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI
专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI
机构 * Department of Artificial Intelligence, Korea University(人工智能系,韩国大学)
专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI
Comments SMC 2025
机构 * University of Arizona(亚利桑那大学) ; Chinese Academy of Sciences(中国科学院)
专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI
Comments 21 pages, 6 figures. Submitted to ACM TKDD
专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.CL
机构 * Korea Advanced Institute of Science and Technology (KAIST)(韩国科学技术院) ; University of Illinois at Urbana-Champaign (UIUC)(伊利诺伊大学厄巴纳-香槟分校)
专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.CL
Comments ICLR 2025
机构 * Massachusetts Institute of Technology(麻省理工学院) ; Amazon Web Services(亚马逊网络服务)
专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV、cs.AI
Comments 8 pages, 5 figures, accepted to the 11th IEEE International Workshop on Computer Vision in Sports (CVSports) at CVPR 2025; supplementary appendix included
机构 * Huazhong University of Science and Technology(华中科技大学) ; La Trobe University(拉特罗布大学)
专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.MM
机构 * Zhejiang University(浙江大学)
专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI
机构 * GVC Lab, Great Bay University(Great Bay大学GVC实验室)
专题命中 视频多模态 :MLLM(abstract);分类 cs.CV、cs.MM
Comments Project Page: https://jayleejia.github.io/FairyGen/ ; Code: https://github.com/GVCLab/FairyGen
专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI
机构 * University of Maryland, College Park(马里兰大学学院市分校) ; Meta Reality Labs(Meta现实实验室) ; Worcester Polytechnic Institute(沃斯特理工学院) ; University of Toronto(多伦多大学) ; FAIR, Meta AI(Meta AI)
专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV、cs.AI
Comments Accepted at ICCV 2025
机构 * Griffith University(格里菲斯大学) ; Data61/CSIRO(Data61/澳大利亚联邦科学与工业研究组织) ; University of New South Wales(新南威尔士大学)
专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI
Comments Accepted for publication in International Journal of Computer Vision (IJCV)
机构 * Google(谷歌) ; Stanford University(斯坦福大学) ; University of California, Berkeley(加州大学伯克利分校)
专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV、cs.AI
Journal ref Workshop on Deepfake Detection, Localization and Interpretability @ IJCAI 2025
机构 * National-Regional Key Technology Engineering Laboratory for Medical Ultrasound, School of Biomedical Engineering, Shenzhen University Medical School, Shenzhen University, Shenzhen, Guangdong, China(国家级医学超声关键技术研发实验室、生物医学工程学院、深圳大学医学院、深圳大学、深圳、广东、中国) ; Medical UltraSound Image Computing (MUSIC) Lab, Shenzhen University, Shenzhen, Guangdong, China(医学超声图像计算(MUSIC)实验室、深圳大学、深圳、广东、中国) ; Shenzhen RayShape Medical Technology Inc.(深圳RayShape医疗科技有限公司) ; Cancer Center, Department of Ultrasound Medicine, Zhejiang Provincial People’s Hospital, Affiliated People’s Hospital of Hangzhou Medical College, Hangzhou, Zhejiang, China(肿瘤中心、超声医学科、浙江省人民医院、杭州医学院附属人民医院、杭州、浙江、中国) ; Department of Health Management Center, Qilu Hospital, Cheeloo College of Medicine, Shandong University, Jinan, Shandong, China(健康管理中心、齐鲁医院、山东大学齐鲁医学院、济南、山东、中国)
专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV、cs.AI
专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV、cs.AI