CoVStream: Edge-Cloud Collaboration for Understanding of Long Video Streams
CoVStream: 面向长视频流的边缘-云协作理解
Xu Liu, Guikun Chen, Zihao Yan, Kanzhi Wu, Wenguan Wang
机构
*
The State Key Lab of Brain-Machine Intelligence, Zhejiang University(浙江大学脑机智能国家重点实验室)
;
vivo Mobile Communication Co., Ltd., Shenzhen, China(深圳 vivo 通信有限公司)
Rethinking Video-Language Model from the Language Input Perspective
从语言输入角度重新思考视频-语言模型
Xiang Fang, Wanlong Fang, Changshuo Wang, Xiaoye Qu, Daizong Liu
机构
*
School of Software Engineering, Huazhong University of Science and Technology(华中科技大学软件学院)
;
Nanyang Technological University, Singapore(新加坡南洋理工大学)
;
University College London(伦敦大学学院)
;
Huazhong University of Science and Technology(华中科技大学)
;
Wuhan University(武汉大学)
Video-o3: Native Interleaved Clue Seeking for Long Video Multi-Hop Reasoning
Video-o3:长视频多跳推理的原生交错线索搜索
Xiangyu Zeng, Zhiqiu Zhang, Yuhan Zhu, Xinhao Li, Zikang Wang, Changlian Ma, Qingyu Zhang, Zizheng Huang, Kun Ouyang, Tianxiang Jiang, Ziang Yan, Yi Wang, Hongjie Zhang, Yali Wang, Limin Wang
机构
*
Nanjing University(南京大学)
;
Shanghai AI Laboratory(上海人工智能实验室)
;
Shanghai Jiao Tong University(上海交通大学)
;
Peking University(北京大学)
;
University of Science and Technology of China(中国科学技术大学)
;
Zhejiang University(浙江大学)
;
SIAT, Chinese Academy of Sciences(中国科学院软件研究所)
机构
*
Gansu Provincial Key Laboratory of Wearable Computing, School of Information Science and Engineering, Lanzhou University(甘肃省可穿戴计算重点实验室,兰州大学信息科学与工程学院)
;
Guangdong-Hong Kong-Macao Joint Laboratory for Emotional Intelligence and Pervasive Computing, Shenzhen MSU-BIT University(粤港澳大湾区情感智能与泛在计算联合实验室,深圳MSU-BIT大学)
;
Department of Computer and Information Engineering, Khalifa University(计算机与信息工程系,哈利法大学)
;
Artificial Intelligence Research Institute, Shenzhen MSU-BIT University(人工智能研究院,深圳MSU-BIT大学)
;
Department of Electrical and Computer Engineering, The University of Hong Kong(电子与计算机工程系,香港大学)
机构
*
Sun Yat-sen University(中山大学)
;
OPPO AI Center(OPPO人工智能中心)
;
Shenzhen Loop Area Institute(深圳河套学院)
;
Guangdong Key Laboratory of Big Data Analysis and Processing(广东大数据分析与处理重点实验室)
VideoStir: Understanding Long Videos via Spatio-Temporally Structured and Intent-Aware RAG
VideoStir:通过时空结构化和意图感知的RAG理解长视频
Honghao Fu, Miao Xu, Yiwei Wang, Dailing Zhang, Jun Liu, Yujun Cai
机构
*
University of Queensland(昆士兰大学)
;
University of California, Merced(加州大学默塞德分校)
;
Institute of Automation, CAS(中国科学院自动化研究所)
;
Lancaster University(兰卡斯特大学)
机构
*
University of California, Irvine(加州大学尔湾分校)
;
University of California, San Diego(加州大学圣地亚哥分校)
;
City University of Hong Kong(香港城市大学)
;
University of Pennsylvania(宾夕法尼亚大学)
;
Adobe Research(Adobe研究院)
机构
*
Bioengineering Department, Imperial College London, London, UK School of Biomedical Engineering \& lmaging Sciences, King's College London, London,UK
专题命中
长视频与时序推理
:video language model(title,abstract);分类 cs.CV
PAS3R: Pose-Adaptive Streaming 3D Reconstruction for Long Video Sequences
PAS3R:面向长视频序列的姿态自适应流式3D重建
Lanbo Xu, Liang Guo, Caigui Jiang, Cheng Wang
机构
*
State Key Laboratory of Human-Machine Hybrid Augmented Intelligence, Institute of Artificial Intelligence and Robotics, Xi'an Jiaotong University, China(人机混合增强智能国家重点实验室,人工智能与机器人研究院,西安交通大学,中国)
;
University of East Anglia, Norwich, NR47TJ, United Kingdom(东安格利亚大学,诺里奇,英国)
机构
*
National Key Laboratory of Human-Machine Hybrid Augmented Intelligence, Institute of Artificial Intelligence and Robotics, Xi’an Jiaotong University(国家人类-机器混合增强智能重点实验室,人工智能与机器人研究院,西安交通大学)
;
Institute of Artificial Intelligence (TeleAI), China Telecom(人工智能研究院(TeleAI),中国电信)
;
University of Science and Technology Beijing(北京科技大学)
FluxMem: Adaptive Hierarchical Memory for Streaming Video Understanding
FluxMem: 用于流视频理解的自适应分层内存
Yiweng Xie, Bo He, Junke Wang, Xiangyu Zheng, Ziyi Ye, Zuxuan Wu
机构
*
Institute of Trustworthy Embodied AI(可信具身人工智能研究所)
;
Fudan University(复旦大学)
;
Shanghai Innovation Institute(上海创新研究院)
;
Shanghai Key Laboratory of Multimodal Embodied AI(上海多模态具身人工智能重点实验室)
;
University of Maryland, College Park(马里兰大学 College Park 分校)
FreshMem: Brain-Inspired Frequency-Space Hybrid Memory for Streaming Video Understanding
FreshMem: 基于大脑的频率-空间混合内存用于流视频理解
Kangcong Li, Peng Ye, Lin Zhang, Chao Wang, Huafeng Qin, Tao Chen
机构
*
College of Future Information Technology, Fudan University, Shanghai, China(复旦大学未来信息科技学院)
;
Shanghai Innovation Institute, Shanghai, China(上海创新研究院)
;
Shanghai Artificial Intelligence Laboratory, Shanghai, China(上海人工智能实验室)
;
The Chinese University of Hong Kong, Hong Kong, China(香港中文大学)