机构
*
Shanghai Jiao Tong University(上海交通大学)
;
Eastern Institute of Technology(东部技术研究院)
;
School of Electronic Information and Electrical Engineering(电子信息与电气工程学院)
;
Li Auto(力汽车)
;
National University of Singapore(新加坡国立大学)
;
Tsinghua University(清华大学)
;
Ningbo Key Laboratory of Spatial Intelligence and Digital Derivative(宁波空间智能与数字衍生实验室)
;
Ningbo Institute of Digital Twin(宁波数字孪生研究院)
机构
*
University of California, Los Angeles(加州大学洛杉矶分校)
;
University of Pittsburgh(匹兹堡大学)
;
Fudan University(复旦大学)
;
University of California, Riverside(加州大学河滨分校)
;
Hong Kong University of Science(香港科学大学)
;
Maharishi International University(玛希拉国际大学)
Hi-GaTA: Hierarchical Gated Temporal Aggregation Adapter for Surgical Video Report Generation
Hi-GaTA:用于外科视频报告生成的分层门控时间聚合适配器
Kedi Sun, Chaohui Dang, Yue Feng, James Glasbey, Theodoros N. Arvanitis, Le Zhang
机构
*
School of Engineering, College of Engineering and Physical Sciences, University of Birmingham, Birmingham, UK(英国伯明翰大学工程学院)
;
School of Computer Science, University of Birmingham, Birmingham, UK(英国伯明翰大学计算机科学学院)
;
Department of Applied Health Sciences, University of Birmingham, Birmingham, UK(英国伯明翰大学应用健康科学系)
BARISTA: A Multi-Task Egocentric Benchmark for Compositional Visual Understanding
BARISTA:一种多任务第一人称视角基准,用于组合视觉理解
Patrick Knab, Orgest Xhelili, Inis Buzi, Drago Andres Guggiana Nilo, Mohd Saquib Khan, Lorenz Kolb, Manuel Scherzer, Kerem Yildirir, Christian Bartelt, Philipp Johannes Schubert
机构
*
Ramblr.ai Research(Ramblr.ai 研究院)
;
Technical University of Clausthal(Clausthal 技术大学)
机构
*
University of Electronic Science and Technology of China(电子科学与技术大学)
;
Peking University(北京大学)
;
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
Zhongguancun Academy(中关村学院)
机构
*
The University of Hong Kong(香港大学)
;
Fudan University(复旦大学)
;
Zhejiang University(浙江大学)
;
Hong Kong University of Science and Technology(香港科学与技术大学)
;
University of Sydney(悉尼大学)
;
Alibaba Group(阿里巴巴集团)
OmniVTG: A Large-Scale Dataset and Training Paradigm for Open-World Video Temporal Grounding
OmniVTG:一种大规模数据集和开放世界视频时间定位的训练范式
Minghang Zheng, Zihao Yin, Yi Yang, Yuxin Peng, Yang Liu
机构
*
Wangxuan Institute of Computer Technology, Peking University(北京大学王轩计算机技术研究所)
;
State Key Laboratory of General Artificial Intelligence, Peking University(北京大学通用人工智能国家重点实验室)
;
Central Media Technology Institute, Huawei Technologies Ltd.(华为技术有限公司中央媒体技术研究所)
;
PKU-WUHAN Institute for Artificial Intelligence, Peking University(北京大学武汉人工智能研究所)
机构
*
Hangzhou Institute for Advanced Study, UCAS(浙江大学杭州高等研究院)
;
Computer Network Information Center, CAS(中国科学院计算机网络信息中心)
;
Department of AI Infrastructure, Bilibili Inc.(B站人工智能基础设施部)
Eevee: Towards Close-up High-resolution Video-based Virtual Try-on
Eevee:迈向基于视频的高分辨率虚拟试衣
Jianhao Zeng, Yancheng Bai, Ruidong Chen, Xuanpu Zhang, Lei Sun, Dongyang Jin, Ryan Xu, Nannan Zhang, Dan Song, Xiangxiang Chu
机构
*
Amap, Alibaba Group(高德,阿里巴巴集团)
;
Tianjin University(天津大学)
;
Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(中国科学院深圳先进技术研究院)
ActivityForensics: A Comprehensive Benchmark for Localizing Manipulated Activity in Videos
活动取证:一种全面的基准,用于定位视频中的篡改活动
Peijun Bao, Anwei Luo, Gang Pan, Alex C. Kot, Xudong Jiang
机构
*
College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院)
;
School of Electrical and Electronic Engineering, Nanyang Technological University(南洋理工大学电气与电子工程学院)
;
School of Computing and Artificial Intelligence, Jiangxi University of Finance and Economics(江西财经大学计算机与人工智能学院)
;
Jiangxi Provincial Key Laboratory of Multimedia Intelligent Processing(江西省多媒体智能处理重点实验室)
;
Faculty of Engineering, Shenzhen MSU-BIT University(深圳北理莫斯科大学工程系)
;
VinUniversity
机构
*
Alaya Studio, Shanda AI Research Tokyo(Alaya工作室,盛大人工智能研究院东京)
;
National Taiwan University(国立台湾大学)
;
The University of Tokyo(东京大学)
;
National Yang Ming Chiao Tung University(国立阳明交通大学)