OmniVTG: A Large-Scale Dataset and Training Paradigm for Open-World Video Temporal Grounding
OmniVTG:一种大规模数据集和开放世界视频时间定位的训练范式
Minghang Zheng, Zihao Yin, Yi Yang, Yuxin Peng, Yang Liu
机构
*
Wangxuan Institute of Computer Technology, Peking University(北京大学王轩计算机技术研究所)
;
State Key Laboratory of General Artificial Intelligence, Peking University(北京大学通用人工智能国家重点实验室)
;
Central Media Technology Institute, Huawei Technologies Ltd.(华为技术有限公司中央媒体技术研究所)
;
PKU-WUHAN Institute for Artificial Intelligence, Peking University(北京大学武汉人工智能研究所)
机构
*
Sun Yat-sen University(中山大学)
;
OPPO AI Center(OPPO人工智能中心)
;
Shenzhen Loop Area Institute(深圳河套学院)
;
Guangdong Key Laboratory of Big Data Analysis and Processing(广东大数据分析与处理重点实验室)
机构
*
AMAP, Alibaba Group(阿里集团AMAP)
;
University of California at Merced(加州大学默塞德分校)
;
University of Queensland(昆士兰大学)
;
Case Western Reserve University(凯斯西储大学)
RANGER: A Monocular Zero-Shot Semantic Navigation Framework through Visual Contextual Adaptation
RANGER:通过视觉上下文适应的单目零样本语义导航框架
Ming-Ming Yu, Yi Chen, Börje F. Karlsson, Wenjun Wu
机构
*
Beihang University(北京航空航天大学)
;
Beijing Academy of Artificial Intelligence(北京人工智能研究院)
;
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
Hangzhou International Innovation Institute, Beihang University(北京航空航天大学杭州国际创新研究院)
专题命中
视频多模态
:multimodal(abstract);multimodal foundation model(abstract);分类 cs.CV
Attend Before Attention: Efficient and Scalable Video Understanding via Autoregressive Gazing
先关注再注意:通过自回归注视实现高效的可扩展视频理解
Baifeng Shi, Stephanie Fu, Long Lian, Hanrong Ye, David Eigen, Aaron Reite, Boyi Li, Jan Kautz, Song Han, David M. Chan, Pavlo Molchanov, Trevor Darrell, Hongxu Yin
机构
*
Carnegie Mellon University(卡内基梅隆大学)
;
Kyung Hee University(Kyung Hee大学)
;
Korea Advanced Institute of Science and Technology(韩国科学技术院)
;
University of Pittsburgh(匹兹堡大学)
CommentsThis version corrects the author affiliation to reflect the accurate institutional information at the time of publication. No technical content of the paper has been changed
EventFlash: Towards Efficient MLLMs for Event-Based Vision
EventFlash: 向基于事件的视觉高效MLLMs迈进
Shaoyu Liu, Jianing Li, Guanghui Zhao, Yunjian Zhang, Wen Jiang, Ming Li, Xiangyang Ji
机构
*
Xidian University(西安电子科技大学)
;
Tsinghua University(清华大学)
;
Beijing Institute of Technology(北京理工大学)
;
Guangdong Laboratory of Artificial Intelligence and Digital Economy(SZ)(广东人工智能与数字经济实验室(深圳))
GTPred: Benchmarking MLLMs for Interpretable Geo-localization and Time-of-capture Prediction
GTPred:评估多模态大语言模型在可解释地理定位和拍摄时间预测中的基准测试
Jinnao Li, Zijian Chen, Tingzhu Chen, Changbo Wang
机构
*
School of Computer Science and Technology, East China Normal University(东华大学计算机科学与技术学院)
;
Institute of Image Communication and Information Processing, Shanghai Jiao Tong University(上海交通大学图像通信与信息处理研究院)
;
School of Humanities, Shanghai Jiao Tong University(上海交通大学人文学院)
;
Shanghai AI Laboratory(上海人工智能实验室)
机构
*
Computer Science Department, University of Crete(塞萨洛尼基大学计算机科学系)
;
Institute of Computer Science (ICS), Foundation for Research & Technology – Hellas (FORTH)(希腊基础研究与技术机构计算机科学研究所)