机构
*
School of AI for Science, Peking University(科学人工智能学院,北京大学)
;
School of Electronic and Computer Engineering, Peking University(电子与计算机工程学院,北京大学)
;
School of Computer Science, Peking University(计算机科学学院,北京大学)
SVAG-Bench: A Large-Scale Benchmark for Multi-Instance Spatio-temporal Video Action Grounding
SVAG-Bench:多实例时空视频动作定位的大规模基准
Tanveer Hannan, Shuaicong Wu, Mark Weber, Suprosanna Shit, Jindong Gu, Rajat Koner, Aljoša Ošep, Laura Leal-Taixé, Thomas Seidl
机构
*
LMU Munich(慕尼黑大学)
;
MCML
;
Technical University of Munich(慕尼黑技术大学)
;
University of Zurich(苏黎世大学)
;
University of Oxford(牛津大学)
;
Amazon(亚马逊)
;
NVIDIA(英伟达)
机构
*
Institute of Information Science and Technologies of the National Research Council (ISTI-CNR)(意大利国家研究理事会信息科学与技术研究所)
;
University of Pisa - Department of Information Engineering(比萨大学信息工程系)
The Grounding Gap: How LLMs Anchor the Meaning of Abstract Concepts Differently from Humans
概念基础差距:LLMs如何以不同于人类的方式锚定抽象概念的含义
Odysseas S. Chlapanis, Orfeas Menis Mastromichalakis, Christos H. Papadimitriou
机构
*
Department of Informatics(信息学院)
;
Athens University of Economics and Business(雅典经济与商业大学)
;
Archimedes, Athena Research Center(阿基米德·雅典研究中心)
;
Instituto de Telecomunicações(电信研究所)
;
Department of Computer Science(计算机科学系)
;
Columbia University(哥伦比亚大学)
机构
*
Faculty of Electronic Engineering and Computer Science, Ningbo University(宁波大学电子工程与计算机科学学院)
;
College of Computer Science, Inner Mongolia University(内蒙古大学计算机学院)
;
College of Computing and Engineering, Hunan Normal University(湖南师范大学计算机与工程学院)
;
Faculty of Computing, Georg-August-Universität Göttingen(哥廷根大学计算机学院)
;
Center for Artificial Intelligence, Mathematical and Data Science, Nagoya University(名古屋大学人工智能、数学与数据科学中心)
Multi-Scale Contrastive Learning for Video Temporal Grounding
多尺度对比学习用于视频时间定位
Thong Thanh Nguyen, Yi Bin, Xiaobao Wu, Zhiyuan Hu, Cong-Duy T Nguyen, See-Kiong Ng, Anh Tuan Luu
机构
*
Institute of Data Science (IDS), National University of Singapore(数据科学研究所(IDS),新加坡国立大学)
;
Tongji University(同济大学)
;
Nanyang Technological University (NTU)(南洋理工大学)
TTL: Test-time Textual Learning for OOD Detection with Pretrained Vision-Language Models
TTL: 用于基于预训练视觉-语言模型的分布外检测的测试时文本学习
Jinlun Ye, Jiang Liao, Runhe Lai, Xinhua Lu, Jiaxin Zhuang, Zhiyong Gan, Ruixuan Wang
机构
*
Sun Yat-sen University(中山大学)
;
China United Network Communications Corporation Limited Guangdong Branch(中国联合网络通信集团有限公司广东分公司)
;
Peng Cheng Laboratory(鹏城实验室)
;
Hong Kong University of Science and Technology(香港科技大学)
;
Key Laboratory of Machine Intelligence and Advanced Computing, MOE(教育部机器智能与高级计算重点实验室)
机构
*
UVLab, Department of Computer Science, University of Warwick(华威大学计算机科学系UVLab)
;
Department of Automation, University of Cambridge(剑桥大学自动化系)
;
Department of Computer Science, The University of Sheffield(谢菲尔德大学计算机科学系)
Vision Language Models versus Machine Learning Models Performance on Polyp Detection and Classification in Colonoscopy Images
视觉语言模型与机器学习模型在结肠镜图像息肉检测与分类中的性能对比
Mohammad Amin Khalafi, Seyed Amir Ahmad Safavi-Naini, Ameneh Salehi, Nariman Naderi, Dorsa Alijanzadeh, Pardis Ketabi Moghadam, Kaveh Kavosi, Negar Golestani, Shabnam Shahrokh, Soltanali Fallah, Jamil S Samaan, Nicholas P. Tatonetti, Nicholas Hoerter, Girish Nadkarni, Hamid Asadzadeh Aghdaei, Ali Soroush
专题命中
视觉定位与Grounding
:vision language model(title);vision-language model(abstract);分类 cs.CV