Seeing the Forest and the Trees: Query-Aware Tokenizer for Long-Video Multimodal Language Models
看清森林与树木:面向长视频多模态语言模型的查询感知分词器
Siyou Li, Huanan Wu, Juexi Shao, Yinghao Ma, Yujian Gan, Yihao Luo, Yuwei Wang, Dong Nie, Lu Wang, Wenqing Wu, Le Zhang, Massimo Poesio, Juntao Yu
机构
*
Queen Mary University of London(伦敦女王学院)
;
University of Sheffield(谢菲尔德大学)
;
Imperial College London(伦敦帝国学院)
;
Pengcheng Laboratory(鹏城实验室)
;
Meta Inc(Meta公司)
;
Meituan Inc(美团公司)
;
Nanjing University of Science(南京理工大学)
;
University of Birmingham(伯明翰大学)
;
Utrecht University(乌得勒支大学)
专题命中
预训练与数据
:language model(title,abstract);large language model(abstract)
机构
*
Shanghai Jiao Tong University(上海交通大学)
;
Washington University in St. Louis(华盛顿大学)
;
University of Science and Technology Beijing(北京科技大学)
;
Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
;
Department of Surgical Oncology and General Surgery, Key Laboratory of Precision Diagnosis and Treatment of Gastrointestinal Tumours, Ministry of Education, The First Hospital of China Medical University(外科肿瘤科和普通外科,国家教育委员会胃肠道肿瘤精准诊断与治疗重点实验室,中国医科大学第一医院)
;
Shanghai Innovation Institute(上海创新研究院)
;
Sensetime Research(商汤科技研究院)
EchoJEPA: A Latent Predictive Foundation Model for Echocardiography
EchoJEPA:一种用于超声心动图的潜在预测基础模型
Alif Munim, Adibvafa Fallahpour, Teodora Szasz, Ahmadreza Attarpour, River Jiang, Brana Sooriyakanthan, Maala Sooriyakanthan, Heather Whitney, Jeremy Slivnick, Barry Rubin, Wendy Tsang, Bo Wang
机构
*
University Health Network(大学健康网络)
;
University of Toronto(多伦多大学)
;
University of Chicago(芝加哥大学)
;
University of California, San Francisco(加州大学旧金山分校)
;
Vector Institute(向量研究所)
;
Cohere Labs(Cohere实验室)
机构
*
Department of Pathology and Laboratory Medicine, University of Pennsylvania(病理学与实验室医学系,宾夕法尼亚大学)
;
Department of Electrical and System Engineering, University of Pennsylvania(电气与系统工程系,宾夕法尼亚大学)
;
Department of Biomedical Engineering, Georgia Institute of Technology and Emory University(生物医学工程系,佐治亚理工学院和埃默里大学)
;
NVIDIA Corporation(NVIDIA公司)
;
Department of Biostatistics, Epidemiology and Informatics, University of Pennsylvania(生物统计学、流行病学与信息学系,宾夕法尼亚大学)
Proprioception Enhances Vision Language Model in Generating Captions and Subtask Segmentations for Robot Task
本体感知增强视觉语言模型在为机器人任务生成描述和子任务分割中的应用
Kanata Suzuki, Shota Shimizu, Tetsuya Ogata
机构
*
Faculty of Science and Engineering, Waseda University(工学部,早稻田大学)
;
Artificial Intelligence Laboratory, Fujitsu Limited(Fujitsu 人工智能实验室)
;
National Institute of Advanced Industrial Science and Technology(国家先进工业科学与技术研究院)
机构
*
Fudan University(复旦大学)
;
Shanghai Innovation Institut(上海创新研究院)
;
University of Southern California(南加州大学)
;
Huawei Technologies Co., Ltd(华为技术有限公司)
;
Ningbo Key Laboratory of Spatial Intelligence and Digital Derivative, Institute of Digital Twin, EIT(宁波空间智能与数字衍生关键实验室,数字孪生研究院,EIT)
CommentsWe ran experiments from mid-August to mid-September 2025, notified affected providers shortly after, and now make our findings public after a 90-day disclosure window