EVE: Towards End-to-End Video Subtitle Extraction with Vision-Language Models
EVE:基于视觉-语言模型的端到端视频字幕提取
Haiyang Yu, Mengyang Zhao, Jinghui Lu, Ke Niu, Yanjie Wang, Weijie Yin, Weitao Jia, Teng Fu, Yang Liu, Jun Liu, Hong Chen
机构
*
College of Computer Science and Artificial Intelligence, Fudan University(复旦大学计算机科学与人工智能学院)
;
ByteDance Inc.(字节跳动公司)
;
College of Electronic and Information Engineering, Tongji University(同济大学电子与信息工程学院)
;
School of Computing and Communications, Lancaster University(兰卡斯特大学计算机与通讯学院)
Bootstrapping Physics-Grounded Video Generation through VLM-Guided Iterative Self-Refinement
通过VLM引导的迭代自优化提升物理导向的视频生成
Yang Liu, Xilin Zhao, Peisong Wen, Siran Dai, Qingming Huang
机构
*
School of Computer Science and Technology, University of Chinese Academy of Sciences(中国科学院大学计算机科学与技术学院)
;
School of Computer Science and Technology, Beijing Institute of Technology(北京理工大学计算机科学与技术学院)
;
Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所)
;
School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络安全学院)
;
Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所)
机构
*
Department of Electrical and Computer Engineering, University of California, Riverside, USA(电气与计算机工程系,加州大学河滨分校)
;
Thomas Lord Department of Computer Science, University of Southern California, USA(汤姆斯·劳德计算机科学系,南加州大学)
Glo-VLMs: Leveraging Vision-Language Models for Fine-Grained Diseased Glomerulus Classification
Zhenhao Guo, Rachit Saluja, Tianyuan Yao, Quan Liu, Yuankai Huo, Benjamin Liechty, David J. Pisapia, Kenji Ikemura, Mert R. Sabuncu, Yihe Yang, Ruining Deng
机构
*
New York University(纽约大学)
;
Cornell Tech(康奈尔科技)
;
Vanderbilt University(范德比尔特大学)
;
Weill Cornell Medicine(韦尔医学院)
;
Northwell Health(北well健康)
From Classification to Cross-Modal Understanding: Leveraging Vision-Language Models for Fine-Grained Renal Pathology
Zhenhao Guo, Rachit Saluja, Tianyuan Yao, Quan Liu, Junchao Zhu, Haibo Wang, Daniel Reisenbüchler, Yuankai Huo, Benjamin Liechty, David J. Pisapia, Kenji Ikemura, Steven Salvatoree, Surya Seshane, Mert R. Sabuncu, Yihe Yang, Ruining Deng
机构
*
New York University(纽约大学)
;
Cornell Tech(康奈尔科技)
;
Vanderbilt University(范德比大学)
;
Carnegie Mellon University(卡内基梅隆大学)
;
University of Regensburg(莱茵河畔大学)
;
Weill Cornell Medicine(韦尔·科恩医学中心)
;
Northwell Health(北well健康)
机构
*
School of Information Science and Technology, University of Science and Technology of China(信息科学与技术学院,中国科学技术大学)
;
Institute of Artificial Intelligence (TeleAI), China Telecom, P. R. China(人工智能研究所(TeleAI),中国电信,中华人民共和国)
;
College of Computer Science, Wuhan University(计算机科学学院,武汉大学)
;
School of Artificial Intelligence, OPtics and ElectroNics (iOPEN), Northwestern Polytechnical University(人工智能学院,光学与电子学(iOPEN),西北工业大学)
Efficient Video Sampling: Pruning Temporally Redundant Tokens for Faster VLM Inference
Natan Bagrov, Eugene Khvedchenia, Borys Tymchenko, Shay Aharon, Lior Kadoch, Tomer Keren, Ofri Masad, Yonatan Geifman, Ran Zilberstein, Tuomas Rintamaki, Matthieu Le, Andrew Tao
机构
*
Shanghai Jiao Tong University(上海交通大学)
;
SenseTime Research(商汤科技研究院)
;
Central South University(中南大学)
;
Tongji University(同济大学)
;
Beihang University(北航)
;
The Chinese University of Hong Kong(香港中文大学)
PiCo: Jailbreaking Multimodal Large Language Models via Pictorial Code Contextualization
Aofan Liu, Lulu Tang, Ting Pan, Yuguo Yin, Bin Wang, Ao Yang
机构
*
School of Artificial Intelligence, Wuhan University(武汉大学人工智能学院)
;
School of Electronic and Computer Engineering, Peking University(北京大学电子与计算机工程学院)
;
Beijing Academy of Artificial Intelligence(北京人工智能研究院)
专题命中
VLM训练与架构
:multimodal large language model(title,abstract);分类 cs.AI
CommentsAccepted to IEEE International Conference on Multimedia and Expo (ICME) 2025
The Security Threat of Compressed Projectors in Large Vision-Language Models
Yudong Zhang, Ruobing Xie, Xingwu Sun, Jiansheng Chen, Zhanhui Kang, Di Wang, Yu Wang
机构
*
Department of Electronic Engineering, Tsinghua University(清华大学电子工程系)
;
Large Language Model Department, Tencent(腾讯大语言模型部门)
;
School of Computer and Communication Engineering, University of Science and Technology Beijing(北京科技大学计算机与通信工程学院)
;
Faculty of Science and Technology, University of Macau(澳门大学科技学院)
专题命中
VLM训练与架构
:vision-language model(title);visual language model(abstract);分类 cs.AI