Skip-It? Theoretical Conditions for Layer Skipping in Vision-Language Models
跳过层?视觉-语言模型中层跳过的理论条件
Max Hartman, Vidhata Jayaraman, Moulik Choraria, Akhil Bhimaraju, Lav R. Varshney
机构
*
Department of Electrical and Computer Engineering, The University of Illinois, Champaign, IL, United States(电气与计算机工程系,伊利诺伊大学香槟分校)
;
Department of Mathematics, The University of Illinois, Champaign, IL, United States(数学系,伊利诺伊大学香槟分校)
;
AI Innovation Institute, Stony Brook University, Stony Brook, NY, United States(人工智能创新研究所,石溪大学)
Better with Less: Tackling Heterogeneous Multi-Modal Image Joint Pretraining via Conditioned and Degraded Masked Autoencoder
更少的协同,更好的表现:通过条件和降质掩码自编码器解决异质多模态图像联合预训练
Bowen Peng, Yongxiang Liu, Jie Zhou, Xiaodong Chen, Tianpeng Liu, Xiaogang Yu, Li Liu
机构
*
College of Electronic Science and Technology, National University of Defense Technology (NUDT)(电子科学与技术学院,国防科技大学)
;
Beijing Institute of Remote Sensing Information(遥感信息研究所)
机构
*
Hangzhou Institute for Advanced Study, University of Chinese Academy of Sciences(中国科学院大学杭州高等研究院)
;
The Chinese University of Hong Kong(香港中文大学)
;
University of Science and Technology of China(中国科学技术大学)
;
University of Chinese Academy of Sciences(中国科学院大学)
;
vivo AI Lab(vivo AI实验室)
FastDiSS: Few-step Match Many-step Diffusion Language Model on Sequence-to-Sequence Generation--Full Version
FastDiSS: 少步匹配多步扩散语言模型在序列到序列生成中的应用——完整版本
Dat Nguyen-Cong, Tung Kieu, Hoang Thanh-Tung
机构
*
FPT Software AI Center, FPT Corporation(FPT软件人工智能中心,FPT公司)
;
Department of Computer Science, Aalborg University(奥尔堡大学计算机科学系)
;
Quantum AI and Cyber Security Institute, FPT Corporation(FPT公司量子人工智能与网络安全研究所)
CommentsThis version strengthens the theoretical and empirical grounding of the CI metric, including explicit analysis of structural dependencies and ranking stability under ablations (e.g., excluding Turn 4). Claims regarding scale and robustness are revised to avoid overgeneralization. The evaluation protocol, jury methodology, and limitations are expanded to clarify assumptions and boundary conditions
Toward Real-Time Surgical Scene Segmentation via a Spike-Driven Video Transformer with Spike-Informed Pretraining
迈向实时手术场景分割的脉冲驱动视频变换器及其基于脉冲的预训练
Shihao Zou, Jingjing Li, Wei Ji, Jincai Huang, Kai Wang, Guo Dan, Weixin Si, Yi Pan
机构
*
Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(深圳先进技术研究院,中国科学院)
;
University of Alberta(阿尔伯塔大学)
;
School of Medicine, Yale University(耶鲁大学医学院)
;
Southern University of Science and Technology(南方科技大学)
;
Nanfang Hospital Southern Medical University(南方医科大学南华医院)
;
School of Biomedical Engineering, Shenzhen University(深圳大学生物医学工程学院)
;
Faculty of Computer Science and Control Engineering, Shenzhen University of Advanced Technology(深圳先进技术研究院计算机科学与控制工程学院)