Comments10 pages, 10 figures. Accepted at the 55th International Conference on Parallel Processing (ACM ICPP 2026), Singapore. Code, training pipeline, and measured data: https://github.com/nobeldhar/NeuroPrefetcher
Every Coin Has Two Sides: On the Dual Nature of Generalization in On-Policy Distillation of Large Language Models
硬币皆有两面:关于大型语言模型的策略内蒸馏中泛化的双重性质
Zhaoyi Li, Deyang Kong, Yuan Wei, Evan Yang, Ranran Shen, Mahardika Krisna Ihsani, Ming Yang, Wei Zhang, Chuan Hao, Jian Yang, Ran Tao, Bryan Dai, Shikun Zhang, Wei Ye, Ying Wei, Defu Lian
机构
*
University of Science and Technology of China(中国科学技术大学)
;
Peking University(北京大学)
;
IQuest Research(IQuest研究院)
;
MBZUAI(穆罕默德·本·扎耶德人工智能大学)
;
Zhejiang University(浙江大学)
专题命中
效率与部署
:large language model(title);language model(title);分类 cs.CL
HiViS: Hiding Visual Tokens from the Drafter for Speculative Decoding in Vision-Language Models
HiViS: 为视觉语言模型中的推测解码隐藏视觉标记
Zhinan Xie, Peisong Wang, Shuang Qiu, Jian Cheng
机构
*
C 2 DL, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
;
AIRIA
;
City University of Hong Kong(香港城市大学)
专题命中
效率与部署
:language model(title,abstract);large language model(abstract);分类 cs.AI、cs.LG
机构
*
Digital Environment Research Institute (DERI), Queen Mary University of London(伦敦大学玛丽女王学院数字环境研究所(DERI))
;
Shanghai Institute of Artificial Intelligence for Education, East China Normal University(华东师范大学上海人工智能教育研究院)
;
National Institute of Education, Nanyang Technological University(南洋理工大学国家教育学院)
;
School of Computer Science and Engineering, Jiangsu University of Science and Technology(江苏科技大学计算机科学与工程学院)