发表机构
Qualcomm AI Research; Johns Hopkins University(高通人工智能研究院; 约翰霍普金斯大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出“最佳可用”KV缓存策略及早期退出修复,使循环语言模型实现动态计算,减少高达30%的FLOPs和KV内存,同时保持性能,并在Ouro及较小模型上验证。
AI 中文摘要
循环语言模型参数高效,并有望实现动态计算(在简单标记上节省内存和浮点运算)。然而,最先进的开源循环语言模型(Ouro模型)虽然以这种动态计算能力进行训练,但在实践中并未实现这一能力,因为每次循环迭代(深度)都需要其自身的KV缓存层级,从而要求执行所有循环计算。此外,Ouro的早期退出先验被平等地施加于每个标记,导致无论难度如何,所有标记都经历静态的较低深度处理。在本工作中,我们提出了一种简单的“最佳可用”KV缓存策略,该策略开箱即用,在性能与深度空间中开辟了新的前沿。我们的方法能够实现高达30%的浮点运算和KV内存减少,同时保持全深度性能,展示了循环语言模型的真正灵活性。此外,以对该KV缓存策略的认知来训练循环语言模型,可提升性能和效率。最后,我们对早期退出先验强制执行目标应用了一个小而有效的修复,使标记能够基于努力程度在真正异构的深度退出。我们的发现已在Ouro模型以及从头预训练的较小循环语言模型上得到验证。
英文摘要
Looped LMs are parameter efficient and promise dynamic computation (saving memory and FLOPs on easy tokens). However, state-of-the-art open Looped LMs trained with this dynamic computation capability (Ouro models) do not realize it in practice as each loop iteration (depth) requires its own level of KV-cache, necessitating all loop computations. Moreover, Ouro's early-exit prior is enforced on each token equally, which results in static lower-depth like processing of all tokens regardless of difficulty. In this work, we propose a simple "best-available" KV caching strategy that works out-of-the-box, creating a new frontier in the performance vs depth space. Our approach enables up to 30% reduction in FLOPs and KV memory while retaining full-depth performance, showing the true flexibility of Looped LMs. Furthermore, training looped LMs with awareness about this KV caching strategy improves performance and efficiency. Finally, we apply a small but effective fix to the early-exit prior enforcement objective that makes tokens exit at truly heterogeneous depths based on effort. Our findings are validated on Ouro models as well as smaller looped LMs pre-trained from scratch.