发表机构
Jay Chaudhry Software Innovation Centre; Indian Institute of Technology (BHU)(Jay Chaudhry 软件创新中心; 印度理工学院(巴纳拉斯印度教大学))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
OrBIT通过从轨道动力学中学习局部几何结构来指导嵌入压缩,实现GPT-2上37.9倍和7B表上23倍以上的压缩,同时保持竞争力的率失真性能。
AI 中文摘要
嵌入表是现代语言模型中最大的组件之一。大多数压缩方法固定一种编码几何结构,如坐标块、低秩子空间或无限制的码本,并在该结构内进行优化。我们转而探究编码几何结构本身是否可以被发现。我们引入了OrBIT,一个结构引导的嵌入压缩框架,它从轨道动力学中学习可复用的局部几何结构,并利用该结构约束一小部分共享码字。全局重建残差随后决定固定的编码预算花费在何处,而冗余的重叠图块使得局部误差在拼接后能够相互补偿。我们的理论展示了紧密图块几何结构如何控制失真,全局残差如何指导顺序分配,以及数据几何引导的细化如何改进编解码器。由此产生的轨道机制被编译掉,留下一个紧凑的解码器,其中学习到的结构决定了存储什么、容量分配在哪里以及局部信息如何在全局范围内组装。在四个LLM嵌入表上,相对于16位存储,OrBIT在GPT-2上实现了37.9倍的压缩,在每个7B表上实现了超过23倍的压缩,同时与已建立的量化和低秩基线相比,提供了具有竞争力的率失真性能。
英文摘要
Embedding tables are among the largest components of modern language models. Most compression methods fix a coding geometry such as coordinate blocks, low-rank subspaces, or unrestricted codebooks, and optimize within it. We instead ask whether the coding geometry can itself be discovered. We introduce \emph{OrBIT}, a structure-guided embedding compression framework that learns reusable local geometry from orbit dynamics and uses it to constrain a small set of shared codewords. The global reconstruction residual then decides where the fixed coding budget is spent, while redundant overlapping charts let local errors compensate one another after gluing. Our theory shows how tight-chart geometry controls distortion, how the global residual directs sequential allocation, and how data-geometry-guided refinement improves the codec. The resulting orbit machinery is compiled away, leaving a compact decoder in which the learned structure governs what is stored, where capacity is allocated, and how local information is assembled globally. Across four LLM embedding tables, OrBIT achieves $37.9\times$ compression on GPT-2 and over $23\times$ on each 7B table relative to 16-bit storage, while delivering competitive rate-distortion performance against established quantization and low-rank baselines.