arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.10010cs.LG

CurveFP:用于语言模型的带闭包乘积的有理基数对数数据类型

CurveFP: Co-Designing Numerical Representation and Product Arithmetic for Language Models

Ye Qiao

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出CurveFP数据类型,通过闭包乘积算术协同设计,在少1位元素位数下提升语言模型性能,降低运算误差,优于FP8且更适合紧凑部署。

中文摘要 AI 辅助

低精度数据类型可降低语言模型的成本,但大多数格式仅优化标量保真度,未改变其乘积带来的算术运算。本文提出CurveFP,一种闭包乘积码本族,在紧凑块尺度下将量化幅值分布在交错对数曲线上。有理基数可调节动态范围与局部分辨率的平衡,均匀曲线索引使每个非零乘积具有代数闭包性。乘积运算变为精确的符号异或与整数索引更新,推导得到的有限相位数决定累加调度。我们将该代数实例化为用于训练的CurveFP 8 E4C3/E5C2和用于紧凑部署的CurveFP 7 E3C3。评估显示,CurveFP 7在4个7B至9B模型上的困惑度优于张量级FP8,且元素位数少1位,仅比原生质量低1.32%;CurveFP 8在全部36对前向和反向GEMM比较中降低了操作数NMSE。在3组匹配的1.283亿参数三元组中,每种模式在每个随机种子下完成30亿token的预训练;CurveFP 8的平均BF16推理困惑度为22.5366,而FP8为22.5407,且在所有3个随机种子下格式诱导的惩罚更低。一个36单元的下游矩阵显示,CurveFP 8训练的检查点在全部12个随机种子-格式比较中,WikiText-103困惑度更低,且存在混合PG-19及任务差异。综上,这些结果表明CurveFP是一种算术协同设计,兼具FP8级数值性能、7位推理能力和大幅简化的乘积路径。

英文摘要

Low-precision formats usually optimize scalar fidelity while inheriting conventional product arithmetic. We introduce CurveFP, a block-scaled family that distributes magnitudes across interleaved logarithmic curves. Uniform curve indices make every nonzero product an exact sign and integer-index update, while a rational radix exposes the finite phase schedule required for accumulation. We instantiate the algebra as CurveFP8 E4C3/E5C2 for training and CurveFP7 E3C3 for compact inference. On four 7B-9B models, CurveFP7 beats tensorwise FP8 perplexity with one fewer element bit and stays within 1.32% of native quality. CurveFP8 lowers error in all 36 paired training-GEMM comparisons. Across three matched 3B-token pretraining triplets, it reaches mean BF16-inference perplexity 22.5366 versus 22.5407 for FP8 and has a lower format penalty in every seed. Downstream evaluation shows transfer parity and a consistent WikiText-103 gain. In a preliminary 4x4 Nangate45 spatial accelerator tile, CurveFP8 uses one fewer product register and 4.6% less area than timing-closing FP8 at 500 MHz. These results support CurveFP as a numerical and arithmetic co-design, while leaving system-level efficiency to future study.

发表机构

  • University of California, Irvine(加州大学欧文分校)

机构由 AI 辅助整理,请以论文原文为准。

↑