发表机构
Geometric Intelligence Research Lab., Department of Electrical Engineering and Computer Science University of Wyoming(几何智能研究实验室,电气与计算机科学系,怀俄明大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
通过分析社区竞赛中的2037个拉取请求和1430个评分提交,构建84种优化技术分类,测量每种技术对bits-per-byte的贡献,发现大多数技术收益在竞争提交中缩小,仅少数方法能持续提升性能。
AI 中文摘要
在严格的工件预算下,语言模型能改进多少?参数高尔夫将此问题作为开放社区挑战提出,参与者训练最佳语言模型,完整工件(训练代码+压缩权重)必须适合16 MB,并在8xH100 SXM GPU上十分钟内完成训练。质量以bits-per-byte (BPB)衡量,即编码未见文本每个字节所需的平均比特数。我们分析了竞赛中的2,037个拉取请求和1,430个干净评分提交,构建了84种优化技术的分类,并测量每种技术对BPB的贡献。验证排行榜分数从第一阶段的1.2244降至第三阶段的1.058 BPB——降低了13.6%,尽管单项技术很少能将BPB改进超过1%。我们表明,大多数技术的收益在竞争提交中缩小,从而分离出少数能在各堆栈中提升性能的方法。
英文摘要
How far can a language model improve under a strict artifact budget? Parameter Golf posed this question as an open community challenge in which participants trained the best language model, with the complete artifact (training code + compressed weights) required to fit within 16 MB and to be trained in under ten minutes on 8xH100 SXM GPUs. Quality was measured in bits-per-byte (BPB), the average number of bits required to encode each byte of unseen text. We analyze 2,037 pull requests and 1,430 clean-scored submissions from the contest, build a taxonomy of 84 optimization techniques, and measure each technique's contribution to BPB. The verified leaderboard score dropped from 1.2244 to 1.058 BPB across three phases, a 13.6% reduction, despite individual techniques rarely improving BPB by more than 1%. We show that most techniques' gains shrink when re-measured among competitive submissions, isolating the few methods that help regardless of the surrounding stack. Code and data are available at https://github.com/PMP56/pmgolf-analysis.
CommentsAccepted at the BabyLM Workshop, EMNLP 2026