Junk DNA Hypothesis: Pruning Small Pre-Trained Weights Irreversibly and Monotonically Impairs "Difficult" Downstream Tasks in LLMs
垃圾DNA假说:修剪小的预训练权重不可逆且单调地损害LLM中的“困难”下游任务
机构 * University of Surrey ; Eindhoven University of Technology ; University of Texas at Austin ; Intel Labs ; University of Oxford
AI总结 该研究提出垃圾DNA假说,指出LLM预训练权重中存在关键知识,修剪小权重会单调损害困难下游任务性能,且即使允许持续训练也无法弥补损失。
Comments Published at ICML 2024