AI 中文总结
该研究指出共享学习率比较参数高效微调方案的缺陷,提出冻结解码器仅修复编码器的低内存方案,在SVD-based KV缓存压缩中实现性能相当且参数与内存开销减半。
AI 中文摘要
在单一共享学习率下比较参数高效微调方案是常见但存在缺陷的做法:当被比较的方案具有差异极大的可训练参数数量时,共享学习率会同时压低参数较多方案的均值并放大其方差,从而为参数最少的方案制造出看似具有多种子显著性的优势,但这并非真实效果。我们在具体场景中记录了这种混淆效应:事后基于SVD的KV缓存压缩,即通过将预训练模型的键/值权重分解为下投影(“编码器”)和上投影(“解码器”),将其转换为低秩(多头潜在注意力风格)缓存,之后进行短时间微调(“修复”)以恢复因截断损失的准确率。在共享学习率下,冻结解码器并仅修复编码器的方案,看起来比修复解码器或同时修复两个投影的方案具有明显优势;但当为每个方案单独调整学习率后,这种表观优势消失,仅修复编码器的方案与其他方案达到性能相当,同时可实现3倍更少的可训练参数和3倍更少的优化器状态内存的实测节省。我们在视觉语言模型Qwen2.5-VL-3B-Instruct上,针对该协议覆盖的一个压缩率,通过每方案学习率调整和每个配置3个种子验证了这种性能相当,并在两个主干网络的纯文本测试平台上复现了该结果。因此,仅修复编码器是训练时改造低秩KV缓存压缩的低内存替代方案,我们记录并纠正的共享学习率陷阱,为比较任何可训练参数数量存在差异的微调方案提供了警示性结果。
英文摘要
Comparing parameter-efficient fine-tuning recipes under a single, shared learning rate is a common but flawed practice: when the arms being compared have very different trainable-parameter counts, a shared rate can simultaneously depress the larger arms' means and inflate their variance, manufacturing a large, seemingly multi-seed-significant advantage for the smallest arm that is not a real effect. We document this confound in a concrete setting: post-hoc SVD-based KV-cache compression, where an already-pretrained model is converted to a low-rank (multi-head-latent-attention-style) cache by factorizing its key/value weights into a down-projection ("encoder") and an up-projection ("decoder"), after which a short fine-tune ("healing") recovers the accuracy lost to truncation. Under a shared learning rate, freezing the decoder and healing only the encoder looks like a clear win over healing the decoder or both factors; once every arm is given its own tuned learning rate, that apparent advantage disappears, and encoder-only healing instead reaches parity with the alternatives, at a real, measured saving of 3x fewer trainable parameters and 3x less optimizer-state memory. We verify this parity with per-arm learning-rate tuning and three seeds per configuration on a vision-language model (Qwen2.5-VL-3B-Instruct), at the one compression ratio this protocol covers, and replicate it on a text-only testbed across two backbones. Encoder-only healing is therefore a lower-memory drop-in recipe for retrofitting low-rank KV-cache compression at training time, and the shared-learning-rate pitfall we document and correct is a cautionary result for comparing any fine-tuning recipes whose arms differ in trainable-parameter count.
CommentsPreprint, Under Review