When Gradient Importance Lies: Adaptive LoRA Rank Allocation Fails Under GRPO
基于GRPO的LoRA秩分配:一项实证研究
机构 * Independent Researcher(独立研究者)
AI总结 本文探讨LoRA秩分配在强化学习中的有效性,发现GRPO环境下梯度重要性无法预测容量需求,非均匀分配反而降低准确性。
Comments Accepted at the Insights from Negative Results in NLP Workshop, EMNLP 2026