过程性知识不是低秩的:为什么LoRA无法内化多步骤过程
Procedural Knowledge Is Not Low-Rank: Why LoRA Fails to Internalize Multi-Step Procedures
浏览论文内容
中文总结 AI 辅助
研究发现对于过程性知识,LoRA在保持效率优势的秩上无法与完全微调匹配。通过实验表明在旅行预订等任务中LoRA配置失败,跨域复制也证实其普遍不佳,分析指出完全微调权重变化的有效秩高,LoRA难以企及,这是智能体应用的基本限制。
中文摘要 AI 辅助
像LoRA这样的参数高效微调方法已成为调整大语言模型的默认方法,在遵循指令、风格迁移和事实适应等方面取得了成功。我们表明,对于过程性知识,即通过条件分支遵循多步骤过程直至终端状态的能力,LoRA在保持效率优势的秩上无法与完全微调相匹配。在一个过程性旅行预订任务(14个节点)上进行系统消融(r = 16 - 128),所有LoRA配置均一致失败(任务成功率<=2.54,而完全微调为4.11,所有p < 0.001),分数在较高秩时下降,尽管保持了95 - 99%的对话完成率。在8B的Zoom支持(14个节点)和保险理赔(55个节点)上的跨域复制证实了这种失败具有普遍性:在r = 32和r = 128时,LoRA平均比完全微调低0.8 - 2.2分,在最复杂的过程中差距最大。将秩从32提高到128四倍仅带来边际改善但未缩小差距。对完全微调产生的权重变化进行奇异值分解分析解释了原因:在3B和8B的三个域中,更新的平均有效秩范围从761到1026,秩128仅捕获了43 - 51%的平方Frobenius范数。这些发现共同表明,对于过程性任务,LoRA远不及完全微调,这是智能体应用的一个基本限制。
英文摘要
Parameter-efficient fine-tuning methods like LoRA have become the default for adapting large language models, succeeding across instruction following, style transfer, and factual adaptation. We show that for procedural knowledge--the ability to follow multi-step procedures with conditional branching through to terminal states--LoRA fails to match full fine-tuning at the ranks where it retains its efficiency advantage. In a systematic ablation (r = 16--128) on a procedural travel booking task (14 nodes), all LoRA configurations fail uniformly (task success <= 2.54 vs. 4.11 for full fine-tuning, all p < 0.001), with scores decreasing at higher ranks--despite maintaining 95--99% conversation completion rates. Cross-domain replication on Zoom support (14 nodes) and insurance claims (55 nodes) at 8B confirms the failure generalizes: LoRA underperforms full fine-tuning by 0.8--2.2 points on average at both r = 32 and r = 128, with the largest gap on the most complex procedure. Quadrupling rank from 32 to 128 provides marginal improvement but does not close the gap. SVD analysis of the weight changes produced by full fine-tuning explains why: across three domains at both 3B and 8B, the mean effective rank of the update ranges from 761 to 1,026, and rank 128 captures only 43--51% of the squared Frobenius norm. Together, these findings establish that for procedural tasks LoRA falls well short of full fine-tuning--a fundamental limitation for agentic applications.
发表机构
- University of Melbourne(墨尔本大学)
机构由 AI 辅助整理,请以论文原文为准。