arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

双线性优化散度:诊断因子约束的LoRA持续学习

Bilinear Optimization Divergence: Diagnosing Factor-Constrained LoRA Continual Learning

YongShun Wang, JianLin Su, Yong Ma

arXiv 2609.23594首次发表:更新:

AI 中文总结

本文提出双线性优化散度(BOD)诊断框架,分析LoRA持续学习中因子约束的保护机制,并据此设计半冻结正交路由(SFOR)与硬保护策略,在Qwen3-8B上显著提升反向迁移并降低遗忘。

AI 中文摘要

LoRA因子中的正交性本身并不指定组合更新保护的内容:答案取决于任务起始状态、参数化方式以及实际优化器位移。我们通过双线性优化散度(BOD)形式化这一问题,这是一种基于锚点的诊断方法,用于评估选定历史特征上的有效更新响应。有限步分析区分了两种情况。在共享适配器中,保护路由位移会通过变化的伴随因子留下学习锚点残差。在新的零输出块中,可行的路由状态可以在两个当前因子保持可训练的同时保护组合更新。这些条件产生了用于共享适配器的半冻结正交路由(SFOR)和用于累积O-LoRA的当前块硬保护;权重残差投影(WRP)在优化器步骤后强制执行所需的位移。受控的两任务轨迹验证了预测的残差路径,将共享族的归一化历史响应从19.12%降低到0.005%,将累积族从7.72%降低到0.002%。在Qwen3-8B上的四任务实验表征了由此产生的权衡:SFOR将反向迁移(BWT)从-2.47提高到-0.86,平均准确率(AA)几乎不变,而O-LoRA硬保护将三阶平均AA从80.27%提高到81.30%,遗忘度量(FM)从2.20降低到0.43。组件控制还表明,更严格的可行性不一定能提高最终任务性能。综合分析和证据提供了基于架构条件的说明,解释应强制执行哪种约束、如何强制执行以及如何解释其经验价值。

英文摘要

Orthogonality in a LoRA factor does not by itself specify what the composed update protects: the answer depends on the task-start state, the parameterization, and the realized optimizer displacement. We formalize this question through Bilinear Optimization Divergence (BOD), an anchor-relative diagnostic of effective-update response on selected historical features. The finite-step analysis distinguishes two cases. In a shared adapter, protecting the routing displacement leaves a learned-anchor residual through the changing companion factor. In a fresh zero-output block, a feasible routing state can protect the composed update while both current factors remain trainable. These conditions yield Semi-Frozen Orthogonal Routing (SFOR) for shared adapters and current-block hard protection for cumulative O-LoRA; Weight Residual Projection (WRP) enforces the required displacement after the optimizer step. Controlled two-task traces verify the predicted residual paths, reducing normalized historical response from 19.12% to 0.005% in the shared family and from 7.72% to 0.002% in the cumulative family. Four-task experiments on Qwen3-8B characterize the resulting trade-offs: SFOR improves backward transfer (BWT) from -2.47 to -0.86 with nearly unchanged average accuracy (AA), while O-LoRA hard protection improves three-order mean AA from 80.27% to 81.30% and forgetting measure (FM) from 2.20 to 0.43. Component controls also show that stricter feasibility need not improve final task performance. Together, the analysis and evidence provide an architecture-conditioned account of which constraint to enforce, how to enforce it, and how to interpret its empirical value.

Comments22 pages, 3 figures. Code is available at [https://github.com/legend91019/My_first](https://github.com/legend91019/My_first)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑