对更新评分而非对令牌评分:面向组合LoRA专家的下降对齐路由
Score the Update, Not the Token: Descent-Aligned Routing for Combinatorial LoRA Experts
浏览论文内容
中文总结 AI 辅助
提出VANE路由方法,通过评分专家更新与下降方向的对齐而非令牌本身,在LoRA专家组合中实现更优性能,以更少参数超越现有基线。
中文摘要 AI 辅助
混合LoRA专家方法通过将每个令牌路由到少量低秩专家来提高低秩适配的容量。几乎所有此类方法都将每个专家的一个输入侧因子与一个输出侧因子绑定,并且几乎所有方法都通过对令牌进行评分来路由:路由器在不查看任何专家将写入什么的情况下选择专家。我们认为路由器应该对更新进行评分。在一阶近似下,将专家的更新添加到层输出会将该更新与输出处负损失梯度的内积降低损失。这种有用性在令牌中是二次的,因此对令牌线性的路由器只能看到通过令牌均值运行的部分,而根据自身激活范数对专家进行排名的路由器永远不会看到输出因子。如果每个专家被拆分为读取器(下投影)和写入器(上投影),则每个读取器-写入器对的有用性变为共享秩$r$空间中的内积,并且所有$N_AN_B$个对可以从$N_A+N_B$个向量进行评分。我们基于这一恒等式构建了VANE。一个低秩罗盘预测每个令牌的下降方向。VANE通过其更新与罗盘之间的对齐来对每个对进行评分,而不形成任何更新,以加性门激活前$k$个对,并为每个对提供其精确的一阶路由器梯度。在单领域常识推理和四领域多任务混合(使用Llama-3.2-3B和Llama-3.1-8B)中,VANE在十二种PEFT和MoE-LoRA基线中取得了最佳平均值,分别高出0.9-1.1和1.3-1.5个百分点,同时可训练参数少于8专家MoE-LoRA的一半。其路由器评分对专家实测有用性的追踪也远比令牌路由器更紧密。
英文摘要
Mixture-of-LoRA-experts methods raise the capacity of low-rank adaptation by routing each token to a few low-rank experts. Nearly all of them tie one input-side factor to one output-side factor per expert, and nearly all of them route by scoring the token: the router picks experts without seeing what any of them would write. We argue that the router should score the update. To first order, adding an expert's update to a layer output lowers the loss by the inner product between that update and the negative loss gradient at the output. This usefulness is quadratic in the token, so a router that is linear in the token sees only the part of it that runs through the token mean, and routers that rank experts by the norm of their own activations never see the output factor. If each expert is split into a reader (down-projection) and a writer (up-projection), the usefulness of every reader--writer pair becomes an inner product in the shared rank-$r$ space, and all $N_AN_B$ pairs can be scored from $N_A+N_B$ vectors. We build VANE on this identity. A low-rank compass predicts the descent direction of each token. VANE scores every pair by the alignment between its update and the compass without forming any update, activates the top-$k$ pairs with additive gates, and gives every pair its exact first-order router gradient. On single-domain commonsense reasoning and a four-domain multi-task mixture with Llama-3.2-3B and Llama-3.1-8B, VANE attains the best average among twelve PEFT and MoE-LoRA baselines, by 0.9--1.1 and 1.3--1.5 points respectively, with less than half the trainable parameters of an 8-expert MoE-LoRA. Its router scores also track the measured usefulness of experts far more closely than token routers do.
发表机构
- University of California, Irvine(加利福尼亚大学尔湾分校)
- University of Washington(华盛顿大学)
机构由 AI 辅助整理,请以论文原文为准。