发表机构
University of Sydney; University of New South Wales(悉尼大学; 新南威尔士大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出梯度对齐路由(GAR),通过负载归一化目标函数按梯度对齐分组观测,在多任务文本分类中比仅任务损失路由准确率提升约1个百分点,并区分了负载均衡与专家专业化。
AI 中文摘要
稀疏专家模型可以在均匀分配流量的同时,仍将不兼容的训练信号分组到同一专家中。我们将路由视为一个梯度划分问题,并引入梯度对齐路由(GAR),其负载归一化的路由器目标函数奖励将具有对齐梯度的观测值分组。在五个多任务文本分类混合数据集上,我们将GAR与仅任务损失路由、梯度组合和梯度冲突方法以及负载均衡损失进行了比较。使用完全可训练的RoBERTa主干和分类头专家,GAR取得了最高的总体验证准确率,比仅任务损失路由高出1.07个百分点。使用冻结的DeBERTa和Qwen3-1.7B主干以及低秩适配器专家,GAR再次排名第一,比仅任务损失路由高出1.10个百分点,同时专家负载更均衡,梯度质量纯度更高(即每个专家梯度范数质量中来自其主导任务的比例);负载均衡损失进一步使负载更平坦,但使纯度保持在接近仅任务损失路由的水平。Top-1路由、可训练的全参数前馈网络(FFN)专家以及更大的主干也显示出正向的总体增益。结果区分了专家负载均衡与基于梯度的路由组织,并表明在多任务文本分类中,基于梯度的路由具有预测价值。
英文摘要
Sparse expert models can distribute traffic evenly while still grouping incompatible training signals within the same experts. We study routing as a gradient-partitioning problem and introduce gradient-aligned routing (GAR), whose load-normalized router objective rewards grouping observations with aligned gradients. On five multi-task text-classification mixtures, we compare GAR with task-loss-only routing, gradient-combination and gradient-conflict methods, and load-balancing losses. With a fully trainable RoBERTa backbone and classification-head experts, GAR has the highest aggregate validation accuracy, 1.07 percentage points above task-loss-only routing. With frozen DeBERTa and Qwen3-1.7B backbones and low-rank adapter experts, it again ranks first, 1.10 points above task-loss-only routing, with better-balanced expert load and higher gradient-mass purity, the share of each expert's gradient-norm mass from its dominant task; the load-balancing losses flatten load further but leave this purity near its task-loss-only level. Top-1 routing, trainable full-parameter feed-forward network (FFN) experts, and a larger backbone also show positive aggregate gains. The results distinguish expert-load balance from gradient-based routing organization and indicate the predictive value of gradient-informed routing in multi-task text classification.
Comments68 pages, 6 figures