在不确定的地方投入专家:用于专家混合LoRA的置信度自适应路由
Spend Experts Where You Are Unsure: Confidence-Adaptive Routing for Mixture-of-Experts LoRA
浏览论文内容
中文总结 AI 辅助
研究针对专家混合LoRA路由中简单与困难令牌服务不均问题,提出置信度自适应路由CARE,它以核心方式接纳专家,校准阈值使活跃专家数匹配目标,是单前向传递规则,在多任务中改进效果显著,还提升了分布外检测能力。
中文摘要 AI 辅助
低秩适应(LoRA)的专家混合(MoE)变体将每个令牌路由到固定数量的k个专家。令牌在模型对其的不确定程度上有所不同,因此单一的k在简单令牌上花费过多,而对困难令牌服务不足。我们观察到路由器的输出分布已经是一个逐令牌的不确定性信号:峰值质量表示置信度,而平坦分布表示模糊性。我们引入了CARE(专家的置信度自适应路由),它以核心方式接纳专家。专家按路由器权重递减被激活,直到其累积质量达到阈值,当被接纳的专家意见不一致时会有小的扩展。一个预算恒温器校准阈值,使活跃专家的平均数量匹配任何目标。CARE是一个无需额外参数的即插即用单前向传递规则。在关于LLaMA - 3.1 - 8B和Qwen2.5 - 7B的八个常识基准测试以及数学、代码和知识任务中,CARE在匹配计算时比固定的top - k MoE - LoRA有所改进,并且在激活更少专家的情况下与固定k = 4基线匹配。相同的置信度和不一致信号也比MSP、熵和多通道代理改进了分布外检测。我们用核心保真度、预算最优性和对不一致的认知解读来支持该设计,并发布了代码。
英文摘要
Mixture-of-Experts (MoE) variants of Low-Rank Adaptation (LoRA) route every token to a fixed number of experts $k$. Tokens differ in how uncertain the model is about them, so a single k over-spends on easy tokens and under-serves hard ones. We observe that the router's output distribution is already a per-token uncertainty signal: peaked mass indicates confidence, while a flat distribution indicates ambiguity. We introduce CARE (Confidence-Adaptive Routing of Experts), which admits experts in a nucleus fashion. Experts are activated in decreasing router weight until their cumulative mass reaches a threshold, with a small extension when the admitted experts disagree. A budget thermostat calibrates the threshold so that the average number of active experts matches any target. CARE is a drop-in, single-forward-pass rule with no extra parameters. Across eight commonsense benchmarks on LLaMA-3.1-8B and Qwen2.5-7B, as well as math, code, and knowledge tasks, CARE improves over fixed top-k MoE-LoRA at matched compute and matches the fixed-k=4 baseline while activating fewer experts. The same confidence and disagreement signals also improve out-of-distribution detection over MSP, entropy, and multi-pass proxies. We support the design with nucleus fidelity, budget optimality, and an epistemic reading of disagreement, and we release code.
发表机构
- University of California, Irvine(加州大学欧文分校)
- University of Washington(华盛顿大学)
机构由 AI 辅助整理,请以论文原文为准。