AI 中文总结
TEXAS 是一种 MoE 大语言模型适配方法,通过对比模型成功与失败实例的专家激活发现任务专家,在微调时提升对应答案 token 权重,在多数基准测试中性能优于基线。
AI 中文摘要
混合专家(MoE)语言模型将每个 token 路由至一小部分专家,这一路由模式可用于在下游适配过程中识别任务相关专家。但现有方法存在两个局限:任务专家通常从反映使用情况的聚合路由统计数据中识别,而非与成功完成任务的关联;任务专家激活作为监督分配信号的作用仍未得到充分探索。我们提出任务专家感知监督方法(TEXAS),它结合了正确性条件下的任务专家发现与 token 级监督分配。TEXAS 对比基础模型成功解决和未解决的实例上的专家激活情况,保留在成功实例中激活更强的专家。在微调过程中,当这些专家被激活时,TEXAS 会提升未解决实例中答案 token 的权重。因此,TEXAS 利用现有路由行为,无需将适配限制在固定专家子集或施加显式目标路由分布。在三个 MoE 模型和六个基准测试中,TEXAS 在 18 个设置中的 17 个取得最佳或并列最佳性能,且比最强基线平均提升 1.3-1.5 个百分点。消融实验和进一步分析验证了所发现的专家及由此产生的监督策略。
英文摘要
Mixture-of-Experts (MoE) language models route each token through a small subset of experts, making routing patterns useful for identifying task-relevant experts during downstream adaptation. Yet current approaches have two limitations: task experts are typically identified from aggregate routing statistics that reflect usage rather than association with successful task completion, and task-expert activations remain underexplored as signals for supervision allocation. We introduce Task-Expert-Aware Supervision (TEXAS), which combines correctness-conditioned task expert discovery with token-level supervision allocation. TEXAS compares expert activations on instances that the base model solves successfully and those it fails to solve, and retains experts more strongly activated on successful instances. During fine-tuning, it upweights answer tokens in failed instances when they activate these experts. TEXAS therefore leverages existing routing behavior without restricting adaptation to a fixed expert subset or imposing an explicit target routing distribution. Across three MoE models and six benchmarks, TEXAS achieves the best or tied-best performance in 17 of 18 settings and improves over the strongest baseline by 1.3--1.5 points on average. Ablations and further analyses validate both the discovered experts and the resulting supervision strategy.