发表机构
Dongyang Mirae University(东阳未来大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对混合专家模型的路由诱导偏差问题,提出子组重加权与门控熵正则化结合的端到端框架,在提升公平性的同时维持了预测性能。
AI 中文摘要
深度学习模型常因性别、年龄等敏感属性相关的训练数据不平衡,在不同人口统计子组间产生性能差异。现有研究探索了公平表示学习、数据重采样、对抗训练等方法,可大致分为两类:单阶段方法通常学习用于公平性的共享表示,但往往难以处理异质子组分布;两阶段方法将表示学习与最终预测任务分开进行,可能导致公平性目标与下游预测失配。本文识别出路由诱导偏差这一失效模式——子组不平衡会驱使门控网络将子组路由至少数专家,进而提出端到端的混合专家(Mixture-of-Experts,MoE)框架以纠正该问题。具体而言,本文采用子组重加权修正数据不平衡,引入门控熵正则化防止路由坍缩至子组属性,使专家利用率既平衡又可解释。除提升公平性外,路由分布还提供了子组在各专家间分配的可解释视图。实验结果表明,所提方法在提升公平性的同时,保持了具有竞争力的预测性能。
英文摘要
Deep learning models often produce performance disparities across demographic groups, due to the training data imbalance with respect to sensitive attributes such as gender or age. To address this problem, existing work has explored fair representation learning, data re-sampling, and adversarial training, which can be broadly categorized into two main approaches. Single-stage methods typically learn a shared representation for fairness, but often struggle to handle heterogeneous subgroup distributions. Two-stage methods learn representations separately from the final prediction task, which can lead to misalignment between fairness objectives and downstream predictions. We identify routing-induced bias, a failure mode in which subgroup imbalance drives the gating network to route subgroups onto a few experts, and propose an end-to-end Mixture-of-Experts (MoE) framework that corrects it. Specifically, we apply subgroup reweighting to correct data imbalance, and introduce gate entropy regularization to prevent routing from collapsing onto subgroup attributes, keeping expert utilization both balanced and interpretable. Beyond improving fairness, the routing distribution offers an interpretable view of how subgroups are allocated across experts. Experimental results demonstrate that the proposed approach improves fairness while maintaining competitive predictive performance.
Journal refAVSS 2026 (22nd International Conference on Advanced Visual and Signal-Based Systems)