发表机构
National University of Singapore; The Great Bay University(新加坡国立大学; 大湾区大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现有路由方法无法处理仅依赖最佳专家的奖励结构及操作约束的问题,提出OMD-Approachability方法,在真实众包数据集上验证了其性能。
AI 中文摘要
我们提出了多项子集路由(Multinomial Subset Routing,MSR),这是一种针对K个专家的新型在线路由框架,其中学习器维持的是多项路由策略,而非专家的确定性子集。每一轮中,学习器从多项策略中独立同分布地采样M个专家,采样得到的不同专家构成路由子集。奖励仅取决于路由子集中表现最佳的专家。这种奖励结构在跨专用模型路由场景中自然出现,但现有组合博弈或子集选择方法无法捕捉,后者优化确定性子集且通常假设奖励可加。我们要求选择满足博弈反馈下的若干长期双向操作约束,每轮仅观测获胜者的奖励。我们提出OMD-Approachability方法,将在线镜像下降与Blackwell可达性相结合,证明其在奖励和约束违反方面均达到O(1/√T)的悔界。我们将该框架应用于实际领域,并在真实众包数据集上进行了实证验证。
英文摘要
We study online routing to subsets of experts under aggregate bandit feedback and long-run operational constraints. Expert contributions can be complementary: for each task dimension, the best selected expert determines the contribution, and the total reward aggregates these dimension-wise maxima. At the same time, capacity, budget, and fairness requirements impose lower and upper bounds on long-run expert activation frequencies. A fixed deterministic subset cannot generally satisfy such heterogeneous requirements, motivating a stochastic routing policy. We formalize this problem as Multinomial Subset Routing (MSR). The learner maintains a distribution \(q\) over \(K\) experts, samples an expert independently \(M\) times from \(q\), and routes each task to the distinct sampled experts. We propose OMD-Approachability, which combines online mirror descent with Blackwell's approachability to optimize MSR under two-sided operational constraints. We establish \(O(T^{-1/2})\) average reward regret and expected constraint violation. We also quantify the approximation gap induced by a linear surrogate, and evaluate the approach on a real-world crowdsourcing dataset.