基于偏好导向的多目标赌博机的选择性集成
Selective Ensemble Based on Preference-Directed Multi-Objective Bandits
- Department of Computer Science, Zhejiang Gongshang University(浙江工商大学计算机科学与技术学院)
- Center for Advanced Intelligence Project, RIKEN(日本理化学研究所先进智能研究中心)
- Graduate School of Frontier Sciences, The University of Tokyo(东京大学新领域创成科学研究科)
- National Key Laboratory for Novel Software Technology, Nanjing University(南京大学计算机软件新技术国家重点实验室)
- School of Artificial Intelligence, Nanjing University(南京大学人工智能学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对有限评估预算下部分偏好指定的选择性集成问题,形式化为偏好导向多目标赌博机,提出PrefUCB算法,实现对数遗憾界,在预训练模型集成和资产分配中验证有效性。
AI中文摘要:
现代机器学习系统的选择性集成需要在有限的评估预算下选择有前途的候选模型,而下游任务通常只指定对准确性、鲁棒性和推理等能力的部分偏好。这种设置自然导致在部分指定的线性偏好下的序列决策问题。我们将其形式化为偏好导向的多目标赌博机(PDMOB),其中可接受的权衡由多面体偏好锥表示。基于此公式,我们引入了帕累托$C$-最优性,它将标准帕累托最优性和单权重标量化作为特例恢复。然后,我们提出了偏好导向的上置信界(PrefUCB)算法,该算法维护方向置信区间以指导探索。我们分析了基于指标和基于间隙的遗憾,并为两个准则建立了实例相关的对数界,在经典特例中恢复了对时间$T$的最优对数依赖。在大型预训练模型选择性集成任务和机构授权下的在线资产分配上的实验验证了我们方法的有效性。
英文摘要:
Selective ensemble for modern machine learning systems requires choosing promising model candidates under limited evaluation budgets, while downstream tasks often specify only partial preferences over capabilities such as accuracy, robustness, and reasoning. This setting naturally gives rise to a sequential decision problem under partially specified linear preferences. We formalize it as preference-directed multi-objective bandits (PDMOB), where admissible trade-offs are represented by a polyhedral preference cone. Based on this formulation, we introduce Pareto $C$-optimality, which recovers standard Pareto optimality and single-weight scalarization as special cases. We then propose the preference-directed upper confidence bound (PrefUCB) algorithm, which maintains directional confidence intervals to guide exploration. We analyze both indicator-based and gap-weighted regret, and establish instance-dependent logarithmic bounds for both criteria, recovering the optimal logarithmic dependence on the horizon $T$ in classical special cases. Experiments on large pre-trained model selective ensemble tasks and online asset allocation under institutional mandates validate the efficacy of our method.