发表机构
Durham University Business School; Nanyang Business School, Nanyang Technological University; Faculty of Business and Economics, The University of Melbourne; International Institute of Finance, School of Management, University of Science and Technology of China(杜伦大学商学院; 南洋理工大学南洋商学院; 墨尔本大学商学院与经济学院; 中国科学技术大学管理学院国际金融研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对上下文优化中分离式学习与优化易致决策有效性下降的问题,提出决策驱动正则化框架,平衡预测精度与代价最小化,可推广SPO+,在合成研究中优于多个基准模型。
AI 中文摘要
在上下文优化中,决策者寻求最优决策以最小化随观测特征变化的代价函数,这类上下文常见于诸多商业应用,涵盖按需配送、零售运营、投资组合优化及库存管理等领域。本文研究学习与优化方法,该方法先学习结果如何由特征产生,再基于这些结果选择最优决策。我们聚焦集成学习与优化文献,发现对预测精度缺乏控制会导致过拟合,且相较于简单的分离式学习与优化模型,决策有效性会下降。为此,我们提出一种平衡预测精度与代价最小化的双目标公式,称为决策驱动正则化,它还通过依赖新超参数的代理解决了代价函数定义的模糊性。我们进一步表明,该问题的其他表述视角,即鲁棒优化与后悔最小化,会产生与我们提出的模型密切相关的模型,因此我们的框架可推广SPO+等模型。在合成研究中,我们的模型在数值上优于OLS、随机森林、XGBoost、SPO+、扰动梯度及Learning and Rank等基准。
英文摘要
In contextual optimization, the decision-maker seeks optimal decisions to minimize a cost function, that varies based on observed features. This context is common in many business applications ranging from on-demand delivery and retail operations to portfolio optimization and inventory management. In this paper, we study the learning and optimization approach, which first learns how outcomes result from the features, and then selects optimal decisions based on these outcomes. We focus on the integrated learning and optimization literature, and identify that a lack of control for prediction accuracy can lead to overfitting and a loss of decision effectiveness against simple separate learning and optimization models. Instead, we propose a bi-objective formulation that balances prediction accuracy and cost minimization, termed decision-driven regularization. It also addresses ambiguity in the definition of the cost function via a surrogate that depends on a new hyperparameter. We additionally show that alternative perspectives for formulating the problem, namely robust optimization and regret minimization, lead to models that are closely related to our proposed model. As a consequence, our framework generalizes models such as SPO+. Our model is shown to be numerically superior to other benchmarks, such as OLS, Random Forest, XGBoost, SPO+, Perturbation Gradient, and Learning and Rank, in our synthetic studies.
Comments42 pages (including appendix), 7 figures in main, journal paper