通过勒让德正则化策略进行具有硬约束的平滑学习
Smooth Learning with Hard Constraints via Legendre-Regularized Policies
- The Chinese University of Hong Kong(香港中文大学)
- The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
- Tongji University(同济大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究从策略类设计角度重审上下文优化,提出勒让德正则化策略,将决策参数化,证明其相关优化器映射的良好性质,建立通用逼近结果并统一了优化器,实验表明该方法能提升规定性性能。
AI中文摘要:
我们从策略类设计的角度重新审视上下文优化。一个理想的策略类应该有足够的表现力来学习丰富的上下文决策关系,应该强制执行硬可行性约束而不是软惩罚项,并且应该足够平滑以便对下游决策损失进行基于梯度的训练。现有方法通常只强调这些要求的一部分。我们提出勒让德正则化策略,它将决策参数化为在原始可行区域上的正则化优化问题的解。这种构造产生的策略在构造上是可行的,并且相对于学习到的潜在参数是可微的。我们证明相关的优化器映射是单值的,映射到可行集的相对内部,允许显式雅可比矩阵,是利普希茨连续的,并且可以任意平滑。我们还建立了一个通用逼近结果,表明所提出的类可以逼近紧凑上下文集上的任何连续可行策略。该框架统一了显式正则化优化器和基于隐式扰动的平滑优化器。在上下文报童和资源分配问题上的实验表明,我们的方法相对于基准方法提高了规定性性能。
英文摘要:
We revisit contextual optimization from the perspective of policy class design. A desirable policy class should be expressive enough to learn rich context-decision relationships, should enforce hard feasibility constraints rather than soft penalty terms, and should remain smooth enough for gradient-based training on downstream decision losses. Existing approaches usually emphasize only part of these requirements. We propose Legendre-regularized policies, which parameterize decisions as solutions of regularized optimization problems over the original feasible region. This construction yields policies that are feasible by construction and differentiable with respect to learned latent parameters. We prove that the associated optimizer map is single-valued, maps onto the relative interior of the feasible set, admits an explicit Jacobian, is Lipschitz continuous, and can be made arbitrarily smooth. We also establish a universal approximation result showing that the proposed class can approximate any continuous feasible policy on compact context sets. The framework unifies explicitly regularized optimizers and implicit perturbation-based smooth optimizers. Experiments on contextual newsvendor and resource allocation problems show that our approach improves prescriptive performance relative to the benchmark methods.