可替代产品的可扩展动态定价:基于结构引导的策略学习
Scalable Dynamic Pricing of Substitutable Products through Structure-Guided Policy Learning
AI总结:
针对可替代产品动态定价中动态规划难以扩展的问题,提出两种MNL引导的策略学习方法,通过决策聚焦学习训练,在小规模实例上最优性差距低于0.4%,在大规模实例上优于基准,并揭示了模型结构对性能的影响。
AI中文摘要:
问题定义:我们研究具有有限、产品特定库存的可替代产品的动态定价问题。顾客替代行为使得不同产品的定价决策相互耦合,而库存状态使得精确动态规划在现实规模下难以处理。方法/结果:我们开发了两种MNL引导的策略学习方法,用从库存状态到定价决策的统计映射替代动态规划。第一种方法直接学习价格,第二种方法学习库存机会成本并利用最优MNL定价规则将其转化为价格。两种策略均采用决策聚焦学习进行训练。我们还提出了一种高效的方法来生成前瞻性顾客选择目标以训练我们的策略。管理启示:我们对所提出的方法进行了广泛的数值评估。在可以计算最优动态规划的小规模实例上,学习策略的平均最优性差距低于0.4%。在更大规模的以航空为背景的实例上,它们持续优于所测试的收益管理基准。两种架构之间的比较也凸显了模型结构的作用:当需求由MNL模型良好描述时,使用MNL定价特征特别有效,而在异构混合MNL需求下,直接学习价格提供了更大的灵活性。
英文摘要:
Problem definition: We study dynamic pricing of substitutable products with finite, product-specific inventories. Customer substitution couples pricing decisions across products, while the inventory state makes exact dynamic programming intractable at realistic scale. Methodology / results: We develop two MNL-guided policy-learning approaches that replace the dynamic program with a statistical mapping from inventory states to pricing decisions. The first learns prices directly, while the second learns inventory opportunity costs and converts them into prices using the optimal MNL pricing rule. Both policies are trained using decision-focused learning. We also propose an efficient method to generate anticipative customer-choice targets to train our policies. Managerial implications: We conduct an extensive numerical evaluation of our approaches. On small instances for which the optimal dynamic program can be computed, the learned policies achieve average optimality gaps below 0.4%. On larger airline-motivated instances, they consistently improve on the tested revenue-management benchmarks. The comparison between the two architectures also highlights the role of model structure: using the MNL pricing characterization is particularly effective when demand is well described by MNL, while directly learning prices provides greater flexibility under heterogeneous mixed-MNL demand.