arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过对偶价格求导:容量约束下的端到端策略学习

Differentiating Through Dual Prices: End-to-End Policy Learning Under Capacity Constraints

Mohammadsaeed Haghi, Mahdi Salmani, Nima Kelidari

arXiv 2608.04669首次发表:更新:

AI 中文总结

本文针对容量约束下的资源分配问题,提出通过对偶价格求导的端到端策略学习方法,在多组数据集上验证其性能优于传统无决策方法,更适配资源稀缺且需保证可行性的场景。

AI 中文摘要

许多社会服务会将住房援助、医院干预等稀缺资源分配给依次到达的个体:每个到达者必须立即获得决策,且每种资源的长期使用量必须保持在其容量范围内。我们研究如何从记录的观测数据中学习此类分配策略。标准流程是无决策的:通过回归为每个“臂”(arm)拟合一个结果模型,根据拟合模型对每种有容量限制的资源定价,再将到达者分配给“预测结果减去价格”最大的臂。我们转而端到端地训练结果模型,通过对偶价格本身对部署策略价值的离策略估计进行求导。我们研究两种形式:一种是精确的非凸形式,另一种是凸松弛形式,其最优解总能在期望上满足容量约束,且次优性最多为与平滑温度线性相关、与臂数对数相关的项。所有方法均在队列模拟中评估,资源以其容量速率补充。在六个数据集上,两种端到端变体在每个延迟成本(包括零延迟成本)下的经部署调整价值指数中均占据前列;当容量具有约束力时,无决策基线经常违反容量约束并导致长得多的队列延迟。在最大的数据集——包含7万名患者的医院队列中,端到端训练还实现了显著更高的策略价值,该优势在容量匹配的神经基线中依然存在。当真实值可测量时,灵活的无决策回归仍是更强的纯预测器;端到端训练最适合资源真正稀缺且可行性至关重要的场景。

英文摘要

Many social services assign scarce resources, such as housing assistance or hospital interventions, to people who arrive one at a time: each arrival must receive a decision immediately, and the long-run usage of every resource must stay within its capacity. We study how to learn such an assignment policy from logged observational data. The standard pipeline is decision-blind: fit one outcome model per arm by regression, price each capacitated resource from the fitted models, and assign each arrival the arm whose predicted outcome minus price is largest. We instead train the outcome models end-to-end, differentiating an off-policy estimate of the deployed policy's value through the dual prices themselves. We study two formulations: an exact nonconvex one, and a convex relaxation whose optimum always satisfies the capacity constraints in expectation and which is suboptimal by at most a term linear in the smoothing temperature and logarithmic in the number of arms. Every method is evaluated in a queueing simulation with resources replenished at their capacity rates. Across six datasets, the two end-to-end variants take the top slots on a deployment-adjusted value index at every delay cost, including zero; when capacities are binding, decision-blind baselines frequently violate them and incur much longer queueing delays. On the largest dataset, a hospital cohort of seventy thousand patients, end-to-end training also achieves significantly higher policy value, a margin that survives a capacity-matched neural baseline. Flexible decision-blind regression remains the stronger pure predictor where ground truth is measurable; end-to-end training is best suited to settings where resources are genuinely scarce and feasibility matters.

Comments15 pages, 7 figures, 2 algorithms. Includes a technical appendix with full proofs, an excess-value decomposition, ablations, and reproducibility details. Code: https://github.com/mahdisalmani/end2end-capacity-constrained-policy-learning

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑