arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.32465cs.LG

PolyStepOR:无需最优决策的学习决策

PolyStepOR: Learning to Decide Without Optimal Decisions

Viet The Nguyen, Gunther Gust, An Thai Le

首次发表
浏览论文内容

中文总结 AI 辅助

PolyStepOR提出一种无需最优决策和导数的决策聚焦学习方法,通过参数扰动和最优传输优化分段常数损失,在多个基准上表现强劲。

中文摘要 AI 辅助

决策聚焦学习(DFL)训练预测器以提升下游决策质量,但通常依赖难以获取的最优参考决策。我们提出PolyStepOR,它直接从已实现的决策成本中训练,无需预先计算的最优解,并通过修复或不可行性惩罚扩展到约束内预测。为处理分段常数损失,PolyStepOR扰动预测器参数,评估由此产生的决策,并利用最优传输偏向低成本方向,无需导数。无需任务特定调参,PolyStepOR在经典优化基准上表现强劲,在预测约束和现实世界问题上具有竞争力。理论上,我们刻画了保持决策的扰动和边界检测,界定了对成本误差的敏感性,并为平滑目标建立了驻点保证。因此,PolyStepOR用前向评估取代了最优参考决策和导数。

英文摘要

Decision-focused learning (DFL) trains predictors for downstream decision quality, but often relies on optimal reference decisions that are expensive to obtain. We present PolyStepOR, which trains directly from realized decision costs without pre-computed optima and extends to in-constraint predictions through repair or infeasibility penalties. To handle piecewise-constant losses, PolyStepOR perturbs predictor parameters, evaluates the resulting decisions, and uses optimal transport to favor lower-cost directions, requiring no derivatives. Without task-specific tuning, PolyStepOR performs strongly on classical optimization benchmarks and competitively on predicted-constraint and real-world problems. Theoretically, we characterize decision-preserving perturbations and boundary detection, bound sensitivity to cost errors, and establish stationarity guarantees for a smoothed objective. PolyStepOR thus replaces optimal reference decisions and derivatives with forward evaluations.

↑