特征值分解成本去噪:最短路径问题中预测后优化的一种替代方法
Eigenvalue-Decomposition Cost Denoising as an Alternative to Predict-then-Optimize for Shortest-Path Problems
浏览论文内容
中文总结 AI 辅助
针对预测后优化在模型设定错误下性能下降的问题,提出利用协方差矩阵特征值分解直接去噪含噪成本向量,在网格最短路径基准上以k=5匹配潜在维度,使去噪Dijkstra优于SPO+。
中文摘要 AI 辅助
预测后优化方法,如Elmachtoub和Grigas(2022)提出的智能“预测后优化”(SPO+),学习从上下文特征到未知边成本的映射,然后在预测成本上求解由此产生的组合问题。这种方法功能强大,但依赖于预测模型被正确设定:当真实成本生成过程在特征上是非线性的而预测器是线性的时,SPO+的性能会随着设定错误程度的增加而下降。我们针对一个特定但常见的场景提出并评估了一种结构上不同的补救措施:当决策者观察到同一底层成本过程的许多含噪实现时,实现成本向量本身可以被视为含噪信号,并在调用任何预测模型之前,通过其协方差矩阵的特征值分解(等价于主成分分析)直接进行去噪。我们在Elmachtoub和Grigas(2022)引入的$5\ imes5$网格最短路径基准上实例化了这一想法,仅保留训练成本协方差矩阵的前$k$个特征向量,并在使用Dijkstra(1959)算法求解之前,将新的含噪成本观测投影到该子空间上。我们发现$k$的选择至关重要:仅保留$k{=}2$个特征向量会丢弃真实信号,甚至不如朴素的含噪成本基线;而设置$k{=}5$以匹配真实潜在特征维度,使得特征值去噪Dijkstra在测试的每个设定错误水平下都是表现最好的方法,在高设定错误下大幅优于SPO+。
英文摘要
Predict-then-optimize methods such as Smart "Predict, then Optimize" (SPO+) of Elmachtoub and Grigas (2022) learn a mapping from contextual features to unknown edge costs and then solve the induced combinatorial problem on the predicted costs. This approach is powerful but relies on the predictive model being well specified: when the true cost-generating process is nonlinear in the features and the predictor is linear, SPO+'s performance degrades as the misspecification grows. We propose and evaluate a structurally different remedy for a specific but common setting: when the decision-maker observes many noisy realizations of the same underlying cost process, the realized cost vectors themselves can be treated as a noisy signal and denoised directly, via eigenvalue decomposition (equivalently, Principal Component Analysis) of their covariance matrix, before ever invoking a predictive model. We instantiate this idea on the $5\times5$ grid shortest-path benchmark introduced by Elmachtoub and Grigas (2022), retaining only the top-$k$ eigenvectors of the training cost covariance matrix and projecting new noisy cost observations onto that subspace prior to solving with Dijkstra's (1959) algorithm. We find that the choice of $k$ is decisive: keeping only $k{=}2$ eigenvectors discards real signal and underperforms even the naive noisy-cost baseline, while setting $k{=}5$ to match the true latent feature dimension makes eigenvalue-denoised Dijkstra the best-performing method at every misspecification level tested, outperforming SPO+ by a wide margin under high misspecification.