发表机构
The George Washington University; Purdue University(乔治华盛顿大学; 普渡大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出LNOQRD方法,通过重塑动作空间减少网络控制优化的候选量,在小实例降75.9%候选、大实例性能优于基线,兼顾效率与效果。
AI 中文摘要
现代网络策略控制将意图映射为顺序放置控制决策。Bellman式策略优化主要关注应优化哪个动作,而约束通常通过惩罚、障碍或拉格朗日机制处理。我们发现,在值函数能验证最佳部署之前,中间信号可能已识别出许多应排除在进一步优化之外的候选对象。这催生了一个互补方向:\textbf{学习不优化}。在值函数足够准确以选择最佳放置控制决策之前,中间信号可能已显示候选对象在状态-意图重标记(商化)下等价、导致未来状态一致更差(支配性)或违反可执行网络定律(残差筛选)。\textbf{LNOQRD}将这些计算或学习到的信号作为影子过程,重塑执行原始策略优化的域,从而减小动作空间。我们证明了在明确等变性和单调性条件下的无损商化与支配性,界定了边界大小和排序成本,并量化了近似证书和原始估计带来的损失。实验表明,\textbf{LNOQRD}可将小实例候选对象减少75.9%,同时保留90.8%的近神谕覆盖率;在大实例上,它实现了最高的效用和意图满意度、最低的硬定律违反率与后生成延迟,且在基于候选对象的基线中平均减少73.0%。
英文摘要
Modern network policy control maps intent to sequential placement-control decisions. Bellman-style policy optimization primarily asks which action to optimize, while constraints are commonly handled through penalty, barrier, or Lagrangian mechanisms. We observe that before a value function can certify the best deployment, intermediate signals may already identify many candidates that should be excluded from further optimization. This motivates a complementary direction: \emph{Learning Not to Optimize}. Before a value function is accurate enough to select the best placement-control decision, intermediate signals may already show that candidates are equivalent under state--intent relabeling (quotienting), lead to a uniformly worse future state (dominance), or violate executable network laws (residual screening). \LNOQRD{} uses these computed or learned signals as a shadow process to reshape the domain on which primal policy optimization is performed, thereby reducing the action space. We prove lossless quotienting and dominance under explicit equivariance and monotonicity conditions, bound frontier size and ranking cost, and quantify losses from approximate certificates and primal estimates. Experiments show that \LNOQRD{} reduces small-instance candidates by $75.9\%$ while retaining $90.8\%$ near-oracle coverage and, on large instances, achieves the highest utility and intent satisfaction, the lowest hard-law violation and post-generation latency, and a $73.0\%$ average reduction among candidate-based baselines.