arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

神经组合优化的稳定K选最优训练

Stabilized Best-of-$K$ Training for Neural Combinatorial Optimization

Melveena Jolly, Midhun Xavier

arXiv 2608.00296首次发表:更新:

AI 中文总结

该研究针对神经组合优化,将POMO训练的二元领导者区分替换为稳定排名信号,在TSP-100测试集上验证了稳定K=8方案可降低K选8成本,且不宣称其为无偏估计器或最优方法。

AI 中文摘要

Leader Reward(领导者奖励)修改了POMO训练,以强调多次推理产生的最优轨迹。我们测试了一个窄范围扩展:用由采样预算K索引的稳定排名信号取代其二元领导者/非领导者区分。在固定POMO架构、3050个epoch的调度和TSP-100测试集的情况下,重新实现的Leader Reward在100次启动、8次增强的贪婪解码下取得了7.7662,与报告的7.766精度一致。在独立采样下,稳定的K=8方案在所有三个配对训练种子中降低了实际K选8的成本:7.7944对7.8136。该观察仅针对估计和特定解码器:三个种子低于六个种子的测试下限,Leader Reward在采样K=1时表现更好,在其原始增强贪婪协议下仍略好。我们不提出无偏估计器、普遍优越性或最先进的主张。

英文摘要

Leader Reward modifies POMO training to emphasize the best trajectory produced by repeated inference. We test a narrow extension: replace its binary leader/non-leader distinction with a stabilized rank signal indexed by a sampling budget $K$. With the POMO architecture, 3,050-epoch schedule, and TSP-100 test set held fixed, the Leader Reward reimplementation obtains $7.7662$ under 100-start, 8-augmentation greedy decoding, matching the reported $7.766$ at its displayed precision. Under independent sampling, the stabilized $K=8$ recipe lowers realized Best-of-8 cost in all three paired training seeds: $7.7944$ versus $7.8136$. This observation is estimation-only and decoder-specific: three seeds are below the six-seed testing floor, Leader Reward is better at sampled $K=1$, and it remains slightly better under its original augmented-greedy protocol. We make no unbiased-estimator, universal superiority, or state-of-the-art claim.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑