采样运气伪装成分配增益:神经组合优化的测试时预算分配审计
Sampling Luck Masquerades as Allocation Gain: Auditing Test-Time Budget Allocation for Neural Combinatorial Optimization
浏览论文内容
中文总结 AI 辅助
该研究审计神经组合优化测试时预算分配的增益,发现样本内存在幻象增益,样本外消失,而分布偏移下分配可提升性能,还提供校正流程并发布相关数据代码。
中文摘要 AI 辅助
神经组合优化(NCO)求解器针对每个实例报告多个采样解中的最优解,按惯例每个实例的采样数量相同。固定总预算的非均匀分配是否能带来收益尚未被衡量,我们对此进行了衡量并审计了衡量过程。首先,在分布内工作负载中,分配的增益空间无法检测:在均匀TSP-100上的三个预训练求解器(POMO、AM、SymNCO)中,基于相同存储样本计算和评估的最优分配报告了2.2-2.6%的增益,且置信区间不包含零;而在样本外测量时,相同的增益与零无差异,分别为0.457%、0.015%、-0.512%。遵循常规的样本内流程,所有三个求解器都会支持已发表的2%级别的不存在的增益。我们针对构造的真实增益为零的实例级零假设校准了这一偏差,在我们测试的范围内,该偏差不会随更多样本或更多实例而缩小。其次,消除幻象增益的校正方法保留了真实增益:在分布偏移(混合均匀和聚类实例的工作负载)下,一项预先注册的验证实验发现,在相同评估预算下,由保留样本统计量引导的分配将k个最优解的性能提升了11.5%(AM,主要终点;95%置信区间[7.4, 19.7])和12.0%(SymNCO,重复实验),且未计入信号获取成本;预先注册的阴性对照(POMO,对偏移的鲁棒性高一个数量级)显示为-0.3%[-0.7, 0.24],该增益超过冻结分布标签基线4.2个百分点[1.9, 7.7];一项探索性策略对相同预算收取20个样本的探测成本后,仍保留了3.4%(AM)和4.6%(SymNCO)的增益。我们提供了校正流程和报告清单,并发布了所有数据、代码和预注册记录。
英文摘要
Neural combinatorial optimization (NCO) solvers report the best of many sampled solutions per instance, and the sample count is, by convention, identical for every instance. Whether a non-uniform allocation of a fixed total budget would buy anything has not been measured. We measure it, and we audit the measurement itself. First, on in-distribution workloads the allocation headroom is not detectable. Across three pretrained solvers (POMO, AM, SymNCO) on uniform TSP-100, an oracle allocation computed and evaluated on the same stored samples reports a 2.2-2.6% gain with intervals excluding zero; measured out of sample the same gain is indistinguishable from zero (0.457, 0.015, -0.512 percent). Following the customary in-sample procedure, all three solvers would have supported a published 2%-level gain that does not exist. We calibrate this bias against an instance-wise null in which the true gain is zero by construction; over the ranges we test it does not shrink with more samples or more instances. Second, the same correction that removes the phantom gains preserves a real one. Under distribution shift (a workload mixing uniform and clustered instances), a pre-registered confirmatory experiment finds that allocation guided by held-out sample statistics improves best-of-k by 11.5% (AM, primary endpoint; 95% CI [7.4, 19.7]) and 12.0% (SymNCO, replication) at equal evaluation budget, with the signal-acquisition cost not charged; a pre-registered negative control (POMO, an order of magnitude more robust to shift) shows -0.3% [-0.7, 0.24]. The gain exceeds a frozen distribution-label baseline by 4.2 points [1.9, 7.7]. An exploratory policy charging a 20-sample probe against the same budget retains 3.4% (AM) and 4.6% (SymNCO). We give a correction procedure and a reporting checklist, and release all data, code, and the pre-registration record.
发表机构
- Hankuk University of Foreign Studies(韩国外国语大学)
机构由 AI 辅助整理,请以论文原文为准。