发表机构
Hankuk University of Foreign Studies(韩国外国语大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文通过预注册实验,在旅行商问题上研究测试时预算重新分配的决定因素,发现实例难度差异而非平均难度是关键,且预算感知策略可恢复大部分改进。
AI 中文摘要
神经组合优化求解器为每个实例生成大量候选解,并报告找到的最佳解,无论实例难度如何,每个实例都使用相同的采样预算。一项配套研究表明,将固定预算重新分配给较难的实例可以提高解的质量,但衡量这种改进的标准方式存在偏差:在同一数据上决定分配并评估该分配,即使实际上没有改进,也可能制造出表面上的收益。这留下了两个未解决的问题:工作负载的什么属性决定了重新分配是否值得,以及一个花费部分预算来决定如何分配剩余预算的策略,在计入该成本后是否仍然划算。\n本文通过预注册的验证性实验——在数据收集之前固定分析和判定标准——回答了这两个问题,实验涉及三个独立训练的求解器和两种在旅行商问题上构建更难工作负载的方式。在我们研究的工作负载中,决定性属性是工作负载内实例在难度上的差异程度,而非工作负载的平均难度:均匀容易或均匀困难的工作负载为重新分配提供的空间很小,而混合工作负载则提供了相当大的空间。一个具有预算意识的策略,为其自身关于实例难度的信息付费,可以恢复大部分(尽管不是全部)在假设该信息免费时可获得的改进。\n每个实验都根据其书面规范独立重新计算,并且对早期版本的每一次修正——包括两次削弱了论文自身主张的修正——都报告了其使结论移动的方向。本文提供了一个具体的实证答案,以及一个验证该答案不是测量方式产物的模板。
英文摘要
Neural combinatorial optimization solvers generate many candidate solutions per instance and report the best one found, using the same sample budget for every instance regardless of difficulty. A companion study showed that reallocating a fixed budget toward harder instances can improve solution quality, but that the standard way of measuring this improvement is biased: deciding an allocation and evaluating it on the same data can manufacture an apparent gain even when none exists. This left open what property of a workload determines whether reallocation is worth doing, and whether a policy that spends part of the budget to decide how to allocate the rest still pays once that cost is counted. This paper answers both questions through pre-registered confirmatory experiments -- analysis and verdict criteria fixed before data collection -- across three independently trained solvers and two ways of constructing harder workloads on the traveling salesman problem. Within the workloads we study, the deciding property is how varied the instances within a workload are in difficulty, not how difficult the workload is on average: a uniformly easy or uniformly hard workload offers little room for reallocation, while a mixed workload offers substantial room. A budget-aware policy that pays for its own information about instance difficulty recovers most, though not all, of the improvement available when that information is assumed free. Every experiment was independently recomputed from its written specification, and every correction to an earlier version -- including two that weakened the paper's own claims -- is reported with the direction it moved the conclusion. The paper offers a specific empirical answer and a template for verifying that answer is not an artifact of how it was measured.
Comments19 pages, 5 figures. Code, cost arrays, and the full pre-registration record (including every amendment and its direction) at https://github.com/nepersoned/best-of-k-allocation