发表机构
Institute of Operations Research and Analytics; National University of Singapore; Wenzhou Buyi Pharmacy; College of Computer Science and Artificial Intelligence; Wenzhou University(运营研究与分析研究所; 新加坡国立大学; 温州布衣药房; 计算机科学与人工智能学院; 温州大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出AdmitOR这一无标注准入机制,通过跨家族团校准阈值实现模型接纳,在优化建模任务中显著提升接纳精度并减少污染接纳,为无标注场景下的模型准入提供了新方案。
AI 中文摘要
用于优化建模的经验学习智能体通过存储已验证的技能来提升性能,但现有学习器需通过与已知答案比对来确认知识,而实际工单流并不提供此类答案。自然的无标注替代方案不可靠:在包含300个问题的无标注工单流中,接纳每个可执行模型会导致约四分之一的接纳结果被污染;而单实例一致性规则会接受仅在一个值上匹配但其他方面存在差异的模型。本文提出AdmitOR,一种基于校准的外部行为证据构建的准入门限。来自三个模型家族、提示策略和求解器栈的候选模型在从提取的参数域重采样的实例上运行;所得价值函数轨迹的一致性通过跨家族团(cross-family clique)进行汇总,校准后的阈值返回接纳、弃权(不执行)或升级三种结果。预先注册的错误发现准则在校准数据上成立,但在实际工单流上不成立。我们完整报告这一负面结果,并将大多数失败归因于未忠实地编码其标注实例的基准文本。在最先进的技能学习器内部的一组日志上比较四种准入判断器,AdmitOR将接纳精度提升至0.927,而多数投票的精度为0.871,执行成功的精度为0.726,分别减少了3.1倍和8.0倍的污染接纳。其库规模最小,且在五个公共基准上达到最高的宏准确率,为58.4,而多数投票为54.8,真实标注库为53.9。与多数投票相比3.5个点的提升得到配对自助法的支持,且在主机端异常校正后依然成立。据我们所知,AdmitOR是首个围绕明确校准的错误发现目标设计的无标注准入机制。迁移失败确定了将其扩展至实际工单流的必要条件。
英文摘要
Agents that learn from experience improve at optimization modeling by storing solved trajectories and reusing them as skills. A wrong trajectory that enters the library can be retrieved again and again, and on a stream of new problems there is no ground-truth answer to decide with. Existing learners admit trajectories by matching known optima or labels, and label-free substitutes such as execution success or agreement at one instance can admit wrong models. We introduce ADMITOR, a label-free admission gate. It generates models from three model families, runs each on the stated problem and on instances with resampled parameters, keeps the largest group of models whose optimal values agree on every instance across families, and applies a threshold fitted on solver-verified problems to accept, abstain, or escalate, with a finite-sample bound on the false-discovery rate among accepted values. Inside a state-of-the-art skill learner, ADMITOR raises candidate-level admission precision to 0.927, against 0.871 for majority vote over the host's own samples and 0.726 for execution success, and its library, the smallest of the four, reaches the highest macro accuracy over five public benchmarks, 58.4 against 54.8 for majority vote. An ablation on the same records shows that the gain comes from the accepted value being external to the learner and unanimous across families; on this stream, resampling never changed an accepted value and only reduced coverage. The false-discovery bound holds on the calibration set but not on the benchmark stream: an audit of every false certificate traces most of them to benchmark texts that omit or round the numbers needed to reproduce the labeled answer, and a label-free check of the extracted numbers against the text flags most of these cases.
CommentsCode and data are available at https://github.com/junbolian/AdmitOR