超越已验证答案:求解器引导的自蒸馏用于引导运筹学语言模型
Beyond Verified Answers: Solver-Informed Self-Distillation for Bootstrapping Operations Research Language Models
浏览论文内容
中文总结 AI 辅助
针对运筹学语言模型训练中监督受限、信用分配粗糙和风格不匹配的问题,提出SOLID框架,利用求解器工件和自蒸馏实现无需验证答案的自我改进,提升解决方案准确性。
中文摘要 AI 辅助
现代大型语言模型(LLM)能够将自然语言描述转化为运筹学(OR)公式。包括强化学习和同策略自蒸馏在内的后训练技术进一步提升了这一能力。然而,在训练用于OR公式的LLM时仍存在三个局限性。首先,训练通常依赖于由人类专家或更强模型验证的合成公式,这限制了可扩展的监督。其次,信用分配要么粗糙要么昂贵:结果奖励对整个轨迹进行评分,而不定位负责的建模决策,而过程级监督则需要额外的评估器。第三,特权自蒸馏可能因使用部署时不可用的求解器上下文而导致风格不匹配。我们发现,模型可以从自身回滚生成的求解器工件反馈中改进,这使得自蒸馏成为一种实用的、无需评估器的密集监督来源。因此,我们提出SOLID:通过自蒸馏进行求解器引导的同策略学习,这是一种新颖的框架,用于在没有验证答案或外部评估器的情况下自我改进OR语言模型。SOLID执行来自多次回滚的候选程序,聚类其目标,并选择多数组的工件作为伪参考。然后,模型使用组相对优势和密集的自监督信号进行更新。在多个OR基准上,SOLID相对于仅结果组相对训练,提高了通用和OR调优模型的解决方案准确性。这些结果表明,求解器工件可以在没有可信答案的情况下支持可扩展的自我改进。
英文摘要
Modern large language models (LLMs) can translate natural-language descriptions into operations research (OR) formulations. Post-training techniques including reinforcement learning and on-policy self-distillation have further improved this capability. However, three limitations remain in training LLMs for OR formulations. First, training commonly relies on synthetic formulations validated by human experts or stronger models, constraining scalable supervision. Second, credit assignment is either coarse or costly: outcome rewards score an entire trajectory without locating the responsible modeling decision, whereas process-level supervision requires an additional evaluator. Third, privileged self-distillation can induce style mismatch by using solver context unavailable at deployment. We find that a model can improve from solver-artifact feedback generated by its own rollouts, making self-distillation a practical, evaluator-free source of dense supervision. Therefore, we propose SOLID: Solver-Informed On-Policy LearnIng through Self-Distillation, a novel framework for self-improving OR language models without verified answers or external evaluators. SOLID executes candidate programs from multiple rollouts, clusters their objectives, and selects a majority-group artifact as a pseudo-reference. The model then performs updates using group-relative advantages and dense self-supervision signals. Across multiple OR benchmarks, SOLID improves solution accuracy for both general-purpose and OR-tuned models over outcome-only group-relative training. These results show that solver artifacts can support scalable self-improvement without trusted answers.
发表机构
- Shanghai Jiao Tong University(上海交通大学)
- Fudan University(复旦大学)
机构由 AI 辅助整理,请以论文原文为准。