arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于多仓库库存分配中运筹学公式选择的大语言模型

Large Language Model for Operations Research Formulation Selection in Multi-Warehouse Inventory Allocation

Jintao Xu, Yingzheng Ma, Jiong Dong, Yongzhi Qi, Jianshen Zhang

arXiv 2607.25956首次发表:更新:

AI 中文总结

研究多仓库库存分配中运筹学公式选择问题,提出求解器引导的大语言模型框架,通过构建SFT记录、转换质量差距为偏好等进行训练,实验表明GRPO能显著提高选择准确性和分配质量。

AI 中文摘要

多仓库库存分配通常被表述为混合整数规划(MIP)问题,但没有单一公式能始终匹配由需求集中、库存不平衡、补货规模、服务约束和预测波动性引起的异构实例级情况。我们将此问题作为实例级运筹学(OR)公式选择来研究,为每个分配实例从候选OR专家库中分配一个可由求解器执行的公式。我们提出了一个用于OR公式选择的求解器引导大语言模型(LLM)框架,其中每个OR专家对应一个编码不同分配优先级的MIP公式。为训练选择器,框架首先构建平衡的专家条件监督微调(SFT)记录用于模式学习,然后使用MIP求解器对历史实例进行评估,将求解器评估的分配质量差距转换为边际加权身份偏好优化(IPO)偏好和每个实例的专家分数元数据,以便在组相对策略优化(GRPO)期间进行奖励查找,为采样响应分配奖励。在中国最大的电子商务公司之一京东的多仓库库存分配实例上进行实验表明,GRPO相对于SFT + IPO选择器显著提高了专家选择准确性,更重要的是,产生的实际分配质量高于偏好训练的选择器和最佳固定公式。使用GRPO,命中率@1从21.45%提高到50.42%,命中率@2从70.47%提高到82.31%。所得选择器比现有基线的分配准确率提高了12.57个百分点,优于SFT + IPO选择器和最佳固定OR专家,并将与事后预言机的差距缩小到4.85个百分点。

英文摘要

Multi-warehouse inventory allocation is typically formulated as a mixed-integer programming (MIP) problem, yet no single formulation consistently matches heterogeneous instance-level regimes induced by demand concentration, inventory imbalance, replenishment scale, service constraints, and forecast volatility. We study this issue as instance-wise operations research (OR) formulation selection, where each allocation instance is assigned to a solver-executable formulation from a candidate OR expert library. We propose a solver-guided large language model (LLM) framework for OR formulation selection, in which each OR expert corresponds to a MIP formulation encoding a distinct allocation priority. To train the selector, the framework first constructs balanced expert-conditioned supervised fine-tuning (SFT) records for schema learning, and then uses MIP solver evaluation on historical instances to convert solver-evaluated allocation-quality gaps into margin-weighted identity preference optimization (IPO) preferences and per-instance expert-score metadata for reward lookup during group relative policy optimization (GRPO) to assign rewards to sampled responses. Experiments on multi-warehouse inventory allocation instances from JD$\mathord{.}$com, one of China's largest e-retailers, demonstrate that GRPO substantially improves expert-selection accuracy relative to the SFT+IPO selector and, more importantly, produces higher realized allocation quality than both the preference-trained selector and the best fixed formulation. With GRPO, Hit Ratio@1 and Hit Ratio@2 increase from 21.45% to 50.42% and from 70.47% to 82.31%. The resulting selector achieves an allocation accuracy gain of 12.57 percentage points over the incumbent baseline, outperforming both the SFT+IPO selector and the best fixed OR expert, and reduces the gap to the ex-post oracle to 4.85 percentage points.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑