多阶段库存控制的信息极限:学习、估值与删失
Information Limits of Multistage Inventory Control: Learning, Valuation, and Censoring
浏览论文内容
中文总结 AI 辅助
本文针对离线多阶段库存控制,建立信息论下界,区分策略学习与成本估值所需信息,并给出删失数据下的学习率及覆盖条件。
中文摘要 AI 辅助
学习如何补充库存与估计其成本所需的信息不同。我们针对具有独立有界需求的离线多阶段库存控制建立了信息论下界。对于总期望成本的固定加性误差,在最坏情况下,即使在全观测和自适应样本分配下,策略学习在时间跨度上需要立方数量级的标量观测。这与已知的经验风险上阶相匹配。平稳性允许二次方的策略学习保证,然而即使最优策略已知,带有继承库存的估值仍可能保持立方级别的困难。删失强化了这一区别:最优决策可以学习,而绝对成本仍无法识别。我们给出了一个携带安全覆盖条件,在该条件下截断能精确保持库存转移和策略间隙。将此简化与匹配的下界相结合,得到了尖锐的删失学习率,原始样本需求与销售日志可用部分成反比。相同的合格观测可以在联合误差保证下检查覆盖并拟合策略。这些结果在显式的条件独立日志记录协议下,分离了决策和估值的信息极限。
英文摘要
Learning how to replenish inventory and estimating the cost of doing so require different information. We establish information-theoretic lower bounds for offline multistage inventory control with independent bounded demand. For fixed additive error in total expected cost, policy learning requires cubically many scalar observations in the horizon in the worst case, even under full observation and adaptive sample allocation. This matches the known empirical-risk upper order. Stationarity permits a quadratic policy-learning guarantee, yet valuation with inherited stock can remain cubically difficult even when the optimal policy is known. Censoring sharpens this distinction: optimal decisions can be learnable while absolute costs remain unidentified. We give a carrysafe coverage condition under which truncation preserves inventory transitions and policy gaps exactly. Combining this reduction with a matching lower bound yields the sharp censored-learning rate, with raw sample requirements inversely proportional to the usable fraction of sales logs. The same qualified observations can check coverage and fit the policy under a joint error guarantee. The results isolate information limits for decisions and valuation under an explicit conditionally independent logging protocol.
发表机构
- National University of Singapore(新加坡国立大学)
- Massachusetts Institute of Technology(麻省理工学院)
- Cornell SC Johnson College of Business(康奈尔SC约翰逊商学院)
机构由 AI 辅助整理,请以论文原文为准。