arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.16891cs.LGcs.AImath.OC

具有重新定位、时钟约束资源的实时卡车满载投标接受的认证差距双价格策略

Certified-Gap Dual-Price Policies for Real-Time Truckload Bid Acceptance with Relocating, Clock-Constrained Resources

  • Bubba AI(Bubba人工智能公司)

机构由 AI 辅助整理,请以论文原文为准。

Aswin Chandrasekaran

AI总结:

研究实时卡车满载投标接受决策问题,基于拉格朗日松弛构建双价格策略,证明其有效性、渐近最优性及证书局限性,实验表明该策略在基准测试中表现良好,决策速度快,证书稳定。

AI中文摘要:

卡车运输公司必须在几秒钟内接受或拒绝每个货物投标。该决策取决于车队状态、服务时长(HOS)时钟和预约窗口。我们将此建模为一个弱耦合动态规划,其中资源会重新定位并携带时钟:服务一个请求会将卡车转移到新市场并消耗其时钟,而一辆卡车是否能服务一个请求取决于其状态。基于占用的可重复使用资源模型无法涵盖此设置。我们从给出问题上界的相同拉格朗日松弛构建了一个实时双价格策略。策略和边界来自同一个对象,所以每次运行都会报告一个认证最优性差距。我们证明了三点。首先,该证书对任何对偶、任何离散化和任何替代质量都是有效的。其次,策略的同时空间梯度规则恰好是流体互补松弛,并且该策略在亚临界流体状态下是渐近最优的;通过线性规划基稳定性,拟合价格在样本路径之间也是可移植的。第三,证书有局限性:在每个车队规模下,每个资源的拉格朗日松弛可以保持远离零。我们展示了一个具有精确有理证书和复制引理的三卡车内核。在一个有三十对种子的公共闭环基准测试中,该策略——无需展开标签,只需一次离线对偶求解——在三种情况中的两种情况下击败了展开训练的替代策略(紧:+2.0个百分点,95%置信区间[+0.5,+3.6],威尔科克森p = 0.023;温和:+3.5个百分点,置信区间[+2.4,+4.5]),并在第三种情况下持平。它在0.04 - 0.09毫秒内做出决策,比蒙特卡罗展开教师快三个数量级。其证书在每个场景的十个有界实例中是稳定的,达到最优值的57 - 64%,与慢1000倍的教师认证值相差3 - 6个百分点。

英文摘要:

A truckload carrier must accept or reject each load tender within seconds. The decision depends on fleet state, hours-of-service (HOS) clocks, and appointment windows. We model this as a weakly coupled dynamic program in which the resources relocate and carry clocks: serving a request moves the truck to a new market and depletes its clocks, and whether a truck can serve a request depends on its state. Occupancy-based reusable-resource models do not cover this setting. We build a real-time dual-price policy from the same Lagrangian relaxation that gives the problem's upper bound. Policy and bound come from one object, so every run reports a certified optimality gap. We prove three things. First, the certificate is valid for any duals, any discretization, and any surrogate quality. Second, the policy's same-time spatial-gradient rule is exactly fluid complementary slackness, and the policy is asymptotically optimal in the subcritical fluid regime; the fitted prices are also portable across sample paths, by linear-programming basis stability. Third, certificates have limits: per-resource Lagrangian slack can stay bounded away from zero at every fleet size. We exhibit a three-truck kernel with an exact rational certificate and a replication lemma. On a public closed-loop benchmark with thirty paired seeds, the policy -- which needs no rollout labels, only one offline dual solve -- beats a rollout-trained surrogate on two of three scenarios (tight: +2.0 pp, 95% CI [+0.5, +3.6], Wilcoxon p = 0.023; mild: +3.5 pp, CI [+2.4, +4.5]) and ties the third. It decides in 0.04-0.09 ms, three orders of magnitude faster than the Monte Carlo rollout teacher. Its certificates are stable across ten bounded instances per scenario, at 57-64% of optimal, within 3-6 points of what the 1000x-slower teacher certifies.

补充信息

↑