arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.21512econ.GNq-fin.EC

飞机分组登机:静态策略基准测试与深度强化学习优化动态分配

Group boarding for airplanes: benchmarking static policies and optimizing dynamic assignment with deep reinforcement learning

Minyu Shen, Weihua Gu, Junqi Ma, Boqian Song, Li Zhen, Gang Kou

首次发表
浏览论文内容

中文总结 AI 辅助

研究飞机登机分组问题,提出将登机组分配视为马尔可夫决策过程并用深度强化学习求解的方法,该动态策略在总登机时间和平均个人登机时间上优于静态策略,还能给出近似帕累托前沿,且策略在多种分布外条件下保持稳健。

中文摘要 AI 辅助

提高登机效率可减少飞机周转时间并提升乘客体验。航空公司通常用基于座位的静态规则将乘客分配到几个连续的登机组,但静态规则忽略了已登机乘客的座位信息。本文提出首个登机组分配的动态公式,将其视为马尔可夫决策过程并用强化学习求解。该策略用卷积神经网络编码登机座位分配状态,通过近端策略优化训练,奖励平衡总登机时间和平均个人登机时间。在内部模拟器中,将该强化学习策略与三种同伴兼容的静态策略进行基准测试,结果显示动态策略在所有布局上均优于静态策略。在一个代表性案例中,强化学习策略在总登机时间上比最优的前后顺序策略最多可提高9.8%,在平均个人时间上最多可提高22.8%。调整奖励权重可得到近似帕累托前沿供运营商选择,训练后的策略在不同负载因子、同伴规模和行李负载等分布外操作条件下仍保持稳健。

英文摘要

Improving boarding efficiency reduces airplane turnaround time and improves passenger experience. Airlines typically assign passengers to a few sequential boarding groups using static seat-based rules. Yet arrivals, seat choices, and luggage are sequential and random, and a static rule ignores the seats earlier passengers have already taken. We propose the first dynamic formulation of boarding group assignment. As each passenger checks in, we observe earlier passengers' seats and groups, the current passenger's seat, and optional luggage information, then assign a group while keeping companions together. We formulate dynamic group assignment as a Markov decision process and solve it with reinforcement learning (RL). The policy uses a convolutional neural network to encode the checked-in seat-assignment state and is trained by proximal policy optimization. The reward balances total boarding time and average individual boarding time. We benchmark the proposed RL policy against three companion-compatible static policies (back-to-front, modified Steffen, and alternating block) in an in-house simulator covering six single- and double-aisle layouts. Back-to-front with optimized group sizes achieves the shortest total boarding time and average individual boarding time among the static benchmarks across all layouts. The dynamic RL policy further outperforms it on both metrics in every layout. On a representative case, the RL policy outperforms the optimal back-to-front by up to 9.8\% in total boarding time and 22.8\% in average individual time. Sweeping the reward weight yields an approximate Pareto frontier for operator choice. Trained policies remain robust under out-of-distribution operating conditions, including varying load factors, companion sizes, and luggage loads.

↑