arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于图推理与拓扑感知的多智能体强化学习用于大规模铁路网络管理

Graph-Based Inference and Topology-Aware Multi-Agent Reinforcement Learning for Large-Scale Railway Network Management

Giacomo Arcieri, Gregory Duthé, Christophe Muller, Konstantinos G. Papakonstantinou, Daniel Straub, Eleni Chatzi

arXiv 2609.30150首次发表:更新:

发表机构

ETH Zürich; University of Oxford; The Pennsylvania State University; Technical University of Munich(苏黎世联邦理工学院; 牛津大学; 宾夕法尼亚州立大学; 慕尼黑工业大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对大规模铁路网络管理,提出结合图推理与拓扑感知多智能体强化学习框架,实现零样本迁移,显著优于现有基线。

AI 中文摘要

现代基础设施资产管理构成一个复杂的序列决策问题,其特点是规划周期长以及系统级交互,如空间劣化相关性和规模经济。虽然深度强化学习在优化维护策略方面显示出潜力,但扩展到现实世界网络仍具挑战性。集中式方法在大规模系统中变得计算上难以处理,而分散式方法往往无法捕捉必要的协调机制。为应对这些挑战,我们提出一个基于图的框架,将准确的环境建模与可扩展的决策支持相结合。首先,我们采用一个层次贝叶斯模型,利用图上的高斯过程核从瑞士联邦铁路提供的真实世界数据中推断出铁路维护规划的现实、空间相关的网络环境。其次,我们通过整合图神经网络和图变换器引入一个拓扑感知的多智能体强化学习(MARL)框架,以优化网络级策略。本工作的一个核心贡献是通过零样本迁移学习展示可扩展性:仅在小型网络部分上训练的基于图的智能体,无需任何重新训练,即可零样本方式成功部署于大规模未见过的网络。数值结果表明,所提出的方法显著优于优化启发式算法和标准MARL基线,在保持大规模网络优越性能的同时减少了计算训练时间。

英文摘要

Modern infrastructure asset management constitutes a complex sequential decision-making problem, characterized by long planning horizons and system-level interactions, such as spatial deterioration correlations and economies of scale. While deep reinforcement learning has shown promise in optimizing maintenance policies, scaling to real-world networks remains challenging. Centralized approaches become computationally intractable in large-scale systems, whereas decentralized approaches often fail to capture essential coordination mechanisms. To address these challenges, we propose a graph-based framework that integrates accurate environment modeling with scalable decision support. First, we employ a hierarchical Bayesian model leveraging a Gaussian Process on Graph kernel to infer a realistic, spatially correlated networked environment of railway maintenance planning from real-world data provided by the Swiss Federal Railways. Second, we introduce a topology-aware Multi-Agent Reinforcement Learning (MARL) framework by integrating graph neural networks and graph Transformers to optimize network-level policies. A central contribution of this work is the demonstration of scalability through zero-shot transfer learning: graph-based agents, trained only on small network portions, are successfully deployed in a zero-shot manner on large-scale unseen networks without any retraining. Numerical results indicate that the proposed method significantly outperforms optimized heuristics and standard MARL baselines, reducing computational training time while maintaining superior performance on large-scale networks.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑