发表机构
Graduate School of Informatics, Kyoto University(京都大学信息学研究科)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出物理信息多智能体协调框架,将BCMP排队网络嵌入分散式强化学习,优化医院患者流,显著降低延迟并满足临床安全。
AI 中文摘要
高效协调各自主医院部门间的患者流对于缓解拥挤和平衡资源利用至关重要。尽管经典排队理论,特别是开放的Baskett-Chandy-Muntz-Palacios(BCMP)网络,为医疗运营提供了可解释的数学拓扑结构,但分析模型依赖于平稳假设和固定路由矩阵,这些在状态依赖的现实动态下会退化。相反,集中式强化学习方法难以适应医院治理的分散结构,其中各临床部门以局部观测、异构资源和不同的运营目标运作。在本文中,我们提出了一个名为“物理信息多智能体协调”的多智能体系统(MAS)框架,该框架将经验校准的BCMP排队拓扑作为物理先验嵌入到分散式多智能体强化学习架构中。该方法被表述为在耦合资源约束下的分散式部分可观测马尔可夫决策过程(Dec-POMDP),使自主部门智能体能够协作协商患者路由和动态服务扩展。为减轻环境非平稳性而不引起过多通信开销,智能体沿网络边缘交换局部动作指纹,并优化空间分解的奖励结构。基于真实MIMIC-IV患者轨迹的实证评估表明,与静态马尔可夫近似、启发式调度和独立多智能体基线相比,这种协作多智能体方法在维持临床安全约束的同时,显著降低了累积系统延迟。
英文摘要
Efficient patient flow coordination across autonomous hospital departments is critical for mitigating overcrowding and balancing resource utilization. While classical queueing theory, specifically open Baskett--Chandy--Muntz--Palacios (BCMP) networks, provides an interpretable mathematical topology for healthcare operations, analytical models rely on stationary assumptions and fixed routing matrices that degrade under state-dependent real-world dynamics. Conversely, centralized reinforcement learning approaches struggle to accommodate the decentralized structure of hospital governance, where individual clinical departments function with localized observations, heterogeneous resources, and divergent operational objectives. In this paper, we present a Multi-Agent Systems (MAS) framework titled \emph{Physics-Informed Multi-Agent Coordination}, which embeds empirically calibrated BCMP queueing topologies as physical priors within a decentralized multi-agent reinforcement learning architecture. Formulated as a Decentralized Partially Observable Markov Decision Process (Dec-POMDP) under coupled resource constraints, our method enables autonomous departmental agents to cooperatively negotiate patient routing and dynamic service scaling. To mitigate environmental non-stationarity without inducing excessive communication overhead, agents exchange localized action fingerprints along network edges and optimize a spatially decomposed reward structure. Empirical evaluations driven by real-world MIMIC-IV patient trajectories indicate that this cooperative multi-agent approach substantially reduces cumulative system delay compared to static Markovian approximations, heuristic dispatching, and independent multi-agent baselines, while maintaining clinical safety constraints.
CommentsAccepted to PRIMA 2026 (Full paper). 16 pages, 2 figures, 3 tables