通过自主多智能体展开实现复制系统中的恢复控制
Recovery Control in Replicated Systems through Autonomous Multiagent Rollout
浏览论文内容
中文总结 AI 辅助
研究复制计算系统的恢复控制问题,将其建模为多智能体结构的POMDP,利用多智能体展开方法及预计算信令信息近似最优控制策略,实验证明该方法可扩展到70个副本系统并降低成本。
中文摘要 AI 辅助
我们研究复制计算系统中的恢复控制。此类系统由副本组成,共同为客户端群体提供服务。这种冗余使系统能承受故障,前提是故障副本的恢复速度快于新故障出现的速度。我们表明,决定何时启动选定副本恢复的问题可表述为具有多智能体结构的部分可观测马尔可夫决策问题(POMDP)。我们利用此结构应用多智能体展开方法来近似最优控制策略。我们的方法使用预先计算的信令信息,减少了副本协调需求并便于并行计算。实验表明,我们的方法可扩展到多达70个副本的系统,且与当前实际使用的恢复策略相比降低了成本。
英文摘要
We study recovery control in replicated computing systems. Such systems consist of replicas that collectively provide a service to a client population. This redundancy enables the system to withstand failures provided that failed replicas are recovered faster than new failures occur. We show that the problem of deciding when to initiate recovery of selected replicas can be formulated as a partially observable Markov decision problem (POMDP) with a multiagent structure. We exploit this structure to apply a multiagent rollout method for approximating optimal control policies. Our method uses precomputed signaling information that reduces the need for replica coordination and facilitates parallel computations. Experiments show that our method scales to systems with up to 70 replicas and reduces costs compared to the recovery policies currently used in practice.