IAPRepair:面向纠删码存储系统的网络内聚合增强型主动修复
IAPRepair: In-Network Aggregation Enhanced Proactive Repair for Erasure-Coded Storage System
查看机构详情
- Guilin University of Electronic Technology(桂林电子科技大学)
- Guilin University of Aerospace Technology(桂林航天工业学院)
- Jiaxing University(嘉兴大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
IAPRepair是面向纠删码存储的网络内聚合增强型主动修复框架,通过联合构建修复批次等设计,在多场景下较现有方法至少缩短47.83%的修复时间,提升了带宽异构下的修复效率。
中文摘要 AI 辅助
纠删码存储提供容错能力,其存储开销远低于全复制,但修复丢失或存在风险的块需要大量跨节点数据传输。现有被动修复方案仅在故障发生后启动,而主动修复方案可在故障前迁移数据,但通常将迁移、重构和网络聚合视为松散耦合的操作。因此,接收端瓶颈、异构可用带宽以及可编程交换机的有限状态仍会限制修复并行性。本文提出IAPRepair,一种面向纠删码存储的网络内聚合增强型主动修复框架。IAPRepair联合构建每个修复批次,根据归一化传输负载分配重构提供方和替换节点,围绕剩余接收容量调度迁移,并选择性地对可能使健康节点过载的重构块启用网络内聚合。这种选择性设计在保留迁移并行性并尊重可配置交换机资源预算的同时,减少了接收端流量。我们使用Tofino可编程交换机和16个存储节点实现IAPRepair,并通过原型测试床和大规模仿真对其进行评估。在不同编码参数、块大小、节点数量以及多STF节点场景下,与所评估的最先进方法相比,IAPRepair将修复时间至少缩短了47.83%。结果表明,将主动修复决策与网络内处理相协调,是在带宽异构条件下提高修复效率的有效途径。
英文摘要
Erasure-coded storage provides fault tolerance with substantially lower storage overhead than full replication, but repairing lost or at-risk blocks requires intensive cross-node data transfer. Existing reactive repair schemes start only after a failure, while proactive schemes can move data before failure but commonly treat migration, reconstruction, and network aggregation as loosely coupled operations. As a result, receiver?side bottlenecks, heterogeneous available bandwidth, and limited programmable-switch state continue to constrain repair paral?lelism. This paper presents IAPRepair, an in-network aggre?gation enhanced proactive repair framework for erasure-coded storage. IAPRepair jointly constructs each repair batch, assigns reconstruction providers and replacement nodes according to normalized transmission loads, schedules migration around the remaining receive capacity, and selectively enables in-network aggregation for reconstruction blocks that would otherwise over?load healthy nodes. The selective design reduces receiver-side traffic while retaining migration parallelism and respecting a configurable switch-resource budget. We implement IAPRepair with a Tofino programmable switch and 16 storage nodes, and evaluate it using both a prototype testbed and large-scale simu?lations. Across coding parameters, block sizes, node populations, and multiple STF-node scenarios, IAPRepair reduces repair time by at least 47.83% compared with the evaluated state-of-the-art methods. The results demonstrate that coordinating proactive repair decisions with in-network processing is an effective way to improve repair efficiency under bandwidth heterogeneity.