arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

HERO:用于联邦持续学习的异构感知基准库

HERO: A Heterogeneity-Aware Benchmark Library for Federated Continual Learning

Thinh T. H. Nguyen, Le-Tuan Nguyen, Minh-Duong Nguyen, Nhi Trinh, Anh Tran Nam Nguyet, Dung D. Le, Kok-Seng Wong

arXiv 2607.08784首次发表:更新:

发表机构

VinUniversity(文大教育集团)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对联邦持续学习评估难的问题,提出HERO异构感知基准库,通过分离任务划分等构建基准流,在CIFAR-100等上评估FCL方法,揭示方法行为变化等问题,还进行案例研究,发布多种资源支持可重复及设置感知的FCL评估。

AI 中文摘要

联邦持续学习(FCL)评估分布式客户端如何在保留先前学习知识的同时从不断变化的数据流中学习。现有评估因常同时改变数据集、任务划分、客户端数据划分、任务顺序、主干、内存假设和报告规则而难以比较。我们引入了HERO,一个用于FCL的异构感知基准库。HERO通过分离通常耦合的三个选择来构建基准流,即任务划分、客户端数据划分和客户端任务序列。在主要可比基准HERO-Core中,α控制客户端数据偏差,ρ控制任务顺序不匹配。我们使用最终平均准确率、平均遗忘率和底部10%客户端准确率在CIFAR-100和TinyImageNet上评估代表性FCL方法。还包括在OGB-MolPCBA上基于图的Domain-IL可移植性案例研究。结果表明方法行为在简单和异构设置中会变化,平均准确率可能掩盖底部客户端的弱性能,任务顺序不匹配有利于与同步评估不同的策略,相同的HERO接口可揭示基于图像的FCIL之外的域转移难度。HERO发布基准流、配置、方法实现和报告脚本以支持可重复和设置感知的FCL评估。

英文摘要

Federated continual learning (FCL) evaluates how distributed clients learn from changing data streams while retaining previously learned knowledge. Existing evaluations are difficult to compare because they often change datasets, task splits, client data splits, task orders, backbones, memory assumptions, and reporting rules simultaneously. We introduce \textbf{HERO}, a heterogeneity-aware benchmark library for FCL. HERO builds benchmark streams by separating three choices that are often coupled, namely the task split, the client data split, and the client task sequence. In HERO-Core, the main comparable benchmark, $α$ controls client data skew and $ρ$ controls task-order mismatch. We evaluate representative FCL methods on CIFAR-100 and TinyImageNet using final average accuracy, average forgetting, and bottom-10\% client accuracy. We also include a graph-based Domain-IL portability case study on OGB-MolPCBA, where scaffold-domain granularity changes the input distribution while the prediction task remains fixed. Our results show that method behavior changes across easy and heterogeneous settings, that average accuracy can hide weak bottom-client performance, that task-order mismatch favors different strategies from synchronized evaluation, and that the same HERO interface can expose domain-shift difficulty beyond image-based FCIL. HERO releases benchmark streams, configurations, method implementations, and reporting scripts to support reproducible and setting-aware FCL evaluation.

Comments30 pages, 10 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑