arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.27819cs.LGcs.DC

当适应有害:联邦可穿戴设备入职中的分裂敏感性与个体级负迁移

When Adaptation Hurts: Split Sensitivity and Person-Level Negative Transfer in Federated Wearable Onboarding

  • Jamia Hamdard(贾米亚哈姆达德大学)
  • Homi Bhabha National Institute(霍米·巴巴国家研究所)
  • Variable Energy Cyclotron Centre(可变能量回旋加速器中心)
  • Gargi Memorial Institute of Technology(加尔吉纪念理工学院)
  • Maulana Abul Kalam Azad University of Technology(毛拉纳·阿布·卡拉姆·阿扎德技术大学)

机构由 AI 辅助整理,请以论文原文为准。

Rahil Aftab, Vineet Kumar Rakesh, Soumya Mazumdar, Tapas Samanta

AI总结:

本研究评估联邦可穿戴设备入职策略,发现个体级负迁移导致平均准确率掩盖失败,提供可审计基准而非普遍优越性声明。

AI中文摘要:

联邦可穿戴模型最终服务于源训练中缺失的人群,但良好的平均准确率并不能证明无标签入职对每个人都有效。我们在泄漏控制的协议下,在五个可穿戴数据集上评估了六种核心入职策略,该协议固定源检查点,仅从源数据估计归一化,将校准与评估记录分离,并对留出的人群而非窗口、设备或随机种子进行推理。完成所有符合条件的HHAR和PAMAP2外部人群轮换,实质性地改变了从原始冻结折中获得的结论。在HHAR上,该单人的平衡准确率在各种方法下为95.6-97.2%,而所有九名用户的准确率为78.3-83.0%,降低了13.8-17.8个百分点(pp)。显示的平均领先者在两个数据集上均发生变化,而配对领先者-运行者自举区间包含零,无法确定更优方法。没有自适应核心机制能在所有五个数据集中结合正的平均增益,且种子平均的个体级损失大于2个百分点(pp)为零。FedBN有一个这样的损失,ATP式适应有八个;仅特征方法在种子平均后没有,但其精确的单侧95%上界为7.6%。一项互补的种子-个体压力审计记录了FedBN、ATP式和仅特征方法在114次中分别有4、22和10次有害实现;这些是重复实现,而非独立参与者。尾部质量、校准可用性和跌倒窗口特异性揭示了平均准确率隐藏的额外失败。因此,该研究提供了一个可审计的开发基准和失败图谱,而非普遍优越性或部署安全性的声明。

英文摘要:

Federated wearable models eventually serve people absent from source training, but favorable average accuracy does not establish that unlabeled onboarding helps each person. We evaluate six core onboarding strategies on five wearable datasets under a leakage-controlled protocol that fixes source checkpoints, estimates normalization from source data only, separates calibration from evaluation recordings, and performs inference over held-out people rather than windows, devices, or random seeds. Completing all eligible HHAR and PAMAP2 outer-person rotations materially changes the conclusion obtained from the original frozen fold. On HHAR, balanced accuracy on that single person is 95.6-97.2% across methods versus 78.3-83.0% over all nine users, a reduction of 13.8-17.8 percentage points (pp). The displayed mean leader changes on both datasets, while paired leader-runner bootstrap intervals include zero and do not resolve a superior method. No adaptive core mechanism combines positive mean gain in all five datasets with zero seed-averaged person-level losses greater than 2 percentage points (pp). FedBN has one such loss and ATP-style adaptation has eight; Feature-only has none after seed averaging, but its exact one-sided 95% upper bound is 7.6%. A complementary seed-person stress audit records 4, 22, and 10 harmful realizations out of 114 for FedBN, ATP-style, and Feature-only, respectively; these are repeated realizations, not independent participants. Tail quality, calibration availability, and fall-window specificity reveal additional failures hidden by mean accuracy. The study therefore provides an auditable development benchmark and failure map rather than a universal-superiority or deployment-safety claim.

补充信息

↑