arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.02957cs.LGcs.CL

理解联邦世界模型学习中的轨迹异质性

Understanding Trajectory Heterogeneity in Federated World Model Learning

Yipan Wei, Zhaokun Yan, Ziming Hong, Jiaqi Wu, Lixu Wang

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过MIMIC-IV临床预测基准,揭示联邦世界模型中轨迹异质性对时间覆盖、算法误差和更新行为的影响,提出多维评估框架。

中文摘要 AI 辅助

世界模型从轨迹中学习状态演化,这使得对时间上下文的访问成为核心训练要求。联邦学习可以利用分布式记录,而轨迹内部的归属边界限制了每个客户端可以构建的样本。本研究通过八个MIMIC-IV疾病队列上的每小时动作条件临床预测来基准测试这种跨时间设置,共包含4087万个转移成员资格。我们指定了基于严重程度的客户端归属、患者分离构建、本地历史和未来窗口规则,以及从1小时到32小时的配对回滚评估。一个由十种联邦算法组成的矩阵覆盖了32种疾病-分区配置,在五轮10%参与率下进行。从现有结果和训练日志中得出三个发现。第一,客户端归属和参与共同限制了长窗口覆盖:在池化可用的32步窗口中,平均仅有7.55%至21.36%的窗口在训练期间具有本地完整的锚点访问。第二,在16个配对比较中,更细的严重程度分区伴随更高的FedAvg误差,而算法增益较小且依赖于预测范围:FedProx将平均误差降低了0.56%,在32步时没有一致的改进。第三,算法标签掩盖了不同的更新行为,包括不活跃的外推和更新规模的量级差异。在相同的基准协议下,缓存更新性能在不同轨迹分区之间也存在显著差异。这些结果确立了时间访问、参与覆盖、优化行为和范围分辨预测作为评估联邦临床世界模型的互补维度。

英文摘要

World models learn state evolution from trajectories, making access to temporal context a central training requirement. Federated learning can use distributed records, while ownership boundaries within a trajectory restrict the examples each client can construct. Our study benchmarks this cross-time setting through hourly action-conditioned clinical prediction on eight MIMIC-IV disease cohorts, comprising 40.87 million transition memberships. We specify severity-based client ownership, patient-separated construction, local history and future-window rules, and paired rollout evaluation from one to 32 hours. A matrix of ten federated algorithms covers 32 disease--partition configurations under five rounds of ten-percent participation. Three findings emerge from existing results and training logs. First, client ownership and participation jointly restrict long-window coverage: only 7.55\%--21.36\% of pooled-available 32-step windows have a locally complete anchor visited during training, averaged across diseases. Second, finer severity partitions accompany higher FedAvg error in 15 of 16 paired comparisons, while algorithm gains are small and horizon-dependent: FedProx reduces mean error by 0.56\%, with no consistent improvement at 32 steps. Third, algorithm labels conceal distinct update behavior, including inactive extrapolation and orders-of-magnitude differences in update scale. Cached-update performance also varies strongly across trajectory partitions under the same benchmark protocol. These results establish temporal access, participation coverage, optimization behavior, and horizon-resolved prediction as complementary dimensions for evaluating federated clinical world models.

发表机构

  • Wuhan University(武汉大学)
  • China Academy of Information and Communications Technology(中国信息通信研究院)
  • The University of Sydney(悉尼大学)
  • Tsinghua University(清华大学)
  • The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))

机构由 AI 辅助整理,请以论文原文为准。

↑