AI 中文总结
FEDWORLD是一种感知范围的智能体世界模型联邦协议,通过交换结构化抽象转移规则,减少冲突动态下的负迁移,提升智能体任务成功率。
AI 中文摘要
大型语言模型(LLM)智能体从本地交互经验中学习世界动态,以支持后续规划与动作选择。然而,单个客户端可用的经验往往不完整,这促使客户端之间共享知识。现有联邦方法主要聚合模型参数,而智能体记忆共享方法通常汇集轨迹、记忆或规则,却未检查这些内容对每个客户端是否仍有效。这一假设存在问题,因为相同的抽象动作在不同策略、环境或异常条件下可能产生不同效果。因此,多数客户端支持的规则可能覆盖少数客户端持有的正确知识。为解决该问题,我们提出FEDWORLD,一种感知范围的世界模型联邦协议,该协议交换结构化抽象转移规则。每个客户端将私有转移转换为标准化规则,服务器对齐相关规则以识别每个规则在客户端间的支持与矛盾证据。所得证据决定规则是共享、集群特定、私有还是未解决。每个目标客户端保留本地规则,仅对推断范围兼容的未覆盖情况接受联邦更新;模糊规则不予采用。在τ-bench和ALFWorld上的实验表明,FEDWORLD减少了冲突动态下的负迁移,同时保留了有用的跨客户端迁移,进而减少了状态回退、重复动作和多余步骤,提高了任务成功率。
英文摘要
Large language model (LLM) agents learn world dynamics from local interaction experience to support subsequent planning and action selection. However, the experience available to a single client is often incomplete, which motivates sharing knowledge across clients. Existing federated methods mainly aggregate model parameters, while agent memory-sharing methods commonly pool trajectories, memories, or rules without checking whether they remain valid for each client. This assumption is problematic because the same abstract action may produce different effects under different policies, environments, or exception conditions. Consequently, a rule supported by most clients may overwrite correct knowledge held by a minority client. To address this problem, we propose FEDWORLD, a scope-aware federated world-model protocol that exchanges structured abstract transition rules. Each client converts private transitions into normalized rules, and the server aligns related rules to identify each rule supporting and contradicting evidence across clients. The resulting evidence determines whether a rule is shared, cluster-specific, private, or unresolved. Each target client retains its local rules and accepts federated updates only for uncovered cases whose inferred scope is compatible; ambiguous rules are withheld. Experiments on $τ$-bench and ALFWorld show that FEDWORLD reduces negative transfer under conflicting dynamics while retaining useful cross-client transfer, leading to fewer state regressions, repeated actions, and excess steps, as well as higher task success.