WFM:面向复杂智能体推理的维基基础模型
WFM: Wiki Foundation Model for Complex Agentic Reasoning
浏览论文内容
中文总结 AI 辅助
针对智能体动态推理中稀疏图表示限制,提出维基基础模型WFM,通过Wiki图模式、查询条件注意力聚合及NCCL边界交换协议,实现可扩展的智能体原生知识表示与检索,在五个基准上性能优异并获10.5倍训练加速。
中文摘要 AI 辅助
现实世界中的智能体从根本上需要持久的非参数知识来进行动态推理,即长期记忆和检索增强生成。尽管图在提供结构化证据方面展现出可靠优势,但稀疏的图表示天然限制了复杂智能体工作流所需的机器可读性和语义密度。受此限制驱动,整个行业正经历从传统稀疏图到LLM Wiki的范式转变,LLM Wiki是一种智能体原生的知识表示,将密集的文档上下文与包含多层拓扑链接的markdown文件相结合。然而,对这种丰富语义进行参数化具有挑战性,因为使用传统的稀疏图嵌入难以编码密集的文本上下文。此外,使用现有的图编码器学习LLM Wiki可能会使分布式系统开销过载,从而阻碍其在大型商业场景中的部署。为此,我们提出了一种新颖的范式——维基基础模型(WFM),专为可扩展的智能体原生表示和检索而设计。具体而言,(i)我们形式化了一个Wiki图模式,无缝桥接细粒度结构与密集上下文,在保持显式拓扑的同时兼具连续语义;(ii)我们定制了一种查询条件注意力聚合机制,用于丰富的维基消息传递和显式的注意力方差正则化;(iii)我们设计了一种基础设施级的NCCL边界交换协议,该协议提升静态分区索引并利用固定形状的GPU到GPU集合通信,绕过了CPU序列化和内存拷贝开销。在五个长期智能体记忆和多跳推理基准上的广泛评估表明,WFM性能卓越,同时在分布式集群上实现了10.5倍的训练加速。
英文摘要
Real-world agents fundamentally require persistent non-parametric knowledge for dynamic reasoning, i.e., long-term memory and retrieval-augmented generation. While graphs have shown reliable advantages in providing structured evidence, the sparse graph representations naturally restrict machine readability and semantic density required for complex agentic workflows. Driven by this limitation, the entire industry is witnessing a paradigm shift from traditional sparse graphs to LLM Wiki, an agent-native knowledge representation that couples dense document contexts with markdown files containing multi-layered topological linkages. However, parameterizing such rich semantics is challenging to encode dense textual contexts using traditional sparse graph embeddings. Moreover, learning LLM Wiki with existing graph encoders could overwhelm distributed system overheads that hinder deployment in large-scale commercial scenarios. To this end, we propose a novel paradigm Wiki Foundation Model, i.e., WFM, tailored for scalable, agent-native representation and retrieval. Specifically, (i) we formalize a Wiki Graph schema that seamlessly bridges fine-grained structures with dense contexts, maintaining explicit topologies alongside continuous semantics; (ii) A query-conditioned attentive aggregation is tailored for rich wiki message passing and explicit attention variance regularization; (iii) We engineer an infrastructural NCCL boundary exchange protocol that hoists static partition indices and leverages fixed-shape GPU-to-GPU collectives, bypassing CPU serialization and memory copy overheads. Extensive evaluations across five long-term agent memory and multi-hop reasoning benchmarks demonstrate the remarkable performance of WFM, while achieving a 10.5 times training acceleration on distributed clusters.
发表机构
- Tencent Youtu Lab(腾讯优图实验室)
- Monash University(蒙纳士大学)
- Hong Kong Baptist University(香港浸会大学)
- Shenzhen University(深圳大学)
机构由 AI 辅助整理,请以论文原文为准。