发表机构
The Chinese University of Hong Kong(香港中文大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
PARSER通过并行子智能体读取和主智能体迭代散射-聚集推理解耦读取与推理,用强化学习优化,在多跳问答中显著提升精度并降低延迟。
AI 中文摘要
顺序记忆智能体通过逐块读取长文档并维护紧凑的记忆状态来处理文档,将文档遍历与推理深度耦合在一起。这种耦合导致对证据位置的敏感性,并使推理延迟与文档长度线性相关。我们提出PARSER,将读取与推理解耦。一组轻量子智能体各自绑定一个块,并行读取整个文档,而主智能体通过迭代的散射-聚集轮次进行深度推理:在每一轮中,它向所有子智能体广播查询,聚合返回的证据,并根据迄今发现的信息制定更深入的后续查询。这种解耦设计将所有可学习行为集中在主智能体中,通过强化学习进行优化,而子智能体保持为冻结的现成模型。在上下文长度从7K到896K token的多跳问答任务中,使用4B骨干网络的PARSER平均比最强的顺序记忆基线高出5.7个百分点,在896K token时高出12.0个百分点。扩展到9B骨干网络后,PARSER比DeepSeek-V4-Pro高出6.3个百分点。控制实验证实,PARSER对证据位置、顺序和距离的扰动具有鲁棒性,而这些条件会导致顺序方法出现较大的精度波动,同时将推理延迟降低多达11倍。
英文摘要
Sequential memory agents process long documents by reading chunks one after another while maintaining a compact memory state, coupling document traversal to reasoning depth. This coupling introduces sensitivity to evidence placement and ties inference latency linearly to document length. We introduce PARSER, which decouples reading from reasoning. A bank of lightweight subagents each bound to a single chunk read the entire document in parallel, while a lead agent reasons in depth through iterative scatter--gather rounds: at each round it broadcasts a query to all subagents, aggregates the returned evidence, and formulates a deeper follow-up query conditioned on what has been found so far. This decoupled design concentrates all learnable behavior in the lead agent, which is optimized with reinforcement learning, while the subagents remain frozen off-the-shelf models. On multi-hop QA with contexts ranging from 7K to 896K tokens, PARSER with a 4B backbone outperforms the strongest sequential memory baseline by 5.7 points on average and by 12.0 points at 896K tokens. Scaling to a 9B backbone, PARSER surpasses DeepSeek-V4-Pro by 6.3 points. Controlled experiments confirm that PARSER is robust to perturbations in evidence position, order, and distance, conditions that cause large accuracy swings in sequential methods, while reducing inference latency by up to 11x.