发表机构
University of Illinois Urbana-Champaign; Imperial College London; City University of Hong Kong(伊利诺伊大学厄巴纳-香槟分校; 伦敦帝国理工学院; 香港城市大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对现有基于文本的过程奖励模型在长多智能体展开中评分成本高的问题,提出KV-PRM,通过读取KV缓存降低评分成本,经理论证明和实验验证,该模型在多基准测试及多种TTS方法下表现优异,大幅减少计算资源消耗。
AI 中文摘要
过程奖励模型(PRMs)已被证明在指导测试时扩展(TTS)方法方面非常有效,可显著提升基于语言模型的多智能体系统的能力。然而,现有PRMs是基于文本的,在长多智能体展开中,评分成本随序列长度L呈二次增长,造成严重计算瓶颈。为此,我们引入KV-PRM,一种高效的过程奖励模型,通过直接读取语言模型生成阶段自然产生的KV缓存来消除繁重的文本重新编码。通过针对预先存在的KV缓存处理单个“验证令牌”,KV-PRM将评分成本从O(L^2)降低到O(L)。我们正式证明KV缓存包含比文本更大的信息容量,且对下游奖励建模更高效。在经验上,在MATH、GSM8K和AIME基准测试中,KV-PRM在各种TTS方法(如波束搜索、蒙特卡洛树搜索和加权投票)下与文本PRMs匹配或严格优于它们,与基于文本的PRMs相比,评分浮点运算减少高达5000倍,延迟减少37倍,每序列内存占用减少34倍。
英文摘要
Process Reward Models (PRMs) have been proven to be highly effective in guiding test-time scaling (TTS) methods, which significantly boost the capabilities of LLM-based multi-agent systems. However, existing PRMs are text-based: they re-encode the entire trajectory text from scratch. In long multi-agent rollouts, the scoring cost, growing quadratically with respect to sequence length L, creates a severe computational bottleneck, severely limiting PRMs' application in long-context scenarios. To resolve this, we introduce KV-PRM, a highly efficient process reward model that eliminates the heavy text re-encoding by directly reading the KV cache produced naturally during the LLM's generation phase. By processing a single "verify token" against the pre-existing KV cache, KV-PRM reduces the scoring cost from O(L^2) to O(L). We formally prove that the KV cache contains strictly greater information capacity than text, and is more efficient for downstream reward modeling. Empirically, across the MATH, GSM8K, and AIME benchmarks, KV-PRM matches or strictly outperforms text-PRMs under various TTS methods such as Beam Search, MCTS, and Weighted Voting, with up to a 5,000x reduction in scoring FLOPs, a 37x reduction in latency, and a 34x reduction in per-sequence memory footprint compared to text-based PRMs.