arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

HSRM:用于测试时验证的隐藏状态奖励模型

HSRM: Hidden-State Reward Models for Test-Time Verification

Xianzhi Li, Xiaodan Zhu

arXiv 2608.30841首次发表:更新:

发表机构

Ingenuity Labs Research Institute(Ingenuity Labs 研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出轻量型隐藏状态奖励模型HSRM,复用生成器内部表征验证数学推理候选解,在四个基准中多数场景性能优于55M参数纯文本验证器,仅需约2M参数,实现高效验证。

AI 中文摘要

大型语言模型(Large Language Models)常能生成看似合理的数学推理过程,但在多个候选解中可靠识别正确解仍是关键挑战。现有测试时推理流水线通常依赖基于文本的验证器,会重新读取每个生成的解,使验证成为推理中开销高昂的环节。不过已有研究表明,大型语言模型会在内部表征中编码与正确性相关的信号,包括对自身答案可能出错的感知。基于该观察,我们提出HSRM,一种轻量型隐藏状态奖励模型,它通过直接读取生成器的内部表征而非重新处理文本,来验证候选解。HSRM在推理步骤边界处从冻结的生成器中提取隐藏状态,并使用小型Transformer编码器对候选解进行排序。它通过自生成的带结果标签的轨迹进行训练,既不需要人工编写的过程监督,也不需要大型预训练验证器。在四个数学推理基准测试中,HSRM在16种生成器-数据集设置里的15种中,性能与55M参数的纯文本能量验证器相当或更优,且仅使用约2M参数,通过复用推理过程中已计算的表征,为纯文本验证提供了一种高效替代方案。

英文摘要

Large language models can often generate plausible mathematical reasoning traces, but reliably identifying the correct solution among multiple candidates remains a key challenge. Existing test-time reasoning pipelines typically rely on text-based verifiers that re-read each generated solution, making verification an expensive component of inference. Prior work has shown, however, that LLMs often encode correctness-related signals in their internal representations, including awareness of when their own answers are likely to be wrong. Building on this observation, we introduce HSRM, a lightweight hidden-state reward model that verifies candidate solutions by directly reading the generator's internal representations rather than re-processing its text. HSRM extracts hidden states from a frozen generator at reasoning-step boundaries and uses a small Transformer encoder to rank candidates. It is trained from self-generated trajectories with outcome labels, requiring neither human-written process supervision nor a large pretrained verifier. Across four mathematical reasoning benchmarks, HSRM matches or outperforms a 55M-parameter text-only energy verifier in 15 of 16 generator--dataset settings while using only about 2M parameters, providing an efficient alternative to text-only verification by reusing representations already computed during generation.

CommentsEMNLP 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑