arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

剖析层级推理模型:一项机制性研究

Dissecting Hierarchical Reasoning Models: A Mechanistic Study

Leo Raphael Rodrigues, Jian Kang

arXiv 2609.22197首次发表:更新:

发表机构

Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过因果干预、线性探针和稀疏自编码器等方法,机制性地剖析了层级推理模型在数独等任务上的推理过程,发现其本质是约束感知的迭代细化,且不依赖紧凑的因果特征集。

AI 中文摘要

我们研究了层级推理模型(HRM),这是一种具有代表性的、基于Transformer的层级潜在推理模型,包含多种变体,并在数独、迷宫和ARC-AGI-2任务上进行了评估。我们从机制层面理解了HRM如何进行推理以及它编码了哪些信息。我们的分析将HRM与带有和不带有循环模块的Transformer基线进行了比较,对循环状态施加了因果干预,并利用线性探针与随机方向消融进行对比,同时使用了带有特征消融的稀疏自编码器。我们的结果揭示了几个关键发现:循环模型优于单次通过基线,而单状态循环Transformer与HRM性能相当。状态干预进一步表明,高层和低层状态的因果贡献因任务特定检查点和推理阶段而异。选定的任务变量可以从HRM的循环状态中线性解码,然而,消融探针方向产生的效果与随机对照相当。SAE消融比探针方向消融产生了更大的行为变化。然而,排名靠前的SAE特征在较大消融规模或跨任务时,并未显示出相对于大小匹配的随机子集的稳定优势;同样的模式在带有步内BPTT的数独对照中也持续存在。综合来看,我们表征出HRM本质上是在谜题特定解状态上实现约束感知的迭代细化,其中不同层级组件的功能贡献各不相同,且不依赖于紧凑、因果重要的特征集。这些结果凸显了研究不同工作机制的必要性,以及开发更适合潜在空间递归推理模型的机制可解释性技术的重要性。

英文摘要

We study Hierarchical Reasoning Model (HRM), a representative hierarchical Transformer-based latent reasoning model with many variants, on Sudoku, Maze, and ARC-AGI-2. We mechanistically understand how HRM reasons and what information it encodes. Our analyses compare HRM against Transformer baselines with and without recurrent modules, apply causal interventions on recurrent states, and utilize linear probes against random-direction ablations, as well as sparse autoencoders with feature ablations. Our results reveal several key findings: recurrent models outperform one-pass baselines, while single-state recurrent Transformers are comparable to HRM. State interventions further show that the causal contributions of the high- and low-level states vary across task-specific checkpoints and inference stages. Selected task variables are linearly decodable from the recurrent states in HRM, yet ablating probe directions produce effects comparable to random controls. SAE ablations yield larger behavioral changes than probe-direction ablations. However, top-ranked SAE features show no stable advantage over size-matched random subsets at larger ablation sizes or across tasks; the same pattern persists in a Sudoku control with within-step BPTT. Together, we characterize that HRM is essentially implementing constraint-aware iterative refinement on a puzzle-specific solution state, in which the functional contributions of components at different levels vary without relying on a compact, causally important feature set. These results highlight the necessity of studying the different working mechanisms and the importance of developing mechanistic interpretability techniques better suited for latent-space, recursive reasoning models.

Comments7 pages, 5 figures, Under Review

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑