发表机构
Temple University; Dickinson College(天普大学; 迪金森学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究序列模型在部分可观测下如何表示潜在随机动态,发现在随机波动率设置中有两阶段计算,Transformer中潜态可解码性在特定阶段出现,输出头替换揭示部分退化原因,为机械可解释性提供有用基准。
AI 中文摘要
机械可解释性主要集中在语言模型和确定性玩具任务上。对于序列模型如何在有噪声、部分可观测的情况下内部表示潜在随机动态知之甚少。我们在一个可控的多变量随机波动率设置中研究这个问题,模型只能观测回报,而研究人员知道真实的潜在波动率状态。这种设置为部分可观测性下的机械可解释性提供了有用的基准。我们发现跨架构、损失和输出头存在两阶段计算的证据。在Transformer中,潜态可解码性在可识别的架构阶段出现,其位置取决于波动率周期。输出头替换表明,噪声MSE训练下的部分退化源于读出失准而非表示失败。这些结果表明,随机波动率模型为有噪声潜态动态和部分可观测性下的机械可解释性提供了有用的基准。
英文摘要
Mechanistic interpretability has largely focused on language models and deterministic toy tasks. Much less is known about how sequence models internally represent latent stochastic dynamics under noisy, partially observed observations. We study this question in a controlled multivariate stochastic volatility setting, where models observe only returns while the ground-truth latent volatility state is known to the researcher. This setting provides a useful benchmark for mechanistic interpretability under partial observability: the latent state is hidden from the model but directly available for evaluation. Across architectures, losses, and output heads, we find evidence for a two-stage computation. Hidden representations encode substantial information about the next latent volatility state, and the output head maps this representation to squared return forecasts. Furthermore, in Transformers, latent-state decodability emerges at identifiable architectural stages whose location depends on the volatility period. In long-cycle regimes, this computation simplifies into an explicit latent-state filter consisting of a learned linear projection followed by $\ell^2$ normalization. Output-head replacement further shows that part of the degradation under noisy MSE training arises from readout misalignment rather than representation failure. These results suggest that stochastic volatility models provide a useful benchmark for mechanistic interpretability under noisy latent dynamics and partial observability.