发表机构
Rutgers University; The University of Hong Kong(罗格斯大学; 香港大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对隐藏规则游戏中Transformer智能体的可解释性,通过在其决策令牌嵌入上训练稀疏自编码器,在双规则任务中恢复结构,单个维度对应可解释策略,为解释智能体行为提供方法。
AI 中文摘要
解释学习到的决策系统的一个核心挑战是确定其内部表示是否包含有助于解释其行为的概念。我们报告了针对隐藏规则游戏(GOHR)中一个分词自回归Transformer智能体的可解释性实验。我们专注于一个紧凑的双规则任务,其中两个隐藏规则都将对象形状映射到目标桶,但排列不同。策略在从这两个隐藏规则采样的情节上进行训练,然后用固定权重进行评估。它从未被给予规则标签,也不使用显式规则分类器;任何规则信息都必须从交互历史中隐式推断。在这种设置下,在智能体尝试一个信息丰富的动作并观察接受/拒绝反馈之前,正确的规则是无法识别的。在智能体的决策令牌嵌入上训练的稀疏自编码器(SAE)恢复了这种结构。当通过诸如所选形状或桶等简单概念对留出的决策进行标记时,对某个概念具有高度选择性的SAE维度涵盖了该概念出现的大多数决策。单个SAE维度还对应于可解释的策略,例如探测一个规则假设并在负面反馈后切换。
英文摘要
A central challenge in interpreting learned decision-making systems is to determine whether their internal representations contain concepts that help explain their behavior. We report interpretability experiments for a tokenized autoregressive Transformer agent in the Game of Hidden Rules (GOHR). We focus on a compact two-rule task in which both hidden rules map object shapes to target buckets, but with different permutations. The policy is trained on episodes sampled from these two hidden rules and then evaluated with fixed weights. It is never given a rule label and does not use an explicit rule classifier; any rule information must be inferred implicitly from interaction history. In this setting, the correct rule is not identifiable before the agent tries an informative move and observes accept/reject feedback. Sparse autoencoders (SAEs) trained on the agent's decision-token embeddings recover this structure. When held-out decisions are labeled by simple concepts such as the chosen shape or bucket, SAE dimensions that are highly selective for a concept cover most decisions where that concept is present. Individual SAE dimensions also correspond to interpretable strategies such as probing one rule hypothesis and switching after negative feedback.